Unstructured Data Segmentation for Privacy Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing and managing large volumes of unstructured data is time-consuming and resource-intensive, and existing methods struggle to efficiently identify and protect sensitive information while meeting privacy regulations and infrastructure costs.
Innovation Solution
A computer-implementable method for segmenting unstructured data sources involves connecting to data sources, extracting metadata, calculating probability using algorithms, segmenting data based on indicator presence, and assessing to confirm indicators, with the option to train algorithms for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If scanning and indexing large amounts of unstructured data is performed, then data processing completeness is improved, but time consumption and resource usage increase significantly
Solution Approach 1:
The patent segments unstructured data into structured representations by extracting metadata and identifying indicators, dividing the large data volume into manageable segments that can be processed more efficiently. This allows the system to scan and index only relevant segments rather than processing entire datasets, reducing time consumption while maintaining processing completeness.
Solution Approach 2:
The patent performs preliminary actions by extracting metadata and calculating probability indicators before full data processing. This preliminary segmentation and classification allows the system to prioritize and process only the most relevant data portions, significantly reducing the time required for complete scanning and indexing while maintaining comprehensive data processing.
2Reliability
If comprehensive data processing is performed to meet privacy regulations, then data protection compliance is improved, but infrastructure costs and overhead processes increase
Solution Approach 1:
The patent extracts and isolates specific data segments that require protection under privacy regulations. By identifying and separating sensitive data portions from the larger dataset, the system can apply targeted protection measures only where needed, reducing overall infrastructure complexity and costs while maintaining compliance with privacy regulations.
Solution Approach 2:
The patent applies different processing and protection qualities to different data segments based on their sensitivity and regulatory requirements. High-value or sensitive data receives enhanced protection and processing, while less critical data receives standard handling, optimizing infrastructure resource allocation and reducing overall system complexity while maintaining compliance.
3Measurement precision
If algorithm training is performed to improve indicator detection accuracy, then detection precision is improved, but processing time and computational resources increase
Solution Approach 1:
The patent applies partial algorithm training by training models only on representative samples of data rather than requiring comprehensive training on all data. This partial action approach achieves sufficient detection accuracy for practical purposes while significantly reducing training time and computational resources required compared to full dataset training.
Data Source
AI summary
Described are computer-implementable method, system and computer-readable storage medium for segmenting data and information of unstructured data sources. The data and information can be documents and files. Connecting is performed to one or more unstructured data sources that store the data and information. Metadata is extracted from the data and information. A probability using one or more algorithms, as to existence of defined indicators of the data and information is performed. Segmentation is performed based on the calculated probability of indicator. Assessing of data and information is performed to confirm the indicators. Training of algorithms if determined is performed.


