Unstructured Data Segmentation for Privacy Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Processing and managing large volumes of unstructured data is time-consuming and resource-intensive, and existing methods struggle to efficiently identify and protect sensitive information while meeting privacy regulations and infrastructure costs.

Innovation Solution

A computer-implementable method for segmenting unstructured data sources involves connecting to data sources, extracting metadata, calculating probability using algorithms, segmenting data based on indicator presence, and assessing to confirm indicators, with the option to train algorithms for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If scanning and indexing large amounts of unstructured data is performed, then data processing completeness is improved, but time consumption and resource usage increase significantly

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidtime to scan and index data
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments unstructured data into structured representations by extracting metadata and identifying indicators, dividing the large data volume into manageable segments that can be processed more efficiently. This allows the system to scan and index only relevant segments rather than processing entire datasets, reducing time consumption while maintaining processing completeness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by extracting metadata and calculating probability indicators before full data processing. This preliminary segmentation and classification allows the system to prioritize and process only the most relevant data portions, significantly reducing the time required for complete scanning and indexing while maintaining comprehensive data processing.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If comprehensive data processing is performed to meet privacy regulations, then data protection compliance is improved, but infrastructure costs and overhead processes increase

Engineering Contradiction:
Improvedata protection complianceVSAvoidinfrastructure cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and isolates specific data segments that require protection under privacy regulations. By identifying and separating sensitive data portions from the larger dataset, the system can apply targeted protection measures only where needed, reducing overall infrastructure complexity and costs while maintaining compliance with privacy regulations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing and protection qualities to different data segments based on their sensitivity and regulatory requirements. High-value or sensitive data receives enhanced protection and processing, while less critical data receives standard handling, optimizing infrastructure resource allocation and reducing overall system complexity while maintaining compliance.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If algorithm training is performed to improve indicator detection accuracy, then detection precision is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveindicator detection accuracyVSAvoidalgorithm training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial algorithm training by training models only on representative samples of data rather than requiring comprehensive training on all data. This partial action approach achieves sufficient detection accuracy for practical purposes while significantly reducing training time and computational resources required compared to full dataset training.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20230214403A1Method and System for Segmenting Unstructured Data Sources for Analysis
Publication Date: 2023.07.06 JETTEE INC
  • US20230214403A1 patent drawing
  • US20230214403A1 patent drawing
  • US20230214403A1 patent drawing

AI summary

Described are computer-implementable method, system and computer-readable storage medium for segmenting data and information of unstructured data sources. The data and information can be documents and files. Connecting is performed to one or more unstructured data sources that store the data and information. Metadata is extracted from the data and information. A probability using one or more algorithms, as to existence of defined indicators of the data and information is performed. Segmentation is performed based on the calculated probability of indicator. Assessing of data and information is performed to confirm the indicators. Training of algorithms if determined is performed.