String Frequency Uniqueness Assessment for Sensitive Data Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Entities face challenges in identifying unique, potentially important, and sensitive information within large collections of documents, as existing techniques fail to effectively characterize and protect such information.
Innovation Solution
The method involves calculating the collection frequency of character strings across multiple document collections, assigning a uniqueness rating, and performing automated actions such as anonymization or protection based on this rating to identify and safeguard unique information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional information processing techniques are used to analyze large document collections, then processing capability is limited, but identifying unique and sensitive information becomes inefficient and inaccurate
Solution Approach 1:
The patent segments the large document collection into multiple smaller collections, each representing a specific entity or time period. By calculating collection frequency at this segmented level, the system can efficiently identify unique character strings without processing the entire large collection at once, thus improving both accuracy and processing efficiency.
Solution Approach 2:
The patent replaces traditional mechanical information processing methods with an automated computational system that calculates collection frequency using string-based algorithms. This substitution enables efficient processing of large document collections by using automated frequency calculations rather than manual or traditional search methods.
2Measurement precision
If all character strings in document collections are analyzed in detail, then identification accuracy improves, but processing time and computational resources increase significantly
Solution Approach 1:
The patent extracts only the essential characteristic of character strings - their collection frequency - rather than analyzing all details of each string. By focusing on this single extracted metric, the system achieves accurate uniqueness assessment without the need for comprehensive detailed analysis of every character string, thus reducing processing time.
Solution Approach 2:
The patent changes the analysis parameter from detailed content examination to collection frequency calculation. This parameter transformation allows the system to assess uniqueness by comparing frequency values rather than performing detailed content analysis, significantly reducing computational time while maintaining assessment accuracy.
3Extent of automation
If manual review methods are used to identify sensitive information, then accuracy can be maintained, but automation level and processing speed decrease
Solution Approach 1:
The patent implements a self-service automated system that independently calculates collection frequency, assigns uniqueness ratings, and identifies sensitive information without requiring manual review. The system serves itself by using the collected document data to automatically generate protection decisions, achieving both high automation and maintained accuracy through algorithmic consistency.
Solution Approach 2:
The patent incorporates feedback mechanisms where the calculated collection frequency directly influences the uniqueness rating assignment, which in turn determines protection actions. This feedback loop enables automated decision-making that maintains accuracy by continuously refining identification based on frequency data, while simultaneously achieving high processing speed through elimination of manual intervention.
Data Source
AI summary
Techniques are provided for assessing uniqueness of information using string-based collection frequency techniques. One method comprises obtaining multiple collections of documents from at least one data source; determining a collection frequency for a given character string based on a number of the collections comprising the given character string relative to a total number of the collections; assigning a uniqueness rating to the given character string based at least in part on a comparison of the collection frequency of the given character string to a collection frequency of one or more additional character strings in one or more of the plurality of collections; and performing an automated action using the given character string based on the assigned uniqueness rating. The automated action may comprise protecting the given character string and/or identifying the given character string as important information satisfying one or more importance criteria.


