Intelligent Term Grouping for Scalable Sensitive Data Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current security systems face challenges in scalable document registration and manual processes for synthesizing sensitive information, making it time-intensive and complex to identify and protect valuable data, especially in non-fixed formats.
Innovation Solution
A system and method for intelligent term grouping using concept builders that select key terms and regular expressions from text mining, automating the process of identifying related terms and forming concepts for data classification, enabling efficient protection of sensitive information across networks without prior knowledge of what needs to be protected.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual processes are used for synthesizing sensitive information, then security professionals can identify and protect valuable data, but the process becomes time-intensive and complex
Solution Approach 1:
The system enables self-service by automatically synthesizing sensitive information through text mining and concept building. The network appliance autonomously identifies sensitive data, extracts key terms, and creates protection rules without requiring continuous manual intervention from security professionals, thus resolving the time-intensive nature of manual processes while maintaining security reliability
Solution Approach 2:
The patent replaces manual mechanical processes with automated computational systems. Text mining algorithms, concept builders, and pattern recognition systems substitute for manual security analysis, automatically processing large volumes of data to identify sensitive information and generate protection rules, thereby eliminating the time-consuming aspects of manual synthesis
2Reliability
If manual registration of documents is used, then sensitive information can be identified, but scalability is limited
Solution Approach 1:
The network appliance performs multiple functions including text mining, concept building, rule generation, and data protection within a single system. This universal approach allows the same automated system to handle diverse document types and sensitive information formats across the enterprise network, enabling scalable deployment without requiring separate manual registration processes for different data types
Solution Approach 2:
Manual document registration is replaced with automated text mining and pattern recognition systems that can process unlimited volumes of documents across the network. The system automatically synthesizes sensitive information from various sources and formats, making the process scalable to enterprise-wide deployments without proportionally increasing manual effort
3Reliability
If comprehensive data protection is implemented across all information assets, then security coverage is improved, but system complexity increases
Solution Approach 1:
The system extracts only the essential elements needed for protection by using text mining to identify key terms and concepts from sensitive documents. The concept builder extracts meaningful patterns and synthesizes them into focused protection rules, rather than attempting to manually configure comprehensive security policies for every possible data element, thus reducing system complexity while maintaining broad security coverage
Solution Approach 2:
The automated system self-generates protection rules by analyzing document content and identifying sensitive information patterns. This self-service capability eliminates the need for complex manual configuration and ongoing management of security policies, allowing comprehensive coverage across all information assets while keeping the system manageable through automation rather than increasing complexity through manual processes
Data Source
AI summary
A method is provided in one example embodiment and it includes identifying a root word for a tree to be used in managing data and creating a word stem to be included in the tree. A query is initiated to determine whether a stem node exists at one or more branch points of the word, and if the stem node does not exist, then the stem node is added to a branch point of the tree. In more specific embodiments, if the stem node does exist, then node statistics are updated. In other embodiments, the method includes updating a branch point list after creating the word stem. In yet other embodiments, the branch point is a word or a combination of words. The tree can be used to identify locations and frequencies within a document set where one or more words are present.


