Document Smart Grouping for User-Specific Consequence Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Evaluating a large dataset, such as a corpus of documents associated with an enterprise computer user, requires significant review time and processing power, and existing automated algorithms struggle to manage dataset variability, necessitating substantial human and computing resources.
Innovation Solution
A computing platform applies unsupervised and supervised machine learning algorithms to create smart groups from user documents, enabling efficient categorization and calculation of user-specific consequence indices through iterative user interaction and labeling, leveraging overlapping group creation for enhanced efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated algorithms are used to evaluate large datasets, then processing speed is improved, but the ability to manage dataset variability deteriorates
Solution Approach 1:
The patent segments the large dataset into multiple clusters or groups using unsupervised machine learning algorithms. This segmentation allows the system to process manageable portions of data while maintaining the ability to handle variability within and across segments, resolving the contradiction between processing speed and variability management.
Solution Approach 2:
The patent employs dynamic clustering algorithms that can adapt to different data distributions and variability patterns. The system dynamically adjusts clustering parameters and approaches based on the specific characteristics of the dataset being processed, enabling both fast processing and effective variability management.
2Measurement precision
If manual review is used to evaluate documents, then accuracy is improved, but review time increases
Solution Approach 1:
The patent implements feedback loops where user interactions with clustered results are used to refine and improve future clustering. This feedback mechanism allows the system to learn from manual review patterns and progressively improve accuracy while maintaining automated processing speed, resolving the contradiction between accuracy and review time.
Solution Approach 2:
The system performs self-improvement by automatically learning from user feedback and refining its clustering algorithms without requiring continuous manual intervention. This self-service capability enables the system to maintain high accuracy while minimizing review time through progressive automation.
3Reliability
If comprehensive document analysis is performed, then threat detection accuracy is improved, but computing resources required increase
Solution Approach 1:
The patent segments the comprehensive analysis task into multiple clustering stages, where unsupervised learning algorithms first group documents by similarity. This segmentation reduces the computational burden by processing documents in organized groups rather than individually, maintaining threat detection accuracy while reducing computing resource requirements.
Solution Approach 2:
The patent applies partial analysis strategies where not all documents require the same level of comprehensive analysis. Superseded document identification and selective clustering allow the system to focus computational resources on the most relevant documents, achieving effective threat detection with reduced computing resources.
Data Source
AI summary
Aspects of the disclosure relate to using a machine learning system to process a corpus of documents associated with a user to determine a user-specific consequence index. A computing platform may load a corpus of documents associated with a user. Subsequently, the computing platform may create a first plurality of smart groups based on the corpus of documents, and then may generate a first user interface comprising a representation of the first plurality of smart groups. Next, the computing platform may receive user input applying one or more labels to a plurality of documents associated with at least one smart group. Subsequently, the computing platform may create a second plurality of smart groups based on the corpus of documents and the received user input. Then, the computing platform may generate a second user interface comprising a representation of the second plurality of smart groups.


