Machine Learning Regulatory Text Classification System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack a consistent solution for managing document changes and assessing their impact across an enterprise, particularly in analyzing regulations embedded in dense documents, which requires deconstructing regulations into line-level requirements, leading to time-consuming and complex compliance assessments.
Innovation Solution
A machine learning-based system that analyzes, classifies, and maps text within documents by associating sentences with location identifiers, classifying sentences based on context, identifying relationships between texts, and generating custom machine learning models to improve speed and accuracy in regulatory compliance assessments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing manual techniques are used to deconstruct regulations into line-level requirements, then measurement precision can be achieved, but loss of time is excessive (taking over a year for 39 regulations with 9,000 requirements)
Solution Approach 1:
The patent replaces manual mechanical analysis with machine learning-based automated text processing. The system uses trained models to automatically deconstruct regulations into line-level requirements, achieving both high precision in identification and dramatically reduced processing time from over a year to minutes or seconds.
Solution Approach 2:
The patent transforms the analysis approach by changing parameters such as using contextual window sizes, similarity thresholds, and model confidence levels to optimize both precision and speed. The system adjusts these parameters dynamically to balance accurate requirement identification with efficient processing.
2Reliability
If dense documents containing thousands of regulations are analyzed manually, then comprehensive coverage can be achieved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent segments dense regulatory documents into manageable units such as sentences, clauses, and individual requirements. The system processes these segmented units independently through machine learning models, ensuring comprehensive coverage while reducing system complexity through modular architecture.
Solution Approach 2:
The patent introduces machine learning models as intermediaries between the raw dense documents and the analysis system. These models pre-process and structure the unstructured text, transforming it into organized data that is easier to analyze and reducing the complexity of subsequent processing steps.
3Manufacturing precision
If regulations with different wording but duplicative content are identified, then manufacturing precision improves, but loss of information increases due to potential misclassification
Solution Approach 1:
The patent performs preliminary actions by training machine learning models on diverse regulatory text before actual classification. The models learn to recognize semantic equivalence across different wordings and contexts, enabling accurate classification while preserving contextual information through features like contextual embeddings and attention mechanisms.
Solution Approach 2:
The patent implements feedback mechanisms where classification results are continuously refined. The system provides feedback loops that allow manual correction and retraining, improving precision in distinguishing duplicative regulations while preserving important contextual information through iterative optimization.
Data Source
AI summary
A device that includes an enterprise data indexing engine (EDIE) configured to receive a set of sentences and to compare the words in the sentences to a set of predefined keywords. The EDIE is further configured to identify one or more sentences that do not contain any of the keywords and to associate the identified sentences with a first classification type. The EDIE is further configured to identify a sentence that contains one or more keywords and to associate the sentence with a second classification type. The EDIE is further configured to link together the sentence that is associated with the second classification type and the sentences that are associated with the first classification type.


