Iterative Learning for Dynamic Data Compliance Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing rules-based classification systems are inadequate for handling complex, unexpected, or large data sets, are static, and ineffective for continuous data streams, and fail to utilize relationships between data sets, leading to inaccurate and costly compliance verification.
Innovation Solution
A system and method that iteratively learns data compliance by storing compliant and non-compliant datasets, extracting meta-data, and generating estimated rules using machine learning algorithms to dynamically verify data compliance against time-variable rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If rules-based classification systems are used to classify data, then classification can be performed with programmed rules, but the system becomes static and cannot continuously verify data compliance or handle unexpected data formats
Solution Approach 1:
The patent transforms the static rules-based system into a dynamic machine learning-based system that continuously learns from compliant and non-compliant data examples. The classification model is iteratively trained and updated to adapt to changing data formats and compliance requirements, enabling the system to handle unexpected data while maintaining verification accuracy.
Solution Approach 2:
The system employs self-supervised learning where the classification model automatically learns compliance rules from labeled examples of compliant and non-compliant data without requiring explicit programming of each rule. The model serves itself by identifying patterns and relationships in the data, reducing the need for manual rule configuration and improving adaptability.
2Productivity
If programmed rules are used for data classification, then the system can operate with defined criteria, but it becomes ineffective at classifying live or continuous streams of data
Solution Approach 1:
The patent performs preliminary training of the machine learning model using historical compliant and non-compliant data before deployment. This pre-training phase allows the model to learn compliance patterns in advance, enabling it to quickly and accurately classify live data streams without requiring complex rule evaluation during real-time operation.
Solution Approach 2:
The system implements continuous feedback loops where classification results are monitored and used to iteratively retrain and improve the model. Compliant and non-compliant examples are fed back into the training process, allowing the model to adapt to new compliance requirements and maintain high accuracy over time while processing continuous data streams.
3Adaptability or versatility
If rules are programmed to define compliance categories, then classification can be performed, but the system cannot easily modify rules when classification criteria change
Solution Approach 1:
The patent changes the fundamental parameter of rule representation from explicit programmed code to learned model parameters. Instead of modifying programming logic when rules change, the system updates the machine learning model by retraining it with new compliant and non-compliant examples, allowing flexible adaptation to changing classification criteria through data-driven parameter adjustment.
4Measurement precision
If rules-based systems are used to classify unknown data, then classification can be attempted, but the system becomes expensive and inaccurate for complex or large data sets
Solution Approach 1:
The patent uses meta-data extracted from compliant and non-compliant data examples as simplified representations (copies) of the actual compliance rules. Instead of directly analyzing complex raw data with expensive rule-based systems, the model learns from compressed meta-data features, reducing computational complexity while maintaining classification accuracy for large and complex data sets.
Data Source
AI summary
An exemplary system, method, and computer-accessible medium can include, for example, establishing a unique rule-identifier in one-to-one correspondence with at least one set of unknown time-variable rules against which data is to be made compliant, obtaining at least one set of data marked compliant against the one or more set of rules, obtaining meta-data from the compliant data, obtaining at least one set of data marked non-compliant against the set of unknown time-variable rules, extracting meta-data from the non-compliant data, joining the set of compliant and non-compliant metadata to generate a set of estimated rules corresponding to the rule-identifier based at least one of (i) the meta-data of the joined set and (ii) machine learning algorithms.


