Contextual Data Classification via Hybrid ML and Rules

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data classification systems, particularly those based on machine learning and artificial intelligence, are inadequate for contextual classification of data, as they struggle to identify business-specific sensitivity classifications and rely heavily on metadata, leading to inefficiencies and lack of transparency in data management and compliance.

Innovation Solution

A hybrid rules-based and machine learning-based system that utilizes user-generated contextual sensitivity hierarchies to classify data, leveraging metadata attributes and enabling iterative classification processes to ensure accurate and context-specific data categorization, thereby overcoming the limitations of existing systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning-based systems are used for data classification, then processing speed and automation are improved, but contextual accuracy and business-specific sensitivity classification capability deteriorate

Engineering Contradiction:
Improvedata processing speedVSAvoidcontextual classification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent combines machine learning-based automated classification with rules-based metadata analysis and user feedback mechanisms into a hybrid system. The ML component handles initial rapid classification while the rules-based component refines accuracy using business-specific metadata, achieving both speed and contextual precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system incorporates user feedback loops where users can correct or verify classifications. This feedback is used to iteratively improve the ML models and refine classification rules, enabling the system to learn from mistakes and continuously enhance contextual accuracy while maintaining automated processing speed.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If rules-based classification systems are used, then contextual accuracy and transparency are improved, but processing speed and automation capability deteriorate

Engineering Contradiction:
Improvecontextual classification accuracyVSAvoiddata processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The classification process is segmented into multiple stages: initial rapid ML-based filtering, followed by targeted rules-based validation on borderline cases, and finally user review only for uncertain classifications. This segmentation allows the system to achieve high accuracy without processing all data through the slowest methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies comprehensive rules-based classification only when necessary (e.g., for high-stakes data or ambiguous cases) rather than to all data uniformly. This partial application of the more accurate but slower method maintains overall processing speed while ensuring accuracy where it matters most.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If comprehensive metadata analysis is performed, then classification accuracy is improved, but system complexity and computational requirements deteriorate

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies different levels of metadata analysis depth to different data types and classification contexts. Not all data requires the same comprehensive analysis - the system tailors the metadata examination scope to the specific classification needs, reducing unnecessary computational overhead while maintaining accuracy where required.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11947574B2System and method for user interactive contextual model classification based on metadata
Publication Date: 2024.04.02 NVISNX INC
  • US11947574B2 patent drawing
  • US11947574B2 patent drawing
  • US11947574B2 patent drawing

AI summary

A system and a method for contextual categorization of data comprises a server having a processor and a non-transitory computer-readable storage medium in electronic communication with the processor and comprising program instructions executable by the processor to access an initial inventory of data set and metadata associated with the initial inventory of data set. The system is then configured to classify the initial inventory of data set by using the metadata into (a) reduced set of data comprising high level sensitivity classification and (b) a remainder data set. The system and method can be further configured for contextual categorization of data that involves receiving an initial data set to be categorized; establishing a library of contextual classifiers, the library comprising (1) a set of predetermined high level sensitivity classifications and (2) a set of user-generated business-specific sensitivity classifications subordinated below the high level sensitivity classifications; identifying and removing redundant, outdated, trivial or abandoned (ROTA) data from the initial data set to create a reduced data set and a remainder data set of ROTA data; applying the user-generated business-specific sensitivity classifications to the reduced data set to create a first set of classified data and a second set of unclassified data; and iteratively applying additional user-generated business-specific sensitivity classifications to the both the first set of classified data and the second set of unclassified data until all data in the reduced data set has been classified in exactly one use-generated business-specific sensitivity classification.