Automated Data Classification Engine Using Compound Feature Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for automated classification of large data sets, particularly unstructured and semi-structured documents, face challenges in efficiently integrating search capabilities with text analysis and require significant human intervention for labeling and training, leading to inconsistencies and inefficiencies.

Innovation Solution

An automated or semi-automated system that converts unclassified data sets into graphic and text compound data sets, uses machine learning classifiers to improve classification performance, generates self-training data sets, and detects inconsistencies, allowing for automated labeling and classification with reduced human effort and improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated classification systems are used for large unstructured data sets, then productivity is improved, but measurement precision deteriorates due to inconsistencies in automated labeling

Engineering Contradiction:
Improveclassification speedVSAvoidlabeling accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback loops where classification results are continuously evaluated and used to refine labeling rules. The classification engine processes documents and feeds results back to update the labeling system, improving accuracy over time while maintaining automated processing speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by pre-processing documents and extracting features before classification. Training data is prepared in advance with standardized labeling rules, so that when actual classification occurs, the system can operate efficiently with consistent precision.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual labeling is used for training data, then measurement precision is improved, but loss of time increases due to significant human intervention required

Engineering Contradiction:
Improvelabeling accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically generating training data through its own classification operations. The labeling system uses pre-defined rules and extracted features to create labeled training examples without requiring manual human intervention, thus maintaining precision while eliminating time loss.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system merges the classification and labeling functions into a unified process. The same feature extraction and analysis mechanisms used for classification are leveraged to generate training labels automatically, combining multiple functions into one efficient system that eliminates the need for separate manual labeling steps.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If existing classification systems are used, then ease of operation is maintained, but adaptability deteriorates due to inability to handle inconsistent data sets from multiple sources

Engineering Contradiction:
Improvesystem usabilityVSAvoiddata source compatibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The labeling system is designed with universal functionality to handle multiple data types and formats from diverse sources. It can process structured data, semi-structured documents, and unstructured text using the same core mechanisms, making the system adaptable to various data sources while maintaining ease of operation through a unified interface.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements dynamic adaptability by allowing labeling rules and parameters to be adjusted based on the characteristics of incoming data. The classification engine can dynamically adapt to different data sources and formats, and these adaptations are automatically incorporated into the labeling process without requiring system reconfiguration or sacrificing usability.

Inventive Principle:
Principle #15Dynamics

4Loss of time

If automated labeling systems are implemented, then loss of time is reduced, but manufacturing precision deteriorates due to inconsistencies in automated data generation

Engineering Contradiction:
Improvelabeling timeVSAvoiddata consistency
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The system maintains precision by carefully controlling and optimizing parameters in the automated labeling process. Feature extraction parameters, similarity thresholds, and classification criteria are tuned to ensure consistent results. The system monitors and adjusts these parameters to maintain manufacturing precision while benefiting from automated processing speed.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11836584B2Data extraction engine for structured, semi-structured and unstructured data with automated labeling and classification of data patterns or data elements therein, and corresponding method thereof
Publication Date: 2023.12.05 SWISS REINSURANCE CO LTD
  • US11836584B2 patent drawing
  • US11836584B2 patent drawing
  • US11836584B2 patent drawing

AI summary

A fully or semi-automated, integrated learning, labeling and classification system and method have closed, self-sustaining pattern recognition, labeling and classification operation, wherein unclassified data sets are selected and converted to an assembly of graphic and text data forming compound data sets that are to be classified. By means of feature vectors, which can be automatically generated, a machine learning classifier is trained for improving the classification operation of the automated system during training as a measure of the classification performance if the automated labeling and classification system is applied to unlabeled and unclassified data sets, and wherein unclassified data sets are classified automatically by applying the machine learning classifier of the system to the compound data set of the unclassified data sets.