Machine Learning Term Prediction for Outlier Data Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing large volumes of electronic data are time-consuming and require manual effort to identify outlier or unexpected data, making it difficult to obtain timely and meaningful results.

Innovation Solution

A system and method using machine learning techniques, specifically building and training a term embedding model and a term prediction model, to determine outlier data by analyzing term patterns between different term types, such as incident, vehicle damage, and injury terms in the context of vehicle accidents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review methods are used to analyze large volumes of electronic data, then accuracy in identifying outlier data can be maintained, but time consumption increases significantly

Engineering Contradiction:
Improveaccuracy in identifying outlier dataVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-analysis where the machine learning model independently processes electronic data, identifies term patterns, predicts expected terms, and detects outliers without requiring manual human review. The model serves itself by automatically learning from historical data and applying patterns to new data points.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical review processes with automated machine learning systems. The term prediction model uses computational algorithms to substitute human analytical judgment, automatically comparing predicted terms against actual terms to identify outliers in electronic data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated machine learning techniques are applied to process electronic data, then processing speed and timeliness improve, but system complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex data analysis task into distinct functional components: term extraction, pattern learning, term prediction, cohesiveness determination, and outlier identification. Each component handles a specific aspect of the analysis, making the overall complex system more manageable and maintainable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary term prediction model that bridges the gap between raw electronic data and final outlier identification. This intermediate layer processes and predicts expected terms before final comparison, simplifying the overall system architecture by breaking down the analysis into manageable stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual analysis of large data volumes is performed, then detailed examination of each data point is possible, but the ability to obtain timely results deteriorates

Engineering Contradiction:
Improvedetailed examination capabilityVSAvoidtimeliness of results
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The machine learning model continuously processes electronic data in real-time or near-real-time, maintaining constant analysis without interruption. The system continuously learns from new data and updates its patterns, enabling timely detection of outliers as they appear without requiring batch processing or manual review cycles.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary term extraction and pattern learning on historical data before analyzing new incoming data. By pre-processing and pre-learning from historical corpora, the model is ready to quickly and accurately identify outliers in new data without requiring time-consuming manual examination of each data point.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12299013B2Predicting outlier data from network of electronic data
Publication Date: 2025.05.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12299013B2 patent drawing
  • US12299013B2 patent drawing
  • US12299013B2 patent drawing

AI summary

A computer-implemented method, computer program product, and/or computing system for determining unexpected data includes: running for a given claim a machine learning term prediction model with first and second term types to predict third type terms in a specific domain that uses a machine learning term embedding model that has learned the correlation between term types in the specific domain; determining cohesiveness between the predicted third type terms and the actual third type terms from the given claim; determining a value score between the predicted third type terms and the actual third type; and running a machine learning propensity model using the value score determined for the given claim and context data for the given claim to determine if the given claim is unexpected.