Machine Learning Term Prediction for Outlier Data Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing large volumes of electronic data are time-consuming and require manual effort to identify outlier or unexpected data, making it difficult to obtain timely and meaningful results.
Innovation Solution
A system and method using machine learning techniques, specifically building and training a term embedding model and a term prediction model, to determine outlier data by analyzing term patterns between different term types, such as incident, vehicle damage, and injury terms in the context of vehicle accidents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review methods are used to analyze large volumes of electronic data, then accuracy in identifying outlier data can be maintained, but time consumption increases significantly
Solution Approach 1:
The system enables automated self-analysis where the machine learning model independently processes electronic data, identifies term patterns, predicts expected terms, and detects outliers without requiring manual human review. The model serves itself by automatically learning from historical data and applying patterns to new data points.
Solution Approach 2:
The patent replaces manual mechanical review processes with automated machine learning systems. The term prediction model uses computational algorithms to substitute human analytical judgment, automatically comparing predicted terms against actual terms to identify outliers in electronic data.
2Productivity
If automated machine learning techniques are applied to process electronic data, then processing speed and timeliness improve, but system complexity increases
Solution Approach 1:
The system segments the complex data analysis task into distinct functional components: term extraction, pattern learning, term prediction, cohesiveness determination, and outlier identification. Each component handles a specific aspect of the analysis, making the overall complex system more manageable and maintainable.
Solution Approach 2:
The patent introduces an intermediary term prediction model that bridges the gap between raw electronic data and final outlier identification. This intermediate layer processes and predicts expected terms before final comparison, simplifying the overall system architecture by breaking down the analysis into manageable stages.
3Measurement precision
If manual analysis of large data volumes is performed, then detailed examination of each data point is possible, but the ability to obtain timely results deteriorates
Solution Approach 1:
The machine learning model continuously processes electronic data in real-time or near-real-time, maintaining constant analysis without interruption. The system continuously learns from new data and updates its patterns, enabling timely detection of outliers as they appear without requiring batch processing or manual review cycles.
Solution Approach 2:
The system performs preliminary term extraction and pattern learning on historical data before analyzing new incoming data. By pre-processing and pre-learning from historical corpora, the model is ready to quickly and accurately identify outliers in new data without requiring time-consuming manual examination of each data point.
Data Source
AI summary
A computer-implemented method, computer program product, and/or computing system for determining unexpected data includes: running for a given claim a machine learning term prediction model with first and second term types to predict third type terms in a specific domain that uses a machine learning term embedding model that has learned the correlation between term types in the specific domain; determining cohesiveness between the predicted third type terms and the actual third type terms from the given claim; determining a value score between the predicted third type terms and the actual third type; and running a machine learning propensity model using the value score determined for the given claim and context data for the given claim to determine if the given claim is unexpected.


