LLM Threshold Configuration for Scalable Data Anomaly Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data anomaly detection methods rely heavily on manual processes, which are time-consuming and difficult to scale, especially when dealing with vast amounts of electronic data containing duplicates, malicious activity, or unexpected events.

Innovation Solution

Utilizing Large Language Models (LLMs) to automate the detection of data anomalies by generating reference datasets, determining thresholds, and identifying deviations through algorithms, thereby reducing manual intervention and improving scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processes are used to detect data anomalies, then detection accuracy can be maintained through human analysis, but the process becomes time-consuming and difficult to scale

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection speed and scalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary system comprising automated anomaly detection algorithms, machine learning models, and processing systems that act as a mediator between raw data and human analysts. This intermediary automatically processes vast datasets, identifies anomalies using multiple detection techniques, and presents refined results to human reviewers, thereby maintaining high detection accuracy while dramatically improving processing speed and scalability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The detection process is segmented into multiple independent stages: data collection from multiple sources, preliminary automated filtering, anomaly detection using various algorithms, validation processes, and final review. This segmentation allows each component to specialize in specific tasks, improving overall efficiency while maintaining accuracy through layered verification.

Inventive Principle:
Principle #1Segmentation

2Productivity

If automated algorithms are used to detect data anomalies, then processing speed and scalability improve, but false positives increase

Engineering Contradiction:
Improvedetection speed and scalabilityVSAvoidfalse positive rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements multi-layered feedback mechanisms where detection results from one algorithm inform and adjust subsequent detection processes. Validation algorithms provide feedback on initial detections, human reviewers provide feedback on false positives and negatives, and this feedback continuously refines detection parameters and reduces false positive rates while maintaining high productivity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Multiple anomaly detection algorithms and validation methods are merged into a unified detection system. By combining the strengths of different detection techniques and cross-validating results through multiple independent algorithms, the system achieves high scalability while maintaining reliability and minimizing false positives through consensus-based detection.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If extensive manual analysis is performed on datasets, then comprehensive understanding of anomalies is achieved, but resource consumption and time requirements increase significantly

Engineering Contradiction:
Improvecomprehensive data understandingVSAvoidanalysis time and computational resources
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary automated actions including data collection, initial filtering, preprocessing, and preliminary anomaly detection before human analysis is required. This preliminary action prepares and refines datasets in advance, ensuring comprehensive data understanding is achieved through automated preprocessing while minimizing the time and resources required for subsequent manual analysis.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250371316A1Data anomaly detection using a large language model
Publication Date: 2025.12.04 PINTEREST INC
  • US20250371316A1 patent drawing
  • US20250371316A1 patent drawing
  • US20250371316A1 patent drawing

AI summary

Disclosed are systems and methods that process a dataset to determine data anomalies in the dataset. The process may receive a query to create the dataset. At least one known data anomaly may be identified in the dataset. An algorithm that models a pattern of the dataset may be selected. The dataset, known data anomaly, and/or algorithm may be sent to a Large Language Model (LLM) with instructions to determine configuration information for data anomaly detection including at least one threshold that indicates additional anomalies in the dataset. The algorithm may create a reference dataset that is compared to the dataset to determine deviations. The threshold may determine which deviations indicate additional anomalies. The LLM may send configuration data, including at least the threshold, to an anomaly detection application, which may be configured with the configuration data and used to determine data anomalies in other, similar, datasets generated with the query or a similar query.