Data Boundary Deriving System for AI Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI-based process monitoring systems face challenges in generating learning data for abnormal states, especially in the initial stages of equipment operation or when data for abnormal states is scarce, making it difficult to establish accurate models for detecting anomalies.
Innovation Solution
A data boundary deriving system that receives sample data, identifies outliers, generates clusters, derives a probability density function, and labels data based on calculated values to create learning data without separate labeling operations, enabling the detection of abnormal states even without explicit abnormal state data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional AI-based process monitoring systems are used to detect abnormal states, then anomaly detection capability is improved, but the requirement for labeled learning data for abnormal states makes the system inapplicable when abnormal state data is scarce or unavailable
Solution Approach 1:
The patent creates synthetic abnormal state data by copying and transforming normal state data through the probability density function. Specifically, it generates artificial abnormal state samples by applying transformations to normal state data, thereby creating counterfeit abnormal state learning data that mimics real abnormal states without requiring actual abnormal state observations
Solution Approach 2:
The patent performs preliminary actions by pre-calculating the probability density function from normal state data before actual anomaly detection is needed. This allows the system to prepare transformation parameters and boundary definitions in advance, enabling synthetic abnormal state data generation without waiting for actual abnormal events to occur
2Quantity of substance
If manual labeling operations are performed to create learning data, then labeled learning data is generated, but the process is time-consuming and labor-intensive
Solution Approach 1:
The patent implements self-service by enabling the system to automatically label data without human intervention. The probability density function automatically computes boundary values and assigns labels to data points based on calculated probability densities, transforming the labeling process from manual to automated and eliminating human labor requirements
Solution Approach 2:
The patent changes parameters by transforming normal state data into abnormal state data through mathematical transformations of the probability density function. It modifies data parameters (mean, variance, covariance) to create synthetic abnormal state samples, thereby automatically generating labeled data through parameter transformation rather than manual annotation
3Reliability
If equipment operates in the initial stage without frequent errors, then normal operation is maintained, but abnormal state data for learning cannot be collected
Solution Approach 1:
The patent copies normal state data and transforms it into synthetic abnormal state data using the probability density function. By applying transformations to normal state samples, it creates artificial abnormal state data that replicates the characteristics of real abnormal states, thereby obtaining learning data without requiring actual abnormal events during normal operation
Solution Approach 2:
The patent introduces the probability density function as an intermediary that bridges normal state data and abnormal state learning requirements. This intermediary transforms normal state observations into synthetic abnormal state data, enabling the system to infer abnormal state characteristics from normal state data through the mediating probability distribution model
Data Source
AI summary
Data boundary deriving system and method include: a sample data reception unit configured to receive a plurality of pieces of sample data having a plurality of characteristic values; a cluster generation unit configured to generate a plurality of clusters by classifying the plurality of pieces of sample data; a probability density function derivation unit configured to derive a probability density function based on the characteristic values of data included in each of the plurality of generated clusters; and a learning data generation unit configured to generate learning data by calculating the values of the probability density function of a cluster including each piece of sample data for each of the plurality of sample data and labeling second sample data based on the calculated values, and an operating method thereof.


