Anomaly Detection via Synthetic Tuplet Permutation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning methods face challenges in identifying anomalous data entities within datasets, especially when the proportion of anomalous data is low, as they often require extensive labeled data and computational resources, and may not perform well with complex or arbitrary data types.
Innovation Solution
The method involves creating dummy data entities through permutation of real data entities, allowing for the identification of characteristic relations that differentiate between normal and anomalous data, and using these relations to train a classifier for improved anomaly detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current machine learning methods are used to identify anomalous data entities, then detection accuracy may be maintained, but extensive labeled data and computational resources are required
Solution Approach 1:
The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.
Solution Approach 2:
The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.
2Measurement precision
If current machine learning methods are used to identify anomalous data entities, then detection capability may be maintained, but computational resources are extensively consumed
Solution Approach 1:
The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.
Solution Approach 2:
The invention changes the parameter composition of data tuplets through permutation operations. By systematically varying the arrangement and combination of data entities within tuplets, the system generates diverse synthetic anomaly examples that improve detection capability while maintaining computational efficiency.
3Quantity of substance
If dummy tuplets are created through permutation of real tuplets, then the proportion of anomalous data is artificially increased, but additional data processing steps are required
Solution Approach 1:
The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.
Solution Approach 2:
The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.
4Measurement precision
If extensive manual labeling is performed to identify anomalous entities, then training data quality may be improved, but time and computational resources are significantly consumed
Solution Approach 1:
The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.
Solution Approach 2:
The system performs self-service by automatically generating its own training data through permutation of existing real data. The dummy tupplet generation process is autonomous, creating synthetic anomaly examples without requiring external manual intervention or labeling efforts.
Data Source
AI summary
There is provided a computer-implemented method of identifying anomalous entities in a dataset, comprising: selecting a subset of training entities from entities of at least one dataset; determining dummy tuplets of entities in the subset by applying a permutation function on real tuplets, wherein the real tuplets represent original and normal data of the at least one dataset, wherein the dummy tuplets represent anomalous data based on artificially created data not found in the original and normal at least one dataset, each one of the real tuplets and dummy tuplets comprises at least two of the training entities; analyzing the dummy tuplets and the real tuplets to identify at least one predefined characteristic relation that statistically differentiates between the real tuplets and the dummy tuplets according to a distinguishing requirement; and outputting the identified at least one predefined characteristic relation to identify a normal entity and/or an anomalous entity.


