Anomaly Detection via Synthetic Tuplet Permutation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning methods face challenges in identifying anomalous data entities within datasets, especially when the proportion of anomalous data is low, as they often require extensive labeled data and computational resources, and may not perform well with complex or arbitrary data types.

Innovation Solution

The method involves creating dummy data entities through permutation of real data entities, allowing for the identification of characteristic relations that differentiate between normal and anomalous data, and using these relations to train a classifier for improved anomaly detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current machine learning methods are used to identify anomalous data entities, then detection accuracy may be maintained, but extensive labeled data and computational resources are required

Engineering Contradiction:
Improveanomaly detection accuracyVSAvoidlabeled data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.

Inventive Principle:
Principle #26Copying

2Measurement precision

If current machine learning methods are used to identify anomalous data entities, then detection capability may be maintained, but computational resources are extensively consumed

Engineering Contradiction:
Improveanomaly detection capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention changes the parameter composition of data tuplets through permutation operations. By systematically varying the arrangement and combination of data entities within tuplets, the system generates diverse synthetic anomaly examples that improve detection capability while maintaining computational efficiency.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If dummy tuplets are created through permutation of real tuplets, then the proportion of anomalous data is artificially increased, but additional data processing steps are required

Engineering Contradiction:
Improveproportion of anomalous dataVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by generating synthetic anomalous data samples before the actual detection process. Dummy tuplets are created by permuting real data entities, establishing a foundation of artificial anomaly examples that can be used for training without requiring extensive manually labeled anomalous data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.

Inventive Principle:
Principle #26Copying

4Measurement precision

If extensive manual labeling is performed to identify anomalous entities, then training data quality may be improved, but time and computational resources are significantly consumed

Engineering Contradiction:
Improvetraining data qualityVSAvoidmanual labeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The invention creates copies of real data entities through permutation to generate dummy tuplets that represent anomalous data. These synthetic copies serve as proxies for actual anomalous examples, allowing the system to train anomaly detection models without needing large volumes of真实 labeled anomaly data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own training data through permutation of existing real data. The dummy tupplet generation process is autonomous, creating synthetic anomaly examples without requiring external manual intervention or labeling efforts.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20170255669A1Systems and methods for detection of anomalous entities
Publication Date: 2017.09.07 SPARKBEYOND
  • US20170255669A1 patent drawing
  • US20170255669A1 patent drawing
  • US20170255669A1 patent drawing

AI summary

There is provided a computer-implemented method of identifying anomalous entities in a dataset, comprising: selecting a subset of training entities from entities of at least one dataset; determining dummy tuplets of entities in the subset by applying a permutation function on real tuplets, wherein the real tuplets represent original and normal data of the at least one dataset, wherein the dummy tuplets represent anomalous data based on artificially created data not found in the original and normal at least one dataset, each one of the real tuplets and dummy tuplets comprises at least two of the training entities; analyzing the dummy tuplets and the real tuplets to identify at least one predefined characteristic relation that statistically differentiates between the real tuplets and the dummy tuplets according to a distinguishing requirement; and outputting the identified at least one predefined characteristic relation to identify a normal entity and/or an anomalous entity.