Text Data Augmentation via Multigraph Translation for Fraud Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current fraud detection technologies face challenges in effectively identifying CEO fraud and other text-based fraudulent schemes due to the scarcity of labeled data, as these scams often require specific context and cannot be automated or scaled like more widespread threats, making it difficult to collect sufficient examples for building a performant fraud detection model.

Innovation Solution

A text data augmentation function is introduced that applies successive transformations such as machine translation, synonym replacement, and misspelling replacement to original text documents, ensuring the augmented text remains semantically similar and relevant, thereby increasing the quantity of labeled data for training fraud detection models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fraud detection models are built using traditional methods with limited labeled data, then the model development process is simple, but the detection accuracy and reliability are insufficient

Engineering Contradiction:
Improvefraud detection accuracyVSAvoidlabeled data quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of fraud detection scenarios by generating artificial text documents that simulate fraudulent communications. These synthetic data copies are created through template-based generation and transformation of existing fraud cases, allowing the model to train on expanded datasets without requiring additional real-world fraud examples.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms existing fraud data by applying parameter changes such as modifying text templates, altering communication patterns, changing temporal parameters, and varying contextual elements. These parameter transformations generate diverse training examples from limited original data, improving model reliability through enhanced data variability.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If more labeled fraud data is collected to improve model performance, then the detection quality improves, but the data collection process becomes more complex and time-consuming

Engineering Contradiction:
Improvemodel performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary data preparation by pre-defining fraud scenario templates, communication patterns, and transformation rules before actual training data generation. This preliminary structuring allows rapid generation of training examples when needed, eliminating the time-consuming process of manual data collection and annotation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service data generation by automatically creating synthetic fraud examples through algorithmic template expansion and transformation. The model can generate its own training data without external intervention, continuously expanding its training dataset automatically without requiring manual data collection efforts.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If synthetic data generation methods are used to augment the dataset, then the data quantity increases, but the data quality and semantic accuracy may deteriorate

Engineering Contradiction:
Improvetraining data quantityVSAvoiddata semantic accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces template structures as intermediary elements that mediate between raw fraud examples and generated synthetic data. These templates act as quality control mechanisms, ensuring that generated data maintains semantic accuracy and structural integrity while still providing diversity for model training.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms where generated synthetic data is evaluated against quality metrics and validation rules. Templates incorporate feedback loops that verify semantic consistency and allow refinement of generation parameters, ensuring that data quality is maintained while quantity increases.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10997366B2Methods, devices and systems for data augmentation to improve fraud detection
Publication Date: 2021.05.04 VADE USA INC
  • US10997366B2 patent drawing
  • US10997366B2 patent drawing
  • US10997366B2 patent drawing

AI summary

A computer-implemented method of generating an augmented electronic text document comprises establishing a directed multigraph where each vertex is associated with a separate language and is connected to at least one other one of the vertices by an oriented edge indicative of a machine translation engine's ability to translate between languages associated with the vertices connected by the oriented edge with acceptable performance. The directed multigraph is then traversed starting at a predetermined origin vertex associated with an original language of the original electronic text document by randomly selecting an adjacent vertex pointed to by an oriented edge connected to the predetermined origin vertex and causing a machine translation engine to translate the original electronic text document from the original language to a language associated with the selected vertex. The directed multigraph is then further traversed as allowed by the oriented edges from the intermediate vertex to successive other next-adjacent connected vertices, each time machine translating a previously-translated electronic text document into a language associated with a randomly-selected next-adjacent vertex until the predetermined origin vertex is selected and the previously translated electronic text document is re-translated into the original language and designated as the augmented electronic text document.