Text Data Augmentation via Multigraph Translation for Fraud Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current fraud detection technologies face challenges in effectively identifying CEO fraud and other text-based fraudulent schemes due to the scarcity of labeled data, as these scams often require specific context and cannot be automated or scaled like more widespread threats, making it difficult to collect sufficient examples for building a performant fraud detection model.
Innovation Solution
A text data augmentation function is introduced that applies successive transformations such as machine translation, synonym replacement, and misspelling replacement to original text documents, ensuring the augmented text remains semantically similar and relevant, thereby increasing the quantity of labeled data for training fraud detection models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fraud detection models are built using traditional methods with limited labeled data, then the model development process is simple, but the detection accuracy and reliability are insufficient
Solution Approach 1:
The patent creates synthetic copies of fraud detection scenarios by generating artificial text documents that simulate fraudulent communications. These synthetic data copies are created through template-based generation and transformation of existing fraud cases, allowing the model to train on expanded datasets without requiring additional real-world fraud examples.
Solution Approach 2:
The patent transforms existing fraud data by applying parameter changes such as modifying text templates, altering communication patterns, changing temporal parameters, and varying contextual elements. These parameter transformations generate diverse training examples from limited original data, improving model reliability through enhanced data variability.
2Reliability
If more labeled fraud data is collected to improve model performance, then the detection quality improves, but the data collection process becomes more complex and time-consuming
Solution Approach 1:
The patent performs preliminary data preparation by pre-defining fraud scenario templates, communication patterns, and transformation rules before actual training data generation. This preliminary structuring allows rapid generation of training examples when needed, eliminating the time-consuming process of manual data collection and annotation.
Solution Approach 2:
The system enables self-service data generation by automatically creating synthetic fraud examples through algorithmic template expansion and transformation. The model can generate its own training data without external intervention, continuously expanding its training dataset automatically without requiring manual data collection efforts.
3Quantity of substance
If synthetic data generation methods are used to augment the dataset, then the data quantity increases, but the data quality and semantic accuracy may deteriorate
Solution Approach 1:
The patent introduces template structures as intermediary elements that mediate between raw fraud examples and generated synthetic data. These templates act as quality control mechanisms, ensuring that generated data maintains semantic accuracy and structural integrity while still providing diversity for model training.
Solution Approach 2:
The system implements feedback mechanisms where generated synthetic data is evaluated against quality metrics and validation rules. Templates incorporate feedback loops that verify semantic consistency and allow refinement of generation parameters, ensuring that data quality is maintained while quantity increases.
Data Source
AI summary
A computer-implemented method of generating an augmented electronic text document comprises establishing a directed multigraph where each vertex is associated with a separate language and is connected to at least one other one of the vertices by an oriented edge indicative of a machine translation engine's ability to translate between languages associated with the vertices connected by the oriented edge with acceptable performance. The directed multigraph is then traversed starting at a predetermined origin vertex associated with an original language of the original electronic text document by randomly selecting an adjacent vertex pointed to by an oriented edge connected to the predetermined origin vertex and causing a machine translation engine to translate the original electronic text document from the original language to a language associated with the selected vertex. The directed multigraph is then further traversed as allowed by the oriented edges from the intermediate vertex to successive other next-adjacent connected vertices, each time machine translating a previously-translated electronic text document into a language associated with a randomly-selected next-adjacent vertex until the predetermined origin vertex is selected and the previously translated electronic text document is re-translated into the original language and designated as the augmented electronic text document.


