Synthetic Code-Switched Data Generation for Language Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in effectively training language models to detect code-switched offensive content due to the scarcity of labeled code-switching data and the nuanced, context-dependent nature of code switching.

Innovation Solution

A method is developed to generate synthetic code-switched data by identifying salient portions in textual content, translating them into another language, and reintegrating them into the original text, thereby creating a dataset for training multilingual classification models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic code-switched data is generated using automated translation methods, then the quantity of training data is improved, but the quality and contextual accuracy of the data deteriorates

Engineering Contradiction:
Improvequantity of training dataVSAvoiddata quality and contextual accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary human-in-the-loop verification process where native speakers review and validate the synthetic code-switched data generated by automated translation. This intermediary step ensures contextual accuracy and linguistic appropriateness while maintaining the efficiency of automated generation, thus resolving the contradiction between data quantity and data quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-training translation models on parallel corpora specific to the target language pair before generating synthetic code-switched data. This preliminary training ensures that the automated translation process produces higher quality output with better contextual accuracy, reducing the need for extensive manual validation while maintaining data quality.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If human-generated code-switched data is used for training, then the quality and contextual accuracy is improved, but the productivity and scalability deteriorates

Engineering Contradiction:
Improvedata quality and contextual accuracyVSAvoidproductivity and scalability
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent merges automated translation methods with human verification processes into a hybrid data generation pipeline. This combination leverages the scalability of automated systems while incorporating human expertise for quality assurance, thus achieving both high productivity and maintained data quality without the limitations of purely manual approaches.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses copying by generating synthetic code-switched data through translation of existing monolingual training data. This approach creates multiple copies and variations of training examples in the target language, dramatically increasing data quantity and scalability while maintaining contextual accuracy through controlled translation processes and optional human validation.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If more diverse language pairs are supported, then the adaptability of the model is improved, but the complexity of data collection and processing increases

Engineering Contradiction:
Improvemodel adaptabilityVSAvoiddata collection and processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements universality by developing a language-agnostic synthetic data generation framework that can handle multiple language pairs through a single automated translation pipeline. The system uses parallel corpora and translation models that are configured for different language pairs, enabling the same infrastructure to support diverse languages without requiring separate manual data collection processes for each language pair, thus maintaining simplicity while improving adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12242820B2Generating synthetic code-switched data for training language models
Publication Date: 2025.03.04 ADOBE INC
  • US12242820B2 patent drawing
  • US12242820B2 patent drawing
  • US12242820B2 patent drawing

AI summary

Techniques for training a language model for code switching content are disclosed. Such techniques include, in some embodiments, generating a dataset, which includes identifying one or more portions within textual content in a first language, the identified one or more portions each including one or more of offensive content or non-offensive content; translating the identified one or more salient portions to a second language; and reintegrating the translated one or more portions into the textual content to generate code-switched textual content. In some cases, the textual content in the first language includes offensive content and non-offensive content, the identified one or more portions include the offensive content, and the translated one or more portions include a translated version of the offensive content. In some embodiments, the code-switched textual content is at least part of a synthetic dataset usable to train a language model, such as a multilingual classification model.