NLP Training Samples with Inflectional Perturbations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing NLP systems trained on standard English data often exhibit linguistic discrimination when interacting with non-native or non-standard English speakers, failing to understand or misrepresent their language due to variability in inflectional morphology.

Innovation Solution

Generating adversarial training samples with inflectional perturbations that preserve semantic meaning, which are used to fine-tune NLP systems, thereby improving their robustness against linguistic discrimination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If NLP systems are trained on standard English data samples, then the systems achieve high accuracy on standard English inputs, but the systems exhibit linguistic discrimination and fail to understand non-native or non-standard English speakers

Engineering Contradiction:
Improveaccuracy on standard EnglishVSAvoidhandling of inflectional variations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by generating adversarial training samples with inflectional perturbations before the actual training process. These perturbed samples, which include variations in verb conjugations, noun plurals, and other morphological changes, are created in advance to expose the model to diverse linguistic patterns during training, thereby improving its adaptability to non-standard English while maintaining standard English accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements parameter changes by systematically modifying linguistic parameters in the training data, specifically inflectional morphology parameters such as verb tenses, noun plurals, and adjective forms. By varying these parameters across training samples while preserving semantic meaning, the model learns to recognize patterns across different inflectional forms, resolving the contradiction between accuracy on standard English and adaptability to inflectional variations

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If NLP systems use standard English training data, then the training process is simple and efficient, but the systems misrepresent minority languages and exhibit bias

Engineering Contradiction:
Improvetraining process simplicityVSAvoidlinguistic discrimination
Core Design Contradiction:
Ease of manufactureVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful effect of linguistic discrimination into a benefit by using adversarial perturbation techniques. Instead of directly addressing bias through complex debiasing methods, the system introduces controlled inflectional variations that mimic the errors made by non-native speakers. This transforms the training process to actively combat linguistic discrimination by making the model robust to such variations, thereby eliminating bias while maintaining training efficiency

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent applies parameter changes by modifying the linguistic parameters of training samples to include inflectional variations. By systematically varying morphological parameters such as verb conjugations and noun forms while keeping the core semantic content unchanged, the training process remains relatively simple yet effectively reduces linguistic discrimination by exposing the model to diverse language patterns

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11256754B2Systems and methods for generating natural language processing training samples with inflectional perturbations
Publication Date: 2022.02.22 SALESFORCE INC
  • US11256754B2 patent drawing
  • US11256754B2 patent drawing
  • US11256754B2 patent drawing

AI summary

Embodiments described herein provide systems and methods for generating an adversarial sample with inflectional perturbations for training a natural language processing (NLP) system. A natural language sentence is received at an inflection perturbation module. Tokens are generated from the natural language sentence. For each token that has a part of speech that is a verb, adjective, or an adverb, an inflected form is determined. An adversarial sample of the natural language sentence is generated by detokenizing inflected forms of the tokens. The NLP system is trained using the adversarial sample.