NLP Training Samples with Inflectional Perturbations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing NLP systems trained on standard English data often exhibit linguistic discrimination when interacting with non-native or non-standard English speakers, failing to understand or misrepresent their language due to variability in inflectional morphology.
Innovation Solution
Generating adversarial training samples with inflectional perturbations that preserve semantic meaning, which are used to fine-tune NLP systems, thereby improving their robustness against linguistic discrimination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If NLP systems are trained on standard English data samples, then the systems achieve high accuracy on standard English inputs, but the systems exhibit linguistic discrimination and fail to understand non-native or non-standard English speakers
Solution Approach 1:
The patent applies preliminary action by generating adversarial training samples with inflectional perturbations before the actual training process. These perturbed samples, which include variations in verb conjugations, noun plurals, and other morphological changes, are created in advance to expose the model to diverse linguistic patterns during training, thereby improving its adaptability to non-standard English while maintaining standard English accuracy
Solution Approach 2:
The patent implements parameter changes by systematically modifying linguistic parameters in the training data, specifically inflectional morphology parameters such as verb tenses, noun plurals, and adjective forms. By varying these parameters across training samples while preserving semantic meaning, the model learns to recognize patterns across different inflectional forms, resolving the contradiction between accuracy on standard English and adaptability to inflectional variations
2Ease of manufacture
If NLP systems use standard English training data, then the training process is simple and efficient, but the systems misrepresent minority languages and exhibit bias
Solution Approach 1:
The patent converts the harmful effect of linguistic discrimination into a benefit by using adversarial perturbation techniques. Instead of directly addressing bias through complex debiasing methods, the system introduces controlled inflectional variations that mimic the errors made by non-native speakers. This transforms the training process to actively combat linguistic discrimination by making the model robust to such variations, thereby eliminating bias while maintaining training efficiency
Solution Approach 2:
The patent applies parameter changes by modifying the linguistic parameters of training samples to include inflectional variations. By systematically varying morphological parameters such as verb conjugations and noun forms while keeping the core semantic content unchanged, the training process remains relatively simple yet effectively reduces linguistic discrimination by exposing the model to diverse language patterns
Data Source
AI summary
Embodiments described herein provide systems and methods for generating an adversarial sample with inflectional perturbations for training a natural language processing (NLP) system. A natural language sentence is received at an inflection perturbation module. Tokens are generated from the natural language sentence. For each token that has a part of speech that is a verb, adjective, or an adverb, an inflected form is determined. An adversarial sample of the natural language sentence is generated by detokenizing inflected forms of the tokens. The NLP system is trained using the adversarial sample.


