Masking Identity Proxies in Language Model Fine-Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models, particularly pre-trained language models, often exhibit biases due to training data, making it challenging to ensure fairness in decision-making processes, especially in domains like credit or employment, where regulatory bodies require bias mitigation, but separate organizations responsible for fairness and model development face cooperation difficulties.

Innovation Solution

A method involving the identification and masking of identity elements and their proxies in downstream data before fine-tuning pre-trained language models, using normalized point-wise mutual information to determine correlations and applying element dropout, which generates masked data for training to reduce bias.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning models are trained on existing training data, then model performance is improved, but bias and unfairness are introduced

Engineering Contradiction:
Improvemodel performanceVSAvoidbias and unfairness
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by identifying and masking identity elements and their proxies in the training data before the fine-tuning process begins. This pre-processing step removes biased information from the data input, allowing the model to learn without inheriting harmful biases from the original training data while still maintaining performance on the downstream task.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and removes identity elements (such as gender, race, age) and their proxy terms from the training data through masking. By taking out these biased elements before training, the model can process the remaining data without the harmful associations that would otherwise be learned from the training corpus.

Inventive Principle:
Principle #2Taking out (Extraction)

2Object-affected harmful factors

If fairness constraints are imposed by separate organizations, then fairness is improved, but cooperation and implementation difficulty increase

Engineering Contradiction:
Improvebias and unfairnessVSAvoidcooperation and implementation
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the model development organization to autonomously identify, mask, and remove biased elements from their own training data without requiring external intervention from fairness enforcement organizations. The system uses automated detection of identity elements and proxies, allowing the organization to comply with fairness requirements independently.

Inventive Principle:
Principle #25Self-service

3Object-affected harmful factors

If identity elements are masked in training data, then bias is reduced, but information loss may occur

Engineering Contradiction:
ImprovebiasVSAvoidinformation loss
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies local quality by selectively masking only identity elements and their proxies while leaving the rest of the training data intact. This targeted approach ensures that biased information is removed without unnecessarily discarding useful information from the data, thereby reducing bias while minimizing information loss.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20230409969A1Providing Fairness in Fine-Tuning of Pre-Trained Language Models
Publication Date: 2023.12.21 ORACLE INT CORP
  • US20230409969A1 patent drawing
  • US20230409969A1 patent drawing
  • US20230409969A1 patent drawing

AI summary

Bias in a language model generated through fine tuning of a pre-trained language model may be mitigated, whether the bias may be incorporated in the pre-trained language model or in fine-tuning data. A pre-trained language model may be fine-tuned using downstream training data. Prior to tuning, elements within the downstream data may be identified that either match or serve as proxies for one or more identity elements associated with training bias sensitivity. Proxy elements may be identified using an analysis of distributions of the downstream elements and distributions of identity elements. Once the elements are identified, instances of the identified elements may be replaced in the downstream data with one or more masking element to generate masked downstream data. A fine-tuned language model with reduced bias may then be generated from the pre-trained language model by tuning the pre-trained language model using the masked downstream data.