Masked Word Inference Using Control Words for Ethical Text Output
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Language models trained on randomly collected data may generate ethically inappropriate sentences due to the reflection of human biases related to race, gender, ethnicity, and culture.
Innovation Solution
An inference device that includes a mask data acquisition unit, word sequence acquisition unit, control information acquisition unit, and inference unit to infer likely candidate words based on adjectival expressions, ensuring ethically appropriate text generation by using a trained model that integrates pointwise mutual information and N-gram models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If language models are trained on randomly collected text data to improve language understanding accuracy, then the model achieves better language processing capability, but the model generates ethically inappropriate sentences due to human biases
Solution Approach 1:
The patent introduces control words as intermediary elements that mediate between the input text and the language model's generation process. These control words explicitly represent ethical constraints and guide the model to generate appropriate outputs while maintaining language understanding accuracy. The control words act as a bridge that translates ethical requirements into actionable constraints for the model.
Solution Approach 2:
The patent applies preliminary action by pre-processing the input text to extract and identify control words before the main language model processing. This preliminary step prepares the ethical constraints in advance, allowing the model to incorporate them during generation. The control words are identified and prepared beforehand to guide the subsequent text generation process.
2Object-affected harmful factors
If control words are integrated into the inference process to ensure ethical appropriateness, then ethical appropriateness is improved, but the device complexity increases
Solution Approach 1:
The patent applies universality by designing the control word mechanism to handle multiple ethical constraints simultaneously through a single unified approach. The same control word extraction and application process works across different types of ethical requirements (race, gender, ethnicity, culture), eliminating the need for separate handling mechanisms for each constraint type.
Solution Approach 2:
The patent segments the inference process into distinct functional units: control word acquisition, control word integration, and text generation. This segmentation allows each component to be independently optimized and maintained, reducing overall system complexity despite the added ethical constraints. The mask data acquisition unit and control information acquisition unit operate as separate modular components.
Data Source
AI summary
The purpose is to obtain an inference device capable of generating an ethically appropriate sentence. The inference device according to the present disclosure includes: a mask data acquisition unit to acquire a character string which includes a masked portion; a word sequence acquisition unit to segment the character string into words and acquires a word sequence including a plurality of words; a control information acquisition unit to acquire an adjectival expression representing a nature or a state of a thing as a control word; and an inference unit to infer a likely candidate word for the masked portion from the control word and the word sequence and output the character string in which the masked portion is replaced with the likely candidate word.


