Language Model Diversity for Logical Inference Text
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for generating natural language sentences from logical formulas representing inference results are limited in diversity, failing to produce sentences that are more varied than the documents used to generate the inference rules, especially in cases involving sensitive information or simpler sentence generation.
Innovation Solution
An information processing apparatus and method that acquires a logical formula representing an inference result and generates a natural language sentence using a language model based on a corpus not containing the document used to generate the inference rule, ensuring the sentence is more diverse and potentially simpler.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a language model is trained on the same document used to generate the inference rule, then the generated natural language sentence accurately reflects the original document's expression, but the diversity of the generated sentence is limited
Solution Approach 1:
The patent introduces a corpus as an intermediary data source that is related to the target document but not identical to it. The language model is trained on this corpus instead of the target document itself, allowing the model to learn relevant linguistic patterns and domain knowledge while avoiding direct copying of the original document's expressions, thus achieving both accuracy and diversity
Solution Approach 2:
The patent creates a linguistic copy through the corpus - a related but distinct text source that captures the essence and domain characteristics of the target document without being the same document. This allows the language model to learn from a representative sample while generating diverse expressions that reflect the original document's meaning
2Adaptability or versatility
If a language model is trained on a large corpus to generate diverse sentences, then the diversity of the generated natural language sentence increases, but the computational resources and training time required increase
Solution Approach 1:
The patent extracts only the necessary linguistic patterns and domain knowledge from a large corpus that are relevant to generating diverse sentences for the specific inference rule. By selectively training on pertinent portions of the corpus rather than the entire corpus, the system achieves diversity while reducing unnecessary computational overhead and training time
3Reliability
If the generated natural language sentence closely follows the original document's expression, then the sentence accurately represents the inference rule, but the sentence becomes more complex and less simplified
Solution Approach 1:
The patent changes the training parameter from the original document itself to a related corpus, which fundamentally alters how the language model learns expressions. This parameter change allows the model to generate sentences that maintain semantic accuracy while using different linguistic structures and vocabularies, thereby reducing complexity while preserving meaning
Data Source
AI summary
An object to generate, from a logical formula representing an inference result, a natural language sentence that is more diverse as compared with a document used to generate an inference rule is attained. In order to attain the object, an information processing apparatus (1) includes: an acquisition section (11) that acquires a logical formula representing an inference result which is based on an inference rule for an observation event; a generation section (12) that generates a natural language sentence from at least one text element with use of a language model generated on the basis of a corpus which does not contain a document used to generate the inference rule, the at least one text element being included in constituents of the logical formula and being text; and an output section (13) that outputs the natural language.


