Language Model Diversity for Logical Inference Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating natural language sentences from logical formulas representing inference results are limited in diversity, failing to produce sentences that are more varied than the documents used to generate the inference rules, especially in cases involving sensitive information or simpler sentence generation.

Innovation Solution

An information processing apparatus and method that acquires a logical formula representing an inference result and generates a natural language sentence using a language model based on a corpus not containing the document used to generate the inference rule, ensuring the sentence is more diverse and potentially simpler.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a language model is trained on the same document used to generate the inference rule, then the generated natural language sentence accurately reflects the original document's expression, but the diversity of the generated sentence is limited

Engineering Contradiction:
Improveaccuracy of reflectionVSAvoidsentence diversity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a corpus as an intermediary data source that is related to the target document but not identical to it. The language model is trained on this corpus instead of the target document itself, allowing the model to learn relevant linguistic patterns and domain knowledge while avoiding direct copying of the original document's expressions, thus achieving both accuracy and diversity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a linguistic copy through the corpus - a related but distinct text source that captures the essence and domain characteristics of the target document without being the same document. This allows the language model to learn from a representative sample while generating diverse expressions that reflect the original document's meaning

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If a language model is trained on a large corpus to generate diverse sentences, then the diversity of the generated natural language sentence increases, but the computational resources and training time required increase

Engineering Contradiction:
Improvesentence diversityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent extracts only the necessary linguistic patterns and domain knowledge from a large corpus that are relevant to generating diverse sentences for the specific inference rule. By selectively training on pertinent portions of the corpus rather than the entire corpus, the system achieves diversity while reducing unnecessary computational overhead and training time

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If the generated natural language sentence closely follows the original document's expression, then the sentence accurately represents the inference rule, but the sentence becomes more complex and less simplified

Engineering Contradiction:
Improveaccuracy of representationVSAvoidsentence complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the training parameter from the original document itself to a related corpus, which fundamentally alters how the language model learns expressions. This parameter change allows the model to generate sentences that maintain semantic accuracy while using different linguistic structures and vocabularies, thereby reducing complexity while preserving meaning

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240273310A1Information processing apparatus, information processing method, and storage medium
Publication Date: 2024.08.15 NEC CORP
  • US20240273310A1 patent drawing
  • US20240273310A1 patent drawing
  • US20240273310A1 patent drawing

AI summary

An object to generate, from a logical formula representing an inference result, a natural language sentence that is more diverse as compared with a document used to generate an inference rule is attained. In order to attain the object, an information processing apparatus (1) includes: an acquisition section (11) that acquires a logical formula representing an inference result which is based on an inference rule for an observation event; a generation section (12) that generates a natural language sentence from at least one text element with use of a language model generated on the basis of a corpus which does not contain a document used to generate the inference rule, the at least one text element being included in constituents of the logical formula and being text; and an output section (13) that outputs the natural language.