Language Model Adaptation via Multi-Layer Auxiliary Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing model adaptation technologies for language models suffer from poor adaptability due to reliance on a single type of neural network and require correct genre labels for appropriate learning, making them ineffective for language models.

Innovation Solution

The proposed solution involves an apparatus with a first neural network unit that transforms input symbols into intermediate states and a second neural network unit that uses these intermediate states and different pieces of auxiliary information to predict the next symbol, with multiple hidden layers receiving distinct auxiliary information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single type of neural network is used for model adaptation, then the device complexity is reduced, but the adaptability of the language model deteriorates

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The neural network is divided into multiple hidden layers, where each layer processes different pieces of auxiliary information. This segmentation allows the model to handle diverse adaptation tasks without requiring a completely different neural network architecture for each task, thus improving adaptability while controlling complexity through modular layer design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple hidden layers that process auxiliary information from different dimensions or perspectives. By adding this dimensional aspect to the neural network structure, the model can capture diverse patterns in adaptation data without fundamentally changing the core network architecture, thereby enhancing adaptability without proportionally increasing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If correct genre labels are required for learning, then the measurement precision of adaptation is improved, but the ease of operation deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces auxiliary information as an intermediary element that bridges the gap between input symbols and output predictions. This auxiliary information acts as a mediator that enables the model to perform adaptation without requiring explicit correct genre labels, thereby maintaining prediction accuracy while improving ease of operation by eliminating the need for labeled data in certain adaptation scenarios.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If diverse auxiliary information is processed in multiple hidden layers, then the adaptability of the language model is improved, but the device complexity increases

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The multiple hidden layers are designed to process different pieces of auxiliary information, making the neural network multi-functional. Each layer can be configured to handle specific types of auxiliary information, allowing a single universal network architecture to perform multiple adaptation tasks. This universality improves adaptability while avoiding the need for separate specialized networks for each adaptation scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12308022B2Apparatus, method, and program for utilizing language model
Publication Date: 2025.05.20 NT T INC
  • US12308022B2 patent drawing
  • US12308022B2 patent drawing
  • US12308022B2 patent drawing

AI summary

Disclosed is a model adaptation technology of a language model with higher adaptability. An aspect of the present disclosure relates to an apparatus includes a first neural network unit that transforms an input symbol and outputs an intermediate state; and a second neural network unit that transforms input auxiliary information and the intermediate state and predicts a symbol following the input symbol, wherein the second neural network unit includes a plurality of hidden layers receiving, as input, the intermediate state and auxiliary information, and pieces of the auxiliary information input to each hidden layer are different from each other.