LLM Distillation for Machine Translation Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current large language models for machine translation require high computational resources and costs, making them inefficient for widespread use.

Innovation Solution

The method involves obtaining a bilingual sentence pair and distilling one of the sentences using a large language model to produce a more efficient and effective translation, thereby reducing computational requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a large language model is used for machine translation, then translation quality is improved, but computational resource consumption and cost increase

Engineering Contradiction:
Improvetranslation qualityVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent creates a distilled translation model that copies the translation capabilities of the large language model but with significantly reduced parameters. The distillation process transfers knowledge from the large model to a smaller model, enabling the smaller model to achieve comparable translation quality while consuming far fewer computational resources.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts the essential translation capabilities from the large language model through the distillation process. By taking out only the necessary translation knowledge and transferring it to a smaller model, the system maintains translation quality while removing the excessive computational overhead of the original large model.

Inventive Principle:
Principle #2Taking out (Extraction)

2Manufacturing precision

If a large language model is used for machine translation, then translation quality is improved, but cost increases

Engineering Contradiction:
Improvetranslation qualityVSAvoidcost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent creates a distilled translation model that copies the translation capabilities of the large language model but with significantly reduced parameters. The distillation process transfers knowledge from the large model to a smaller model, enabling the smaller model to achieve comparable translation quality while consuming far fewer computational resources.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the expensive large language model with a cheaper distilled model that can be deployed more economically. The distilled model, having been trained on distilled data, provides comparable translation quality at a fraction of the computational cost and deployment expense.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS20250094739A1Training method and apparatus for full atomic structure prediction model, and electronic device
Publication Date: 2025.03.20 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20250094739A1 patent drawing
  • US20250094739A1 patent drawing
  • US20250094739A1 patent drawing

AI summary

An information processing method. The method includes obtaining a first bilingual sentence pair, in which the first bilingual sentence pair comprises a source language sentence and a target language sentence; and obtaining a distilled second bilingual sentence pair by distilling a first language sentence in the first bilingual sentence pair based on a large language model (LLM), in which the first language sentence is the source language sentence or the target language sentence.