Simultaneous Interpretation Model Using External Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional real-time simultaneous interpretation systems face challenges in providing prompt and accurate interpretation results, as they typically wait until a speaker finishes a sentence unit before delivering the interpretation, leading to frustration and inefficiency, especially when dealing with languages that have significant differences in word order.

Innovation Solution

A method and system that utilize external alignment information to train a real-time simultaneous interpretation model, allowing for the generation of intermediate interpretation results by determining read and write actions based on word alignment between source and target languages, enabling immediate output of interpretation as soon as a word is spoken.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system waits until the speaker finishes a sentence unit before delivering interpretation, then the accuracy of interpretation is improved, but the interpretation delay increases

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidinterpretation delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the interpretation process into intermediate results and final results. Intermediate interpretation results are generated for partial sentence units or phrases before the complete sentence is finished, while the final result provides the complete interpretation after sentence completion. This segmentation allows listeners to receive timely partial information without sacrificing final accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary interpretation actions by generating intermediate interpretation results during the speech input phase, before the speaker completes the full sentence unit. These preliminary results are based on partial input analysis and provide early feedback to listeners, reducing the overall interpretation delay while maintaining accuracy through subsequent refinement.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the system provides intermediate interpretation results during speech input, then the interpretation promptness is improved, but the interpretation accuracy may deteriorate

Engineering Contradiction:
Improveinterpretation promptnessVSAvoidinterpretation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements a feedback mechanism where intermediate interpretation results are generated and provided to listeners during speech input, then refined and corrected based on the complete sentence unit input. The final interpretation result incorporates feedback from the full context, ensuring accuracy while maintaining promptness through the intermediate output.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The interpretation system operates dynamically by adjusting its output strategy based on the completion status of sentence units. During partial input phases, the system dynamically generates intermediate results for promptness; upon sentence completion, it dynamically refines the output to ensure accuracy. This dynamic behavior allows the system to adapt between speed and precision requirements.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12019997B2Method of training real-time simultaneous interpretation model based on external alignment information, and method and system for simultaneous interpretation based on external alignment information
Publication Date: 2024.06.25 ELECTRONICS & TELECOMM RES INST
  • US12019997B2 patent drawing
  • US12019997B2 patent drawing
  • US12019997B2 patent drawing

AI summary

Provided is a method of training a real-time simultaneous interpretation model based on external alignment information, the method including: receiving a bilingual corpus having a source language sentence as an input text and a target language sentence as an output text; generating alignment information corresponding to words or tokens (hereinafter, words) in the bilingual corpus; determining a second action following a first action in the simultaneous interpretation model on the basis of the alignment information to generate action sequence information; and training the simultaneous interpretation model on the basis of the bilingual corpus and the action sequence information, wherein the first action and the second action represent a read action of reading the word in the input text or a write action of outputting an intermediate interpretation result corresponding to the read action performed up to a present.