Real-Time Interpretation Unit Extraction for Continuous Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic translation and interpretation devices fail to accurately translate real-time continuous speech due to their reliance on sentence units, which often do not align with the natural flow of speech in scenarios like phone conversations or lectures, where pauses are not reliable indicators of sentence boundaries.

Innovation Solution

A device and method that utilize a voice recognition module to recognize voice units as sentence units, a real-time interpretation unit extraction module to form these units into interpretation units, and a real-time interpretation module to perform accurate translation, considering characteristics like lexical, morphological, acoustic, and time features to handle continuous speech effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If translation is performed for each sentence unit using conventional methods, then translation accuracy for standard sentences is improved, but translation accuracy for real-time continuous speech deteriorates

Engineering Contradiction:
Improvetranslation accuracyVSAvoidadaptability to real-time speech
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments real-time continuous speech into voice units based on pause detection, then dynamically groups these voice units into interpretation units. This segmentation approach allows the system to handle both standard sentences and continuous speech by creating appropriate boundaries based on actual speech patterns rather than relying on pre-defined sentence structures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The interpretation unit formation is dynamic rather than static. The system adjusts the grouping of voice units into interpretation units based on real-time speech characteristics, allowing the same speech to be processed differently depending on whether it contains pauses, is spoken continuously, or contains multiple sentences. This dynamic adaptation resolves the contradiction between maintaining accuracy for standard sentences and adapting to varied real-time speech patterns.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If pause is used as a standard for determining sentence units, then sentence segmentation is simplified, but correct translation of continuous speech without reliable pauses deteriorates

Engineering Contradiction:
Improvesentence segmentation simplicityVSAvoidsentence unit identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system introduces voice units as an intermediary layer between raw speech and final interpretation units. Voice units are initially segmented using simple pause detection (maintaining ease of operation), then multiple voice units are grouped into interpretation units using contextual analysis (improving accuracy). This intermediary approach allows the system to combine simple pause-based segmentation with more sophisticated grouping logic to resolve the contradiction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary segmentation into voice units based on pauses, then performs a second-level grouping into interpretation units. This preliminary action allows the system to first apply simple pause-based segmentation and then refine the results by combining voice units that belong together semantically, even if they lack clear pause boundaries. This two-stage approach maintains operational simplicity while improving identification accuracy.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If speech is divided into multiple voice units based on pauses, then processing of continuous speech is facilitated, but correct formation of interpretation units deteriorates when speech lacks proper pause structure

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidinterpretation unit formation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system merges multiple voice units into interpretation units based on contextual and semantic analysis. When speech lacks proper pause structure, the system combines adjacent voice units that form a coherent semantic unit, using linguistic features and contextual information to determine appropriate boundaries. This merging process maintains processing efficiency by working with pre-segmented voice units while improving accuracy through intelligent combination strategies.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses feedback from contextual analysis and semantic evaluation to adjust the formation of interpretation units from voice units. When initial pause-based segmentation produces inappropriate boundaries, the system applies feedback mechanisms to re-evaluate and re-group voice units, using linguistic features and contextual information to correct segmentation errors and form accurate interpretation units.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10366173B2Device and method of simultaneous interpretation based on real-time extraction of interpretation unit
Publication Date: 2019.07.30 HYUNDAI MOTOR CO LTD
  • US10366173B2 patent drawing
  • US10366173B2 patent drawing
  • US10366173B2 patent drawing

AI summary

The present invention relates to a device of simultaneous interpretation based on real-time extraction of an interpretation unit, the device including a voice recognition module configured to recognize voice units as sentence units or translation units from vocalized speech that is input in real time, a real-time interpretation unit extraction module configured to form one or more of the voice units into an interpretation unit, and a real-time interpretation module configured to perform an interpretation task for each interpretation unit formed by the real-time interpretation unit extraction module.