Streaming TTS Translation Using Early Recognition Segments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing translation devices and software suffer from inefficiencies in communication due to prolonged waiting times for translated speech, especially in continuous speech, as they rely on final recognition results and synthesis after speech pauses, mimicking manual simultaneous interpretation is challenging.

Innovation Solution

A speech processing method that performs partial recognition and translation in a streaming manner by detecting uncertain recognition results, breaking sentences using punctuation marks, and adjusting speech speed dynamically to reduce waiting times.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the recognition translation engine waits for speech pauses to return final certain recognition results, then the translation accuracy is improved, but the user waiting time is prolonged

Engineering Contradiction:
Improvetranslation accuracyVSAvoiduser waiting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary recognition and translation actions by processing uncertain recognition results as they become available during continuous speech, rather than waiting for speech pauses. The user terminal detects uncertain recognition texts in real-time and triggers translation and speech synthesis before the speaker finishes, thereby reducing waiting time while maintaining acceptable accuracy through continuous improvement of recognition results.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the system processes translation after speech pauses, then the translation reliability is improved, but the communication efficiency is reduced

Engineering Contradiction:
Improvetranslation reliabilityVSAvoidcommunication efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system maintains continuous useful action by continuously processing speech recognition and translation during ongoing speech without interruption for pauses. The user terminal continuously receives uncertain recognition results from the recognition engine and continuously triggers translation and speech synthesis, eliminating idle waiting time and maintaining high communication efficiency while the recognition system continuously refines accuracy.

Inventive Principle:
Principle #20Continuity of useful action

3Loss of time

If the system uses uncertain recognition results for streaming translation, then the user waiting time is reduced, but the recognition stability is challenged

Engineering Contradiction:
Improveuser waiting timeVSAvoidrecognition stability
Core Design Contradiction:
Loss of timeVSStability of the object's composition

Solution Approach 1:

The system implements feedback mechanisms where the user terminal continuously monitors uncertain recognition results and uses this feedback to dynamically adjust translation timing and speech synthesis. The system compares uncertain recognition texts against previously obtained certain recognition texts to determine when to trigger translation, using this feedback loop to balance early translation benefits with recognition stability requirements.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4679417A1Speech processing method for realizing streaming tts
Publication Date: 2026.01.14 SHENZHEN TIMEKETTLE TECH CO LTD
  • EP4679417A1 patent drawingFigure 1
  • EP4679417A1 patent drawingFigure 2
  • EP4679417A1 patent drawingFigure 3

AI summary

The present disclosure relates to a processing method for implementing a streaming TTS speech text, which comprises: a user beginning to speak and performing voice wake-up on a user terminal; the user terminal sending a streaming voice data packet to a recognition engine for speech recognition; the recognition engine continuously responding with uncertain recognition texts to the user terminal; the user terminal detecting and recognizing the uncertain recognition texts, and obtaining certain recognition texts in advance; the user terminal sending the predetermined certain recognition texts to a translation engine to perform translation and complete speech synthesis; the user terminal playing the synthesized speech. By partially determining recognition results in advance and performing translation and speech synthesis on the determined recognition results in a streaming manner, the method in the disclosure reduces the waiting time of the user before hearing the translated speech, thereby achieving the effect of manual simultaneous interpretation.