Speech Translation Preserving Prosodic Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-to-speech machine translation processes lose valuable non-text information such as emotional expressions and prosodic features during translation, resulting in a lack of understanding of the original speaker's intended meaning.

Innovation Solution

A method and apparatus that extract non-text information from source speech, translate the speech while preserving this information, and adjust the translated speech to maintain the original non-text features, including emotional expressions and prosodic information, to create a more accurate target speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If speech synthesis technology is used to translate text into target speech, then text translation can be achieved, but non-text information such as emotional expressions and prosodic features is lost

Engineering Contradiction:
Improvenon-text informationVSAvoidtranslation system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The translation system is segmented into multiple independent modules: a non-text information extraction module that separates prosodic features and emotional expressions from the speech signal, a text translation module that handles linguistic translation, and a speech synthesis module that reconstructs target speech. This segmentation allows each module to specialize in preserving specific information types without increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary extraction of non-text information (prosodic features and emotional expressions) from the source speech before the translation process begins. This extracted information is then preserved and applied during the synthesis stage, ensuring that emotional and prosodic characteristics are maintained in the target speech without interfering with the text translation process.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If only translated text information is used for speech synthesis, then translation process is simple, but the real meaning of the speaker cannot be fully understood

Engineering Contradiction:
Improvetranslation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The processing system is divided into distinct functional segments: one dedicated to extracting and analyzing non-text information (emotional expressions and prosodic features), another for text translation, and a final segment for synthesizing target speech that integrates both text and non-text information. This segmentation enables comprehensive meaning preservation while maintaining manageable system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The translation system is designed with multi-functionality, handling both text translation and non-text information preservation through integrated modules. The speech synthesis module, for example, serves dual purposes: generating target speech from translated text while simultaneously incorporating preserved emotional expressions and prosodic features, thereby enhancing translation reliability without requiring entirely separate systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If non-text information is extracted and preserved during speech translation, then speaker meaning is better understood, but translation process becomes more complex

Engineering Contradiction:
Improveinformation preservationVSAvoidtranslation system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The translation system is segmented into specialized modules: a non-text information extraction module that identifies and separates prosodic features and emotional expressions from the speech signal, a text translation module for linguistic processing, and a synthesis module that integrates both information types. This segmentation preserves comprehensive speaker meaning while managing complexity through modular, independent processing units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary extraction and analysis of non-text information (emotional expressions and prosodic features) from the source speech before the main translation process. This extracted information is stored and then applied during the target speech synthesis stage, ensuring complete information preservation without adding complexity to the core translation workflow.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9342509B2Speech translation method and apparatus utilizing prosodic information
Publication Date: 2016.05.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9342509B2 patent drawing
  • US9342509B2 patent drawing

AI summary

A method and apparatus for speech translation. The method includes: receiving a source speech; extracting non-text information in the source speech; translating the source speech into a target speech; and adjusting the translated target speech according to the extracted non-text information so that the target speech preserves the non-text information in the source speech. The apparatus includes: a receiving module for receiving source speech; an extracting module for extracting non-text information in the source speech; a translation module for translating the source speech into a target speech; and an adjusting module for adjusting the translated target speech according to the extracted non-text information so that the target speech preserves the non-text information in the source speech.