Automatic Interpreter Voice Characteristic Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic interpreters fail to accurately replicate the voice characteristics of an original speaker, particularly in conveying emotions and intentions, as they primarily generate neutral synthetic sounds or limited emotional expressions, which do not faithfully represent the speaker's speech.
Innovation Solution
An automatic interpretation system that includes a speech recognition module to extract pitch, vocal intensity, speech speed, and vocal tract characteristics from the original speaker's voice, and a speech synthesis module to generate synthetic sounds that mimic these characteristics, ensuring the translated speech retains the original speaker's emotional and intentional nuances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing automatic interpreters generate neutral synthetic sounds for translated sentences, then the translation process is simple and fast, but the synthetic speech does not faithfully represent the original speaker's emotions and intentions
Solution Approach 1:
The speech recognition module segments the original speech signal into multiple characteristic components including pitch information, vocal intensity information, speech speed information, and vocal tract characteristic information. This segmentation allows each emotional dimension to be extracted and preserved independently, enabling faithful representation of the original speaker's emotions and intentions in the synthetic output.
Solution Approach 2:
The system changes multiple speech parameters simultaneously to preserve emotional characteristics: pitch contour is extracted and replicated, vocal intensity ratios are calculated and applied to the synthetic speech, speech speed variations are measured and transferred, and vocal tract characteristics are analyzed and reproduced. This multi-parameter approach ensures comprehensive emotional fidelity while maintaining systematic processing.
2Adaptability or versatility
If advanced speech synthesis technologies distinguish voice by gender only, then the synthesis process remains relatively simple, but the emotional expression capability is limited
Solution Approach 1:
The system transitions from one-dimensional gender-based voice synthesis to multi-dimensional emotional synthesis by adding dimensions for pitch contour, vocal intensity, speech speed, and vocal tract characteristics. This dimensional expansion enables rich emotional expression while maintaining a structured approach to complexity management through modular extraction and synthesis processes.
Solution Approach 2:
The speech recognition module performs preliminary extraction of all emotional characteristic information from the original speech signal before the translation process. By pre-extracting pitch information, vocal intensity information, speech speed information, and vocal tract characteristic information, the system prepares emotional templates that guide the synthetic speech generation, ensuring emotional fidelity without adding complexity during the translation phase.
3Reliability
If vocal characteristics of conversational partner are used for emotional synthesis, then emotional information can be added to synthetic sound, but this differs from automatic interpretation that should translate and synthesize the speaker's own speech characteristics
Solution Approach 1:
The system creates an accurate copy of the original speaker's vocal characteristics by extracting pitch information, vocal intensity information, speech speed information, and vocal tract characteristic information from the original speech signal. These extracted characteristics are then replicated in the synthetic speech generation process, producing a faithful copy of the speaker's voice identity and emotional expression without requiring the conversational partner's vocal data.
Data Source
AI summary
Provided are an automatic interpretation system and method for generating a synthetic sound having characteristics similar to those of an original speaker's voice. The automatic interpretation system for generating a synthetic sound having characteristics similar to those of an original speaker's voice includes a speech recognition module configured to generate text data by performing speech recognition for an original speech signal of an original speaker and extract at least one piece of characteristic information among pitch information, vocal intensity information, speech speed information, and vocal tract characteristic information of the original speech, an automatic translation module configured to generate a synthesis-target translation by translating the text data, and a speech synthesis module configured to generate a synthetic sound of the synthesis-target translation.


