Speech Synthesis Pitch Shift for Natural Response Tone
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice interaction systems often produce responses that sound unnatural to users, leading to a perception that a machine is speaking, and may deteriorate in auditory quality due to excessive pitch shifting.
Innovation Solution
A speech synthesis device and method that detects the pitch of a representative portion of a user's question and shifts the pitch of the response to maintain a consonant-interval relationship, ensuring the response voice is synthesized with a natural and high-quality tone by adjusting the pitch shift amount within predetermined limits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If pitch shifting is applied to synthesize response voice, then the naturalness of the response is improved, but the auditory quality deteriorates
Solution Approach 1:
The patent applies parameter changes by adjusting the pitch shift amount dynamically based on the relationship between question pitch and response pitch. Instead of using a fixed pitch shift amount, the system calculates an optimal shift amount that maintains consonant intervals, thereby improving naturalness while preserving auditory quality through controlled parameter adjustment.
2Ease of operation
If pitch is shifted to match consonant-interval relationship, then the comfort feeling is improved, but the pitch accuracy deteriorates
Solution Approach 1:
The system employs feedback by continuously monitoring the pitch relationship between question and response, calculating the appropriate pitch shift amount to maintain consonant intervals. This feedback mechanism ensures that pitch adjustments enhance comfort feeling while preserving pitch accuracy through dynamic adaptation to the specific question-response pair.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This invention is an improvement of technology for automatically generating response voice to voice uttered by a speaker (user), and is characterized by controlling a pitch of the response voice in accordance with a pitch of the speaker's utterance. A voice signal of the speaker's utterance (e.g., question) is received (102), and a pitch (e.g., highest pitch) of a representative portion of the utterance is detected (106). Voice data of a responsive to the utterance is acquired (110, 124), and a pitch (e.g., average pitch) based on the acquired response voice data is acquired (112). Then, a pitch shift amount for shifting the acquired pitch to a target pitch having a particular relationship to the pitch of the representative portion is determined (114). When response voice is to be synthesized on the basis of the response voice data, the pitch of the response voice to be synthesized is shifted in accordance with the pitch shift amount (116).