Voice Processing Apparatus Adaptive Prosody Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing techniques struggle to appropriately control prosody based on the characteristics of a voice signal, as they rely on fixed reference ranges that fail to account for individual signal variations, leading to inadequate modulation of volume and pitch.
Innovation Solution
A voice processing apparatus that extracts character amounts from a voice signal, calculates difference values relative to reference values, and generates processing values to dynamically control prosody, allowing for emphasis or depression of voice characteristics through variable functions and coefficients.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If fixed reference ranges are used to control volume and pitch, then the processing is simple and consistent, but the prosody control cannot adapt to different voice signal characteristics
Solution Approach 1:
The patent transforms the static fixed reference range approach into a dynamic adaptive system. It calculates the standard deviation of character amounts from the voice signal to dynamically determine reference ranges, allowing the system to adapt to different voice characteristics while maintaining processing efficiency through automated calculations.
Solution Approach 2:
The patent changes the parameter determination method from fixed predetermined values to dynamically calculated values based on signal characteristics. By computing standard deviation and using it to set adaptive reference ranges, the system modifies processing parameters according to the actual voice signal properties, resolving the contradiction between adaptability and complexity.
2Manufacturing precision
If fixed reference ranges are applied to all voice signals, then the processing method is uniform, but no prosody change occurs when character amounts fall within reference ranges
Solution Approach 1:
The patent applies different processing strategies based on local characteristics of the voice signal. By calculating standard deviation and comparing character amounts against dynamically determined reference ranges, it selectively applies prosody control only where needed, improving precision while avoiding unnecessary processing uniformity.
Solution Approach 2:
The system dynamically adjusts reference ranges based on the calculated standard deviation of character amounts. This allows the processing to be effective precisely when needed (when character amounts deviate from the adaptive reference range) and avoids unnecessary processing when already within appropriate bounds, resolving the contradiction between precision and effectiveness.
3Adaptability or versatility
If standard deviation calculation and adaptive reference ranges are used, then prosody control adapts to voice characteristics, but the processing becomes more complex
Solution Approach 1:
The system performs self-characterization by automatically calculating the standard deviation of its own input signal and using this information to adapt its processing parameters. This self-service approach enables adaptability without requiring external configuration or complex manual tuning, as the system autonomously determines appropriate reference ranges from the signal itself.
Solution Approach 2:
The patent implements a feedback mechanism where the calculated standard deviation of character amounts feeds back into the determination of reference ranges. This closed-loop approach allows the system to continuously adapt to voice signal characteristics, achieving high adaptability through a relatively simple feedback-based adjustment rather than complex multi-parameter control.
Data Source
Figure 1~3
Figure 4~5
Figure 6~8
AI summary
Character extraction section (22) extracts character amounts (F), pertaining to a prosody of voice, from a voice signal sequentially in a time-serial manner. Difference value calculation (26) calculates a difference value (D) between each of the extracted character amounts (F) and a reference value (R). Processing values (C), corresponding to the individual character amounts (F), are generated in accordance with the respective difference values (D), and a voice processing section (30) controls the individual character amounts (F) of the voice signal in accordance with the processing values (C) corresponding to the character amounts and thereby generates an output signal having a prosody changed from the prosody of the voice signal.