Text-to-Speech Style Adaptation via Dynamic Weight Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies, such as text-to-speech (TTS), are inadequate for natural dialogue in personal assistant agents and robotics as they fail to adapt to user intent, feelings, and context, leading to insufficient customization of output speech.
Innovation Solution
An electronic device equipped with a processor that acquires text from user speech, determines parameter information for style adjustment using multiple TTS databases and a trained AI model, and synthesizes output speech based on identified weight sets to provide customized and context-aware responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple TTS databases are used to provide various output speech types, then the versatility of speech synthesis is improved, but the device complexity and database size increase significantly
Solution Approach 1:
The patent changes parameters of existing TTS databases dynamically based on user context, feelings, and intent rather than maintaining separate databases for each speech type. By adjusting synthesis parameters (pitch, speed, tone) of a base TTS system, the patent achieves multiple output speech styles without proportionally increasing database size, thus resolving the contradiction between versatility and complexity.
2Adaptability or versatility
If TTS database is expanded to cover more user intents and context information, then the adaptability to user needs is improved, but the difficulty of detecting and measuring appropriate speech style increases
Solution Approach 1:
The patent implements feedback mechanisms that continuously monitor user responses, emotional states, and contextual information to dynamically adjust speech synthesis style. By using feedback from user reactions and context sensors, the system automatically selects appropriate speech styles without requiring complex manual detection and measurement of numerous speech type parameters, thus resolving the contradiction between adaptability and detection difficulty.
Data Source
AI summary
An electronic device and a controlling method of the electronic device are provided. The electronic device acquires text to respond on a received user's speech, acquires a plurality of pieces of parameter information for determining a style of an output speech corresponding to the text based on information on a type of a plurality of text-to-speech (TTS) databases and the received user's speech, identifies a TTS database corresponding to the plurality of pieces of parameter information among the plurality of TTS databases, identifies a weight set corresponding to the plurality of pieces of parameter information among a plurality of weight sets acquired through a trained artificial intelligence model, adjusts information on the output speech stored in the TTS database based on the weight set, synthesizes the output speech based on the adjusted information on the output speech, and outputs the output speech corresponding to the text.


