Speech Synthesizer Voice Selection Using User Satisfaction Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesizers in computing devices typically use a single voice or voice model, resulting in static voice characteristics that do not dynamically adapt to user preferences or situations, potentially reducing user satisfaction.
Innovation Solution
A computing device dynamically selects a voice based on user satisfaction data collected during communications, adjusting voice characteristics such as emotion, intonation, and tone to enhance user satisfaction and achieve desired outcomes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single voice or voice model is used in the speech synthesizer, then the device complexity is reduced and ease of operation is improved, but the adaptability and user satisfaction deteriorate due to static voice characteristics
Solution Approach 1:
The patent implements dynamic voice selection by transitioning from a static single-voice system to a dynamic multi-voice system that automatically selects appropriate voices based on real-time context analysis. The speech synthesizer now includes a voice selection module that dynamically chooses from multiple voice models based on detected communication context, user preferences, and situational factors, making the system adaptable without requiring complex manual configuration
Solution Approach 2:
The system performs self-service by automatically selecting appropriate voices without requiring explicit user commands or manual voice selection. The speech synthesizer autonomously analyzes the communication context, user satisfaction feedback, and situational factors to dynamically choose the most appropriate voice, reducing the operational burden on users while improving adaptability
2Adaptability or versatility
If multiple voice models are provided for user selection, then adaptability is improved, but the ease of operation deteriorates due to requiring user configuration and selection
Solution Approach 1:
The system automatically manages voice selection without requiring user intervention. The speech synthesizer includes contextual analysis components that autonomously determine the appropriate voice based on communication type, user feedback, and situational context, eliminating the need for users to manually configure or select from multiple voice models while maintaining access to diverse voice options
Solution Approach 2:
The system incorporates user satisfaction feedback mechanisms that automatically monitor user responses to voice output and dynamically adjust voice selection accordingly. By analyzing user feedback in real-time, the system learns user preferences and automatically selects voices that maximize user satisfaction, removing the need for explicit user configuration while providing personalized voice experiences
3Manufacturing precision
If the speech synthesizer uses recorded speech from a single voice actor, then manufacturing precision and consistency are improved, but adaptability and user satisfaction deteriorate
Solution Approach 1:
The patent implements a multi-functional voice system where a single speech synthesizer device can dynamically employ multiple voice models with different characteristics (emotion, intonation, gender, pitch, accent, phrasing, tone). This universal system allows the device to adapt to various communication contexts and user preferences while maintaining consistent quality through controlled synthesis parameters across all voice models
Solution Approach 2:
The system dynamically changes voice parameters such as emotion, intonation, gender, pitch, accent, phrasing, and tone by selecting from multiple pre-trained voice models. Each voice model maintains consistent synthesis quality while offering different characteristic parameters, allowing the system to adapt voice characteristics to match communication context and user preferences without sacrificing manufacturing precision
Data Source
AI summary
A computing device having the capability to dynamically select a voice that will be used by a speech synthesizer in creating synthesized speech for use in communicating with a user of the computing device is provided. For example, in some embodiments, the computing device: i) employs the speech synthesizer to have a first audible communication with the user using a first voice; ii) stores user satisfaction data that can be used to determine a user's satisfaction with an action the user took in response to the first audible communication; and iii) determines whether a different voice should be used during a second audible communication with the user based on the stored user satisfaction data.


