Personalized Audio Generation Using User Speech Pattern Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems fail to create personalized audio content that simulates the speech features of individual users, lacking emotional intimacy and trust, as they typically use standardized audio for all users.
Innovation Solution
Obtain speech pattern data from users, including pronunciation, speech rate, and grammatical features, and convert broadcast text into audio content with these personalized features, enhancing the audio with facial and behavioral expressions to match the user's characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If standardized audio is used for all users, then device complexity is reduced and ease of manufacture is improved, but adaptability and user personalization deteriorate
Solution Approach 1:
The audio content delivery system is segmented into multiple voice models corresponding to different user types. Instead of using a single standardized audio for all users, the system divides the audio resources into distinct segments (different voice models) that can be selectively applied based on user characteristics, thereby maintaining ease of manufacture while improving adaptability.
Solution Approach 2:
Different voice models with distinct characteristics (gender, age, emotion, accent) are assigned to different user segments. This local quality approach ensures that each user group receives audio content tailored to their specific preferences and characteristics, enhancing personalization without requiring complete customization for every individual user.
2Adaptability or versatility
If personalized speech features are extracted and applied, then adaptability and user intimacy are improved, but device complexity and data processing requirements increase
Solution Approach 1:
Speech pattern data is extracted and stored in advance during user interaction phases. By performing this data extraction and analysis beforehand, the system avoids the need for complex real-time processing during audio content delivery, thereby reducing device complexity while maintaining high adaptability through pre-prepared personalized audio models.
Solution Approach 2:
Instead of implementing complex real-time speech synthesis and analysis capabilities in the device, the system creates copies of user speech patterns as pre-processed voice models. These copied speech features are stored and directly applied during audio content delivery, significantly reducing the computational complexity required at the device level while maintaining personalized adaptability.
3Measurement precision
If speech pattern data is collected and analyzed, then measurement precision of user characteristics is improved, but loss of user information and privacy concerns increase
Solution Approach 1:
The system extracts only the essential speech pattern features (pronunciation characteristics, speech rate, accent patterns) needed for audio personalization, separating these from other user information. By taking out only the necessary elements for the specific purpose of audio content adaptation, the system achieves high measurement precision for speech characteristics while minimizing the collection and potential loss of broader user information.
Data Source
AI summary
A data processing method is provided. The method includes: obtaining a speech pattern data of a target user based on a speech information of the target user, where the speech pattern data indicates a speech feature of the target user; and converting a broadcast text into an audio content based on the speech pattern data, where a text of the audio content corresponds to the broadcast text, and the audio content has the speech feature.


