Personalized Audio Generation Using User Speech Pattern Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems fail to create personalized audio content that simulates the speech features of individual users, lacking emotional intimacy and trust, as they typically use standardized audio for all users.

Innovation Solution

Obtain speech pattern data from users, including pronunciation, speech rate, and grammatical features, and convert broadcast text into audio content with these personalized features, enhancing the audio with facial and behavioral expressions to match the user's characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If standardized audio is used for all users, then device complexity is reduced and ease of manufacture is improved, but adaptability and user personalization deteriorate

Engineering Contradiction:
Improveease of manufactureVSAvoidadaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The audio content delivery system is segmented into multiple voice models corresponding to different user types. Instead of using a single standardized audio for all users, the system divides the audio resources into distinct segments (different voice models) that can be selectively applied based on user characteristics, thereby maintaining ease of manufacture while improving adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different voice models with distinct characteristics (gender, age, emotion, accent) are assigned to different user segments. This local quality approach ensures that each user group receives audio content tailored to their specific preferences and characteristics, enhancing personalization without requiring complete customization for every individual user.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If personalized speech features are extracted and applied, then adaptability and user intimacy are improved, but device complexity and data processing requirements increase

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Speech pattern data is extracted and stored in advance during user interaction phases. By performing this data extraction and analysis beforehand, the system avoids the need for complex real-time processing during audio content delivery, thereby reducing device complexity while maintaining high adaptability through pre-prepared personalized audio models.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of implementing complex real-time speech synthesis and analysis capabilities in the device, the system creates copies of user speech patterns as pre-processed voice models. These copied speech features are stored and directly applied during audio content delivery, significantly reducing the computational complexity required at the device level while maintaining personalized adaptability.

Inventive Principle:
Principle #26Copying

3Measurement precision

If speech pattern data is collected and analyzed, then measurement precision of user characteristics is improved, but loss of user information and privacy concerns increase

Engineering Contradiction:
Improvemeasurement precisionVSAvoidloss of information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts only the essential speech pattern features (pronunciation characteristics, speech rate, accent patterns) needed for audio personalization, separating these from other user information. By taking out only the necessary elements for the specific purpose of audio content adaptation, the system achieves high measurement precision for speech characteristics while minimizing the collection and potential loss of broader user information.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12499868B2Data processing method
Publication Date: 2025.12.16 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12499868B2 patent drawing
  • US12499868B2 patent drawing
  • US12499868B2 patent drawing

AI summary

A data processing method is provided. The method includes: obtaining a speech pattern data of a target user based on a speech information of the target user, where the speech pattern data indicates a speech feature of the target user; and converting a broadcast text into an audio content based on the speech pattern data, where a text of the audio content corresponds to the broadcast text, and the audio content has the speech feature.