Personalized Voice Model for Text Message Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems lack personalization in delivering text messages, failing to provide personalized audio playback and user customization, particularly in producing voices based on limited audio input.

Innovation Solution

A system that receives a plurality of speech inputs from a user to create a voice model, which is then used to provide personalized audio output for messages received from another user, allowing the message to be played back in the voice of the sender.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional text message delivery systems are used, then system simplicity is maintained, but user engagement and personalization are insufficient

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing speech samples from users in advance and pre-training voice models before messages are delivered. This allows the system to have personalized voice models ready when needed, rather than creating them at the moment of message delivery, thus reducing real-time complexity while maintaining personalization capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates simplified copies of users' voices through voice models that can be stored and reused. Instead of requiring complex real-time voice synthesis, the system uses pre-generated voice copies to transform text messages into personalized audio, reducing the complexity of on-demand voice generation while maintaining personalization

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If personalized voice models are created from limited audio input, then user customization is enabled, but voice model accuracy may be compromised

Engineering Contradiction:
Improveuser customization capabilityVSAvoidvoice model accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary voice model training using available speech samples before deployment. By pre-processing and pre-training with limited audio input in advance, the system maximizes the quality of voice models from constrained data, improving accuracy without requiring extensive real-time audio collection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system adjusts training parameters and model architecture based on the amount of available audio input. When limited audio data is provided, the system modifies training parameters to optimize model performance under data constraints, balancing personalization capability with voice accuracy through parameter optimization

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If voice models are stored and transmitted between devices, then personalized playback is achieved, but data storage and transmission requirements increase

Engineering Contradiction:
Improvepersonalized playback functionalityVSAvoiddata storage requirement
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The system extracts only the essential voice characteristics from full speech recordings to create compact voice models. By taking out only the critical phonetic and vocal features needed for voice synthesis rather than storing complete audio files, the system achieves personalized playback with significantly reduced data storage requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system stores voice models locally on user devices rather than maintaining centralized repositories of all voice data. This local storage approach enables personalized playback functionality while reducing overall data transmission requirements, as each device only needs to store and process its own user's voice model

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12170089B2Personalized voices for text messaging
Publication Date: 2024.12.17 APPLE INC
  • US12170089B2 patent drawing
  • US12170089B2 patent drawing
  • US12170089B2 patent drawing

AI summary

Systems and processes for operating an intelligent automated assistant are provided. In one example, a plurality of speech inputs is received from a first user. A voice model is obtained based on the plurality of speech inputs. A user input is received from the first user, the user input corresponding to a request to provide access to the voice model. The voice model is provided to a second electronic device.