Audio Playback Device Character Voice Model Assignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio playback devices offer a fixed and monotonous audio presentation, limiting user engagement and interest over long-term use, as they lack the ability to customize voice models for characters in text-based content.

Innovation Solution

An audio playback device and method that allow users to select and assign a target voice model from multiple options to specific characters in a text, transforming sentences into speech according to the chosen voice model, enabling customizable audio presentations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a fixed audio playback mode is used, then the device structure is simple and easy to manufacture, but the audio presentation becomes monotonous and user interest decreases

Engineering Contradiction:
Improvedevice structure simplicityVSAvoidaudio presentation customization
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The audio presentation is segmented into multiple independent voice models, each representing different characters or styles. The system divides the audio output into separate channels that can be independently configured, allowing users to select and combine different voice models for different characters in the text, thereby achieving customization without complicating the overall device structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameters of audio presentation by providing multiple voice models with different characteristics (pitch, tone, speed, etc.). Users can select different voice models and adjust their parameters to create customized audio presentations, transforming the fixed playback mode into a flexible, parameter-adjustable system while maintaining relatively simple device architecture.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If multiple voice models are provided for character customization, then user engagement and interest are enhanced, but the device complexity increases

Engineering Contradiction:
Improvevoice model selectionVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio playback device is designed with multi-functionality by integrating multiple voice models into a single unified system. The device can play back audio using different voice models for different characters, and the same hardware infrastructure supports various playback modes (fixed and customized), reducing the need for separate dedicated components for each function and thereby managing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Instead of creating entirely new hardware components for each voice model, the system uses software-based voice model copies that can be loaded and switched. Multiple voice models are stored as data files or software modules that can be instantiated without duplicating physical hardware, allowing extensive customization while maintaining a relatively simple physical device structure.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If text is transformed into audio with speech of target character using voice model, then audio presentation becomes dynamic and engaging, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio customizationVSAvoidaudio generation time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Voice models are pre-processed and prepared in advance, with their acoustic characteristics and parameters stored in optimized formats. When text needs to be converted to speech, the system can quickly retrieve and apply the pre-prepared voice models without performing complex real-time synthesis, significantly reducing the time required for audio generation while maintaining high customization capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11049490B2Audio playback device and audio playback method thereof for adjusting text to speech of a target character using spectral features
Publication Date: 2021.06.29 INSTITUTE FOR INFORMATION INDUSTRY
  • US11049490B2 patent drawing
  • US11049490B2 patent drawing
  • US11049490B2 patent drawing

AI summary

An audio playback device receives an instruction from a user to select a target voice model from a plurality of voice models and assigns the target voice model to a target character in a text. The audio playback device also transforms the text into a speech, and during the process of transforming the text into the speech, transforms sentences of the target character in the text into the speech of the target character according to the target voice model.