Emotion-Based Sound Effect Matching for Audiobook Production

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The manual selection of sound effects for audio books is labor-intensive and time-consuming, leading to high repetition rates and inconsistent quality, especially in mass production scenarios.

Innovation Solution

An automated method using an emotion determination model to identify statement emotion labels, determine emotion offset values, and select target sound effects based on emotion probability distributions, allowing for efficient and consistent addition of sound effects to audio texts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual selection of sound effects is used, then sound effect quality can be controlled, but labor cost and time consumption increase significantly

Engineering Contradiction:
Improvesound effect qualityVSAvoidproduction efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs automatic sound effect selection through emotion determination models and probability distribution calculations, enabling the process to serve itself without manual intervention. The automated selection based on text emotion analysis eliminates the need for human operators to manually choose sound effects, thereby resolving the contradiction between quality control and production efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical selection process with an automated computational system that uses emotion determination models and probability distributions. This substitution of human manual operation with an automated algorithmic system maintains sound effect quality while dramatically improving productivity by eliminating labor-intensive manual selection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manual selection of sound effects is used, then sound effect selection can be customized, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvesound effect customizationVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The automated system analyzes text emotion and automatically selects appropriate sound effects without requiring manual customization input. The system serves itself by determining emotion labels, calculating probability distributions, and selecting sound effects based on predefined criteria, thereby maintaining adaptability while eliminating time loss associated with manual selection.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary emotion analysis and sound effect selection automatically before the actual audio production process. By pre-determining the appropriate sound effects based on text emotion labels and probability distributions, the system eliminates the need for time-consuming manual customization during the production phase.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual selection of sound effects is used, then sound effect selection can be precise, but the repetition rate of sound effects becomes very high

Engineering Contradiction:
Improvesound effect selection precisionVSAvoidsound effect quality consistency
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The system uses emotion determination models that analyze text content and provide feedback through emotion labels and probability distributions. This feedback mechanism ensures that sound effect selections are precisely matched to the emotional context of each text segment, preventing repetitive use of the same sound effects and maintaining quality consistency across different audio productions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the selection parameter from subjective manual judgment to objective emotion-based probability distribution. By using emotion labels and calculated probability values as selection criteria, the system achieves precise and consistent sound effect selection that adapts to different text emotions, eliminating the high repetition rate problem while maintaining quality precision.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If automated sound effect selection is implemented, then productivity improves, but the complexity of the system increases

Engineering Contradiction:
Improvesound effect addition efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The automated system is segmented into distinct functional modules: emotion determination model, probability distribution calculation module, and sound effect selection module. This segmentation allows each component to perform a specific function independently, improving overall productivity while managing system complexity through modular design that can be developed and maintained separately.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12511485B2Sound effect adding method and apparatus, storage medium, and electronic device
Publication Date: 2025.12.30 DOUYIN VISION CO LTD
  • US12511485B2 patent drawing
  • US12511485B2 patent drawing
  • US12511485B2 patent drawing

AI summary

The present invention relates to a sound effect adding method and apparatus, a storage medium, and an electronic device. The method comprises: determining, on the basis of an emotion judgment model, a statement emotion label of each statement of a text to be processed; determining an emotion offset value of said text on the basis of the type of emotion labels which are largest in quantity among the multiple statement emotion labels; for each paragraph of said text, determining an emotion distribution vector of the paragraph according to the statement emotion label of at least one statement corresponding to the paragraph; determining emotion probability distribution of the paragraph on the basis of the emotion offset value and the emotion distribution vector corresponding to the paragraph; determining, according to the emotion probability distribution of the paragraph and sound effect emotion labels of multiple sound effects in a sound effect library, a target sound effect matching the paragraph; and adding the target sound effect to an audio position corresponding to the paragraph in an audio file corresponding to said text. Thus, the effect of automatically selecting and adding sound effects can be implemented, and the efficiency of adding sound effects can be improved.