Automatic Soundtrack Generation for Speech Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The production of soundtrack-enhanced audiobooks is hindered by the significant time and cost associated with manually selecting and compiling music files to overlay narration, making such enhancements rare due to complexity and cost.

Innovation Solution

A method for automatically generating digital soundtracks synchronized with speech audio using natural language processing and semantic analysis to identify emotional profiles in text data, matching audio regions with corresponding emotional profiles from an audio database, and dynamically adjusting soundtrack playback to maintain synchronization with varying narration speeds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If manual selection and compilation of music files is used, then soundtrack quality can be controlled, but production time and cost increase significantly

Engineering Contradiction:
Improveproduction timeVSAvoidmanual selection process
Core Design Contradiction:
Ease of manufactureVSExtent of automation

Solution Approach 1:

The system enables self-service by automatically analyzing the audiobook content through NLP and semantic analysis to generate emotional profiles, then autonomously selecting and compiling appropriate music files without human intervention. The automated system processes text data, generates emotional category identifiers, and matches them with corresponding music tracks, eliminating the need for manual soundtrack creation while maintaining quality through algorithmic decision-making.

Inventive Principle:
Principle #25Self-service

2Productivity

If automatic soundtrack generation is implemented, then production time is reduced, but the complexity of the system increases

Engineering Contradiction:
Improvesoundtrack creation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex soundtrack generation process into distinct manageable modules: (1) NLP processing to extract text data from audiobooks, (2) semantic analysis to generate emotional profiles and category identifiers, (3) music database querying to select appropriate tracks, and (4) audio compilation to generate the final soundtrack. This segmentation allows each component to be developed, tested, and optimized independently, reducing overall system complexity while maintaining high productivity.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If emotional profile matching is used, then soundtrack relevance to content is improved, but processing time increases

Engineering Contradiction:
Improvesoundtrack synchronization accuracyVSAvoidemotional analysis processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing the audiobook text data through NLP and semantic analysis before music selection. Emotional profiles and category identifiers are generated in advance and stored, allowing the music database to be queried efficiently using these pre-computed emotional categories. This preliminary processing enables fast matching between content emotions and music tracks during the actual soundtrack generation, reducing real-time processing time while maintaining high synchronization accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10698951B2Systems and methods for automatic-creation of soundtracks for speech audio
Publication Date: 2020.06.30 BOOKTRACK HLDG
  • US10698951B2 patent drawing
  • US10698951B2 patent drawing
  • US10698951B2 patent drawing

AI summary

A method of automatically generating a digital soundtrack intended for synchronised playback with associated speech audio, the method executed by a processing device or devices having associated memory. The method comprises syntactically and/or semantically analysing text representing or corresponding to the speech audio at a text segment level to generate an emotional profile for each text segment in the context of a continuous emotion model. The method further comprises generating a soundtrack for the speech audio comprising one or more audio regions that are configured or selected for playback during corresponding speech regions of the speech audio, and wherein the audio configured for playback in the audio regions is based on or a function of the emotional profile of one or more of the text segments within the respective speech regions.