Dynamic Speech Timing via Music Progression Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing systems fail to enhance user experience by only providing navigation information at specific intervals during music reproduction, limiting the entertainment value and realism.

Innovation Solution

A speech processing apparatus and method that dynamically determines output time points based on music progression data, allowing for diverse speech outputs at various time points during music playback, using templates with attribute values and timing data to generate dynamic speech content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech is output only at fixed intervals during music reproduction, then navigation information can be provided to users, but user experience and entertainment value are limited

Engineering Contradiction:
Improvespeech output timing flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system transitions from fixed interval speech output to dynamic timing based on music progression analysis. The determination unit dynamically selects output timing by analyzing music progression data to identify appropriate moments (verses, choruses, bridges) rather than using predetermined fixed intervals, thereby improving adaptability while managing complexity through automated analysis

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of speech output timing from static (fixed intervals) to variable (music progression-dependent). By utilizing music progression data containing temporal information about musical structures, the system adjusts speech timing parameters to match the dynamic characteristics of different music segments, enhancing user experience without requiring manual configuration

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speech is output at multiple time points during music progression, then entertainment value and realism are improved, but speech timing precision requirements increase

Engineering Contradiction:
Improvespeech content diversityVSAvoidspeech output timing precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system introduces music progression data as an intermediary between the music reproduction and speech output processes. This data structure contains pre-analyzed temporal information about musical segments (verses, choruses, bridges), serving as a mediator that guides speech timing decisions without requiring real-time complex analysis, thus achieving precise timing while maintaining system feasibility

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The music progression data is prepared in advance, containing predetermined temporal information about musical structures. By performing the analysis of appropriate speech timing points beforehand and storing this information in the music progression data, the system avoids the need for complex real-time decision-making during music reproduction, achieving high timing precision through pre-computed guidance

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10229669B2Apparatus, process, and program for combining speech and audio data
Publication Date: 2019.03.12 SONY GROUP CORP
  • US10229669B2 patent drawing
  • US10229669B2 patent drawing
  • US10229669B2 patent drawing

AI summary

There is provided a speech processing apparatus including: a data obtaining unit which obtains music progression data defining a property of one or more time points or one or more time periods along progression of music; a determining unit which determines an output time point at which a speech is to be output during reproducing the music by utilizing the music progression data obtained by the data obtaining unit; and an audio output unit which outputs the speech at the output time point determined by the determining unit during reproducing the music.