Audio Processing System for Synchronized Sound Effects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for triggering sound effects during reading are prone to errors, fail to adapt to the reader's emotions and reading style, and do not account for physical movements, leading to a lack of immersion and personalization in the reading experience.

Innovation Solution

A process that synchronizes sound effects with real-time reading by determining a speaker's position index and considering prosody and movement data, using preconfigured algorithms to adapt sound effects dynamically, allowing for a more immersive and personalized experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If playback speed estimation is used to trigger sound effects, then the triggering is simple, but the accuracy of sound effect timing is poor

Engineering Contradiction:
Improvesimplicity of sound effect triggeringVSAvoidaccuracy of sound effect timing
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system uses audio data from the reader's actual reading to provide feedback on reading speed and position, dynamically adjusting sound effect triggering timing based on real-time detection of phoneme recognition and prosody analysis, rather than relying on predetermined playback speed estimates

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces simple playback speed estimation with an audio processing system that analyzes phoneme recognition, prosody characteristics, and reading position to determine sound effect triggering, substituting a mechanical timing approach with an intelligent audio analysis approach

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If keyword recognition is used to trigger sound effects, then specific sound effects can be triggered, but recognition errors and tracking errors increase

Engineering Contradiction:
Improveability to trigger specific sound effectsVSAvoidaccuracy of reading tracking
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system segments the text into phonemes and groups phonemes into words, using hierarchical recognition where phoneme-level accuracy feeds into word-level recognition, reducing overall error rates compared to direct keyword recognition

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system continuously monitors reading position and phoneme recognition confidence, providing feedback to adjust tracking accuracy and reduce errors in sound effect triggering based on actual reading progress

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If linear music with loops is used for sound effects, then the implementation is simple, but the reading experience lacks adaptability and personalization

Engineering Contradiction:
Improvesimplicity of sound effect implementationVSAvoidadaptability to reader emotions and style
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system transitions from static loop-based sound effects to dynamic sound effect selection that adapts in real-time based on detected reading emotions, prosody characteristics, and reading position, allowing the audio experience to evolve with the reader's emotional state

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes sound effect parameters based on detected prosody characteristics and reading emotions, selecting from multiple sound effect options and adjusting their properties to match the emotional context of the reading passage

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If phoneme-based position tracking is used, then reading position accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improveaccuracy of reading positionVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments phoneme recognition and position tracking into discrete, manageable units that can be processed sequentially, reducing computational complexity by breaking down the continuous audio stream into discrete phoneme events

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary phoneme recognition and position determination before sound effect triggering decisions are made, pre-processing the audio data to extract relevant position information that simplifies subsequent sound effect selection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4506924A1Method for receiving and processing audio data and triggering of associated sound effects according to prosody and/or movement
Publication Date: 2025.02.12 POÉTIE
  • EP4506924A1 patent drawingFigure 1~3
  • EP4506924A1 patent drawingFigure 4~5
  • EP4506924A1 patent drawing

AI summary

The invention relates to a method for receiving and processing audio data comprising speech corresponding to the real-time reading of a source text for triggering sound effects synchronized with said reading of the text, characterized in that it comprises a step (110) of determining an index (12) of the speaker's position in the source text, by detecting a correspondence between the received audio data and the source text, a step (112) of receiving at least one data representative of a prosody data (10), and/or of receiving at least one motion data, a step (140) of determining the sound effect to be triggered as a function of the position index (12), and as a function of the prosody data and/or the motion data, a step (150) of triggering the determined sound effect.