Audio Processing System for Synchronized Sound Effects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for triggering sound effects during reading are prone to errors, fail to adapt to the reader's emotions and reading style, and do not account for physical movements, leading to a lack of immersion and personalization in the reading experience.
Innovation Solution
A process that synchronizes sound effects with real-time reading by determining a speaker's position index and considering prosody and movement data, using preconfigured algorithms to adapt sound effects dynamically, allowing for a more immersive and personalized experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If playback speed estimation is used to trigger sound effects, then the triggering is simple, but the accuracy of sound effect timing is poor
Solution Approach 1:
The system uses audio data from the reader's actual reading to provide feedback on reading speed and position, dynamically adjusting sound effect triggering timing based on real-time detection of phoneme recognition and prosody analysis, rather than relying on predetermined playback speed estimates
Solution Approach 2:
The patent replaces simple playback speed estimation with an audio processing system that analyzes phoneme recognition, prosody characteristics, and reading position to determine sound effect triggering, substituting a mechanical timing approach with an intelligent audio analysis approach
2Adaptability or versatility
If keyword recognition is used to trigger sound effects, then specific sound effects can be triggered, but recognition errors and tracking errors increase
Solution Approach 1:
The system segments the text into phonemes and groups phonemes into words, using hierarchical recognition where phoneme-level accuracy feeds into word-level recognition, reducing overall error rates compared to direct keyword recognition
Solution Approach 2:
The system continuously monitors reading position and phoneme recognition confidence, providing feedback to adjust tracking accuracy and reduce errors in sound effect triggering based on actual reading progress
3Ease of manufacture
If linear music with loops is used for sound effects, then the implementation is simple, but the reading experience lacks adaptability and personalization
Solution Approach 1:
The system transitions from static loop-based sound effects to dynamic sound effect selection that adapts in real-time based on detected reading emotions, prosody characteristics, and reading position, allowing the audio experience to evolve with the reader's emotional state
Solution Approach 2:
The system changes sound effect parameters based on detected prosody characteristics and reading emotions, selecting from multiple sound effect options and adjusting their properties to match the emotional context of the reading passage
4Measurement precision
If phoneme-based position tracking is used, then reading position accuracy is improved, but computational complexity increases
Solution Approach 1:
The system segments phoneme recognition and position tracking into discrete, manageable units that can be processed sequentially, reducing computational complexity by breaking down the continuous audio stream into discrete phoneme events
Solution Approach 2:
The system performs preliminary phoneme recognition and position determination before sound effect triggering decisions are made, pre-processing the audio data to extract relevant position information that simplifies subsequent sound effect selection
Data Source
Figure 1~3
Figure 4~5
AI summary
The invention relates to a method for receiving and processing audio data comprising speech corresponding to the real-time reading of a source text for triggering sound effects synchronized with said reading of the text, characterized in that it comprises a step (110) of determining an index (12) of the speaker's position in the source text, by detecting a correspondence between the received audio data and the source text, a step (112) of receiving at least one data representative of a prosody data (10), and/or of receiving at least one motion data, a step (140) of determining the sound effect to be triggered as a function of the position index (12), and as a function of the prosody data and/or the motion data, a step (150) of triggering the determined sound effect.