Audio Playback Repositioning Using Sentence Segmentation Tags
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio playback technologies face issues with inaccurate positioning, high operational cost, and low efficiency when re-listening to specific audio segments, particularly when the speech rate is slow or variable, leading to cumbersome and inefficient user operations.
Innovation Solution
An audio playing method that recognizes audio files as text files with sentence segmentation symbols, generates corresponding tags, and determines a target play point based on user triggers and sentence segmentation tags, allowing precise repositioning without repetitive user interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio playback positioning methods are used, then the system is simple to implement, but the positioning accuracy is poor and operational efficiency is low
Solution Approach 1:
The patent introduces sentence segmentation tags as an intermediary layer between the audio file and the playback control system. These tags mark specific positions in the audio corresponding to sentence boundaries, enabling precise positioning without requiring complex manual intervention. The tags act as mediators that translate semantic structure into playable positions.
Solution Approach 2:
The system performs preliminary sentence segmentation and tag generation before audio playback. By pre-processing the audio file to identify and mark sentence boundaries with tags, the system prepares positioning information in advance, eliminating the need for complex real-time analysis during playback and improving both accuracy and efficiency.
2Productivity
If manual positioning methods are used, then the operation cost is high and user operations are cumbersome, but the system requires minimal processing
Solution Approach 1:
The system automatically performs sentence segmentation, tag generation, and positioning without requiring user intervention. The audio processing system serves itself by identifying sentence boundaries and creating playable positions autonomously, eliminating cumbersome manual operations and significantly improving operational efficiency.
Solution Approach 2:
Sentence segmentation tags serve as intermediaries that automate the positioning process. Instead of users manually searching for positions, the tags provide ready-made anchors that the playback system can automatically jump to, reducing operational complexity while maintaining ease of use.
3Adaptability or versatility
If fixed return time is used, then the playback control is simple, but it does not meet the user's required timing and flexibility is low
Solution Approach 1:
The system transitions from fixed return time to dynamic positioning based on sentence segmentation tags. The playback control can now adapt to varying sentence lengths and structures by jumping to tag-marked positions rather than using fixed time intervals, providing flexibility while keeping the control mechanism simple through automatic tag-based navigation.
4Measurement precision
If speech rate variation is not considered, then the processing is straightforward, but the positioning accuracy deteriorates with varying speech rates
Solution Approach 1:
The system performs preliminary sentence segmentation and tag generation that accounts for speech rate variations. By analyzing and marking sentence boundaries before playback, the system creates positioning information that is independent of playback speed, ensuring accurate positioning regardless of how fast or slow the audio is played back.
Data Source
AI summary
The present application provides an audio playing method, an electronic device, and a computer readable storage medium. The method comprises: recognizing an audio file to be played as a text file containing sentence segmentation symbols; generating respective sentence segmentation tags at positions corresponding to the sentence segmentation symbols in the audio file, according to a correspondence relationship between the audio file and the text file; in response to a trigger operation, determining a target play point according to a current play position of the audio file and respective positions of the sentence segmentation tags; and playing the audio file from the target play point.


