Audio Pronunciation Correction via Segmented Prediction Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing methods struggle with low accuracy in predicting pronunciation sequences for text, particularly when dealing with neutral tones and consecutive third tones.
Innovation Solution
The proposed solution involves an audio processing method and apparatus that utilize multiple pronunciation prediction systems to obtain and correct pronunciation sequences. This includes obtaining audio and text, predicting an initial pronunciation sequence, and then correcting neutral tones and third tones after tone sandhi using separate correction systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single pronunciation prediction system is used, then the system complexity is low, but the accuracy of predicting pronunciation sequences for neutral tones and consecutive third tones is low
Solution Approach 1:
The patent divides the pronunciation prediction task into multiple specialized systems: a first pronunciation prediction system for general pronunciation, a second pronunciation prediction system specifically for neutral tones, and a third pronunciation prediction system specifically for third tones after tone sandhi. Each system focuses on predicting specific tone types, allowing them to be optimized independently for their respective functions, thereby improving overall accuracy without requiring each individual system to handle all complexity alone.
Solution Approach 2:
The patent creates a multi-functional architecture where the first pronunciation prediction system serves as a general-purpose system that can handle various pronunciation prediction tasks, while the second and third systems are specialized subsystems that can be activated when specific tone types are detected. This universal approach allows the system to leverage the strengths of each specialized system while maintaining a cohesive overall structure that handles diverse pronunciation scenarios.
2Measurement precision
If multiple pronunciation prediction systems are used to improve accuracy, then the prediction accuracy for neutral tones and third tones is improved, but the system complexity increases
Solution Approach 1:
The patent segments the pronunciation prediction functionality into distinct systems, each responsible for specific tone types. The first system handles general pronunciation prediction, the second system specifically predicts neutral tones, and the third system predicts third tones after tone sandhi. This segmentation allows each system to be trained and optimized for its specific function, improving overall accuracy while keeping each individual system relatively simple and manageable.
Solution Approach 2:
The patent introduces an intermediary mechanism that coordinates between the multiple prediction systems. When the system detects neutral tones or third tones after tone sandhi in the input text, it activates the appropriate specialized system (second or third system) to provide corrected predictions. This intermediary coordination layer manages the complexity of having multiple systems by providing clear rules for when each system should be invoked, ensuring that the system complexity remains organized and controllable.
Data Source
AI summary
Embodiments of the present disclosure provide an audio processing method and apparatus, an electronic device, and a storage medium. The method includes: obtaining first audio and first text corresponding to the first audio; predicting a first pronunciation sequence for the first text by a first pronunciation prediction system based on the first audio and the first text, where tones of pronunciations of characters in the first text that are labeled in the first pronunciation sequence include neutral tones and/or third tones after tone sandhi; and the first third tone in two consecutive third tones in the first text is labeled as a third tone after tone sandhi in the first pronunciation sequence; and correcting a neutral tone in the first pronunciation sequence by a second pronunciation prediction system, and/or correcting a third tone after tone sandhi in the first pronunciation sequence by a third pronunciation prediction system.


