Audio Pronunciation Correction via Segmented Prediction Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing methods struggle with low accuracy in predicting pronunciation sequences for text, particularly when dealing with neutral tones and consecutive third tones.

Innovation Solution

The proposed solution involves an audio processing method and apparatus that utilize multiple pronunciation prediction systems to obtain and correct pronunciation sequences. This includes obtaining audio and text, predicting an initial pronunciation sequence, and then correcting neutral tones and third tones after tone sandhi using separate correction systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single pronunciation prediction system is used, then the system complexity is low, but the accuracy of predicting pronunciation sequences for neutral tones and consecutive third tones is low

Engineering Contradiction:
Improvepronunciation sequence prediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the pronunciation prediction task into multiple specialized systems: a first pronunciation prediction system for general pronunciation, a second pronunciation prediction system specifically for neutral tones, and a third pronunciation prediction system specifically for third tones after tone sandhi. Each system focuses on predicting specific tone types, allowing them to be optimized independently for their respective functions, thereby improving overall accuracy without requiring each individual system to handle all complexity alone.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a multi-functional architecture where the first pronunciation prediction system serves as a general-purpose system that can handle various pronunciation prediction tasks, while the second and third systems are specialized subsystems that can be activated when specific tone types are detected. This universal approach allows the system to leverage the strengths of each specialized system while maintaining a cohesive overall structure that handles diverse pronunciation scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple pronunciation prediction systems are used to improve accuracy, then the prediction accuracy for neutral tones and third tones is improved, but the system complexity increases

Engineering Contradiction:
Improvepronunciation sequence prediction accuracyVSAvoidnumber of prediction systems
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the pronunciation prediction functionality into distinct systems, each responsible for specific tone types. The first system handles general pronunciation prediction, the second system specifically predicts neutral tones, and the third system predicts third tones after tone sandhi. This segmentation allows each system to be trained and optimized for its specific function, improving overall accuracy while keeping each individual system relatively simple and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism that coordinates between the multiple prediction systems. When the system detects neutral tones or third tones after tone sandhi in the input text, it activates the appropriate specialized system (second or third system) to provide corrected predictions. This intermediary coordination layer manages the complexity of having multiple systems by providing clear rules for when each system should be invoked, ensuring that the system complexity remains organized and controllable.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250124916A1Audio processing method and apparatus, electronic device, and storage medium
Publication Date: 2025.04.17 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250124916A1 patent drawing
  • US20250124916A1 patent drawing
  • US20250124916A1 patent drawing

AI summary

Embodiments of the present disclosure provide an audio processing method and apparatus, an electronic device, and a storage medium. The method includes: obtaining first audio and first text corresponding to the first audio; predicting a first pronunciation sequence for the first text by a first pronunciation prediction system based on the first audio and the first text, where tones of pronunciations of characters in the first text that are labeled in the first pronunciation sequence include neutral tones and/or third tones after tone sandhi; and the first third tone in two consecutive third tones in the first text is labeled as a third tone after tone sandhi in the first pronunciation sequence; and correcting a neutral tone in the first pronunciation sequence by a second pronunciation prediction system, and/or correcting a third tone after tone sandhi in the first pronunciation sequence by a third pronunciation prediction system.