Speech Processing Abnormal Segment Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current instant message applications face inefficiencies in speech recording due to abnormal situations like stammering or pauses, requiring users to re-record and re-send speech, which disrupts communication flow.
Innovation Solution
A method and apparatus for processing speech that involves speech recognition to identify and process abnormal segments, such as blank or elongated tone segments, by converting them into preset symbols, deleting or revising these segments, and smoothing the final speech to generate a coherent and natural output, thereby avoiding re-recording.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If users directly send recorded speech without processing, then the speech recording process is simple, but abnormal segments (stammering, pauses) require re-recording which reduces productivity
Solution Approach 1:
The speech is divided into multiple segments based on abnormal detection (stammering, pauses, elongated tones). Each segment is independently identified and processed, allowing targeted removal or replacement of only the problematic portions rather than requiring complete re-recording, thus improving productivity while maintaining manageable processing complexity
Solution Approach 2:
The system performs preliminary speech recognition and abnormal segment detection on the original speech before final processing. By identifying problematic segments in advance through speech-to-text conversion and analysis, the system prepares for efficient processing by knowing exactly which portions need modification, reducing the need for re-recording and improving overall efficiency
2Reliability
If users re-record speech when abnormalities occur, then speech quality can be maintained, but communication flow is disrupted and time is lost
Solution Approach 1:
The system extracts and removes only the abnormal segments (stammering, pauses, elongated tones) from the original speech while preserving the normal segments. This selective extraction maintains the overall speech quality and natural flow without requiring users to re-record the entire message, thereby reducing time loss while maintaining reliability
Solution Approach 2:
The system creates a processed version of the speech by copying and concatenating the normal segments after removing abnormal portions. This copying approach preserves the original speech quality and intent while eliminating problematic parts, avoiding the need for complete re-recording and reducing time loss
3Stability of the object's composition
If abnormal segments are removed from speech, then coherence is improved, but synchronization between speech and text becomes more complex
Solution Approach 1:
Both the speech and its corresponding text are segmented together based on abnormal segment detection. When an abnormal segment is identified in the speech, the corresponding text segment is identified and removed simultaneously. This joint segmentation approach maintains coherence by ensuring speech and text remain synchronized while managing complexity through systematic pairing of speech-text segments
4Measurement precision
If speech recognition is performed on original speech, then abnormal segments can be identified, but processing time increases
Solution Approach 1:
The system performs speech recognition selectively on segments that may contain abnormalities rather than processing the entire speech uniformly. By applying speech recognition only where needed to detect stammering, pauses, and elongated tones, the system achieves adequate detection accuracy while reducing overall processing time compared to full-speech recognition
Data Source
AI summary
Embodiments of the present disclosure provide a method and apparatus for processing a speech. The method may include: acquiring an original speech; performing speech recognition on the original speech, to obtain an original text corresponding to the original speech; associating a speech segment in the original speech with a text segment in the original text; recognizing an abnormal segment in the original speech and/or the original text; and processing a text segment indicated by the abnormal segment in the original text and/or the speech segment indicated by the abnormal segment in the original speech, to generate a final speech. A speech segment in the original speech is associated with a text segment in the original text to realize visual processing of the speech.


