Speech Processing Abnormal Segment Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current instant message applications face inefficiencies in speech recording due to abnormal situations like stammering or pauses, requiring users to re-record and re-send speech, which disrupts communication flow.

Innovation Solution

A method and apparatus for processing speech that involves speech recognition to identify and process abnormal segments, such as blank or elongated tone segments, by converting them into preset symbols, deleting or revising these segments, and smoothing the final speech to generate a coherent and natural output, thereby avoiding re-recording.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users directly send recorded speech without processing, then the speech recording process is simple, but abnormal segments (stammering, pauses) require re-recording which reduces productivity

Engineering Contradiction:
Improvespeech recording efficiencyVSAvoidspeech processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The speech is divided into multiple segments based on abnormal detection (stammering, pauses, elongated tones). Each segment is independently identified and processed, allowing targeted removal or replacement of only the problematic portions rather than requiring complete re-recording, thus improving productivity while maintaining manageable processing complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary speech recognition and abnormal segment detection on the original speech before final processing. By identifying problematic segments in advance through speech-to-text conversion and analysis, the system prepares for efficient processing by knowing exactly which portions need modification, reducing the need for re-recording and improving overall efficiency

Inventive Principle:
Principle #10Preliminary action

2Reliability

If users re-record speech when abnormalities occur, then speech quality can be maintained, but communication flow is disrupted and time is lost

Engineering Contradiction:
Improvespeech qualityVSAvoidtime for re-recording
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts and removes only the abnormal segments (stammering, pauses, elongated tones) from the original speech while preserving the normal segments. This selective extraction maintains the overall speech quality and natural flow without requiring users to re-record the entire message, thereby reducing time loss while maintaining reliability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates a processed version of the speech by copying and concatenating the normal segments after removing abnormal portions. This copying approach preserves the original speech quality and intent while eliminating problematic parts, avoiding the need for complete re-recording and reducing time loss

Inventive Principle:
Principle #26Copying

3Stability of the object's composition

If abnormal segments are removed from speech, then coherence is improved, but synchronization between speech and text becomes more complex

Engineering Contradiction:
Improvespeech coherenceVSAvoidspeech-text synchronization complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

Both the speech and its corresponding text are segmented together based on abnormal segment detection. When an abnormal segment is identified in the speech, the corresponding text segment is identified and removed simultaneously. This joint segmentation approach maintains coherence by ensuring speech and text remain synchronized while managing complexity through systematic pairing of speech-text segments

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If speech recognition is performed on original speech, then abnormal segments can be identified, but processing time increases

Engineering Contradiction:
Improveabnormal segment detection accuracyVSAvoidspeech recognition processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs speech recognition selectively on segments that may contain abnormalities rather than processing the entire speech uniformly. By applying speech recognition only where needed to detect stammering, pauses, and elongated tones, the system achieves adequate detection accuracy while reducing overall processing time compared to full-speech recognition

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11488603B2Method and apparatus for processing speech
Publication Date: 2022.11.01 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11488603B2 patent drawing
  • US11488603B2 patent drawing
  • US11488603B2 patent drawing

AI summary

Embodiments of the present disclosure provide a method and apparatus for processing a speech. The method may include: acquiring an original speech; performing speech recognition on the original speech, to obtain an original text corresponding to the original speech; associating a speech segment in the original speech with a text segment in the original text; recognizing an abnormal segment in the original speech and/or the original text; and processing a text segment indicated by the abnormal segment in the original text and/or the speech segment indicated by the abnormal segment in the original speech, to generate a final speech. A speech segment in the original speech is associated with a text segment in the original text to realize visual processing of the speech.