Speech Signal Recognition Accuracy via Audio-Text Feature Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech interaction technologies face challenges in accurately distinguishing non-speech inputs like environmental noise, leading to false recognition, and existing acoustic confidence technologies either lack granularity or fail to effectively utilize both audio and text information for improved recognition accuracy.

Innovation Solution

A method that generates speech feature representations from a received speech signal, then uses these features to create source and target text feature representations, determining a match degree with predefined reference features to enhance recognition accuracy, thereby improving the judgment of speech signals and human-machine interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing acoustic confidence technologies are used, then speech recognition can be performed, but the accuracy is insufficient due to inability to effectively utilize both audio and text information

Engineering Contradiction:
Improverecognition accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges audio feature representations and text feature representations into a unified processing framework. The neural network model simultaneously processes both modalities to generate confidence scores, combining the strengths of acoustic information and textual information to achieve more accurate speech recognition while avoiding the complexity of separate processing systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network model serves multiple functions: it processes audio features, processes text features, and generates confidence scores for recognition results. This multi-functional approach allows the system to leverage both audio and text information for improved accuracy without requiring separate specialized systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If speech signals are processed without feature representation generation, then processing is faster, but the ability to distinguish non-speech inputs like environmental noise is reduced

Engineering Contradiction:
Improveability to distinguish speech from non-speechVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary feature extraction and representation generation on both audio and text data before the final recognition and confidence scoring. By preparing feature representations in advance, the system can quickly and accurately distinguish speech from non-speech inputs without adding significant processing time to the overall workflow.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If text recognition is performed without comparing against reference features, then the process is simpler, but the accuracy of recognizing text from speech signals is insufficient

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidfeature comparison and matching complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses predefined reference text feature representations as a feedback mechanism to evaluate and improve text recognition accuracy. By comparing generated text features against these references, the system can identify and correct recognition errors, enhancing accuracy without requiring complex real-time verification systems.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11322151B2Method, apparatus, and medium for processing speech signal
Publication Date: 2022.05.03 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11322151B2 patent drawing
  • US11322151B2 patent drawing
  • US11322151B2 patent drawing

AI summary

According to embodiments of the disclosure, a method and an apparatus for processing a speech signal, and a computer-readable storage medium are provided. The method includes obtaining a set of speech feature representations of a speech signal received. The method also includes generating a set of source text feature representations based on a text recognized from the speech signal, each source text feature representation corresponding to an element in the text. The method also includes generating a set of target text feature representations based on the set of speech feature representations and the set of source text feature representations. The method also includes determining a match degree between the set of target text feature representations and a set of reference text feature representations predefined for the text, the match degree indicating an accuracy of recognizing of the text.