Speech Signal Recognition Accuracy via Audio-Text Feature Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech interaction technologies face challenges in accurately distinguishing non-speech inputs like environmental noise, leading to false recognition, and existing acoustic confidence technologies either lack granularity or fail to effectively utilize both audio and text information for improved recognition accuracy.
Innovation Solution
A method that generates speech feature representations from a received speech signal, then uses these features to create source and target text feature representations, determining a match degree with predefined reference features to enhance recognition accuracy, thereby improving the judgment of speech signals and human-machine interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing acoustic confidence technologies are used, then speech recognition can be performed, but the accuracy is insufficient due to inability to effectively utilize both audio and text information
Solution Approach 1:
The patent merges audio feature representations and text feature representations into a unified processing framework. The neural network model simultaneously processes both modalities to generate confidence scores, combining the strengths of acoustic information and textual information to achieve more accurate speech recognition while avoiding the complexity of separate processing systems.
Solution Approach 2:
The neural network model serves multiple functions: it processes audio features, processes text features, and generates confidence scores for recognition results. This multi-functional approach allows the system to leverage both audio and text information for improved accuracy without requiring separate specialized systems for each function.
2Reliability
If speech signals are processed without feature representation generation, then processing is faster, but the ability to distinguish non-speech inputs like environmental noise is reduced
Solution Approach 1:
The system performs preliminary feature extraction and representation generation on both audio and text data before the final recognition and confidence scoring. By preparing feature representations in advance, the system can quickly and accurately distinguish speech from non-speech inputs without adding significant processing time to the overall workflow.
3Measurement precision
If text recognition is performed without comparing against reference features, then the process is simpler, but the accuracy of recognizing text from speech signals is insufficient
Solution Approach 1:
The system uses predefined reference text feature representations as a feedback mechanism to evaluate and improve text recognition accuracy. By comparing generated text features against these references, the system can identify and correct recognition errors, enhancing accuracy without requiring complex real-time verification systems.
Data Source
AI summary
According to embodiments of the disclosure, a method and an apparatus for processing a speech signal, and a computer-readable storage medium are provided. The method includes obtaining a set of speech feature representations of a speech signal received. The method also includes generating a set of source text feature representations based on a text recognized from the speech signal, each source text feature representation corresponding to an element in the text. The method also includes generating a set of target text feature representations based on the set of speech feature representations and the set of source text feature representations. The method also includes determining a match degree between the set of target text feature representations and a set of reference text feature representations predefined for the text, the match degree indicating an accuracy of recognizing of the text.


