Audio Recognition Using Voting Matrix and Characteristic Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio recognition technologies face low matching success rates and recognition accuracy due to poor noise immunity, as extremum points are not stable and can be interfered with by noise.
Innovation Solution
A method for audio recognition that divides audio data into frames, calculates characteristic values based on audio variation trends, and matches these values with a pre-established comparison table, using a voting matrix to determine recognition results, effectively resisting noise interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If extremum points are extracted from sonogram for audio recognition, then recognition process is simplified and efficiency is improved, but noise immunity deteriorates and matching success rate decreases
Solution Approach 1:
The audio signal is divided into multiple frames, and each frame is processed independently to extract characteristic values. This segmentation allows the system to analyze local audio characteristics while maintaining overall recognition accuracy, resolving the contradiction between processing efficiency and recognition reliability.
Solution Approach 2:
The patent transforms the audio recognition approach from extremum point extraction to characteristic value calculation based on audio variation trends. By changing the parameter representation from fixed extremum points to dynamic characteristic values derived from adjacent frames, the system achieves both efficiency and improved noise immunity.
2Device complexity
If extremum points are used for audio recognition, then processing complexity is reduced, but noise immunity worsens and recognition accuracy decreases
Solution Approach 1:
The patent changes the parameter representation from extremum points to characteristic values that capture audio variation trends. This parameter transformation maintains processing simplicity while significantly improving recognition accuracy by representing audio characteristics in a noise-resilient manner.
Solution Approach 2:
The patent introduces characteristic values as an intermediary representation between raw audio data and recognition results. These characteristic values serve as a mediator that simplifies processing while improving accuracy by filtering out noise through trend-based calculation.
3Ease of operation
If extremum points are extracted for matching, then recognition process is simplified, but stability of characteristic points deteriorates under noise interference
Solution Approach 1:
The patent transforms the characteristic representation from fixed extremum points to dynamic characteristic values based on audio variation trends. This parameter change maintains operational simplicity while achieving stability under noise by using relative changes rather than absolute positions.
Solution Approach 2:
The patent introduces dynamic characteristic values that adapt to changing audio conditions rather than relying on static extremum points. This dynamic approach allows the recognition system to maintain simplicity while achieving stability by continuously adapting to noise conditions through trend-based calculation.
Data Source
AI summary
A method may include dividing input audio into frames and calculating a characteristic value for each of the frames. The method may include establishing a voting matrix having a first dimension representing a quantity of segments of sample audio and a second dimension representing a quantity of frames of each segment. The method may include marking voting labels in the voting matrix corresponding to frames of the sample audio when the characteristic values of corresponding frames of the input audio and sample audio match. The method may include determining a frame to be a recognition result when a sum of the voting labels at a corresponding position is higher than a threshold.


