Audio Recognition via Phase Channel Peak Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio recognition methods require manual input of basic information, which is often unavailable or incorrect, limiting their effectiveness in recognizing unknown audio documents, such as songs or music, especially in scenarios where users do not know the details of the audio being played.
Innovation Solution
A method and device for audio recognition that collects an audio document, performs time-frequency analysis to generate phase channels, extracts peak value characteristic points, and calculates audio fingerprints, allowing for automatic recognition by comparing these fingerprints with a pre-established database to identify matching audio documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual input of basic information is required for audio recognition, then the recognition process can be initiated, but the system becomes unusable when users do not know the basic information
Solution Approach 1:
The system automatically extracts basic information (artist name, song title, album) from the audio file itself using audio fingerprinting technology, eliminating the need for manual user input. The audio document serves itself by providing the information needed for recognition through its own content analysis.
2Measurement precision
If manual input of basic information is required, then some recognition can be performed, but recognition accuracy decreases when input information is incorrect
Solution Approach 1:
The patent replaces the mechanical process of manual information entry with an automated acoustic analysis system. The system uses audio fingerprinting and content analysis to automatically extract and verify basic information from the audio signal itself, substituting human manual input with automated signal processing.
3Device complexity
If conventional audio recognition methods are used with manual input, then the system structure remains simple, but the system cannot recognize audio documents when basic information is unavailable
Solution Approach 1:
The system performs preliminary extraction and storage of basic information from audio documents during the indexing phase, before recognition is needed. Audio fingerprints and metadata are pre-computed and stored in a database, enabling rapid automated recognition without requiring manual input at the time of use.
Data Source
AI summary
A method and device for performing audio recognition, including: collecting a first audio document to be recognized; initiating calculation of first characteristic information of the first audio document, including: conducting time-frequency analysis for the first audio document to generate a first preset number of phase channels; and extracting at least one peak value characteristic point from each phase channel of the first preset number of phrase channels, where the at least one peak value characteristic point of each phase channel constitutes the peak value characteristic point sequence of said each phase channel; and obtaining a recognition result for the first audio document, wherein the recognition result is identified based on the first characteristic information, and wherein the first characteristic information is calculated based on the respective peak value characteristic point sequences of the preset number of phase channels.


