Audio Recognition via Phase Channel Peak Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio recognition methods require manual input of basic information, which is often unavailable or incorrect, limiting their effectiveness in recognizing unknown audio documents, such as songs or music, especially in scenarios where users do not know the details of the audio being played.

Innovation Solution

A method and device for audio recognition that collects an audio document, performs time-frequency analysis to generate phase channels, extracts peak value characteristic points, and calculates audio fingerprints, allowing for automatic recognition by comparing these fingerprints with a pre-established database to identify matching audio documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual input of basic information is required for audio recognition, then the recognition process can be initiated, but the system becomes unusable when users do not know the basic information

Engineering Contradiction:
Improveease of useVSAvoidapplicability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system automatically extracts basic information (artist name, song title, album) from the audio file itself using audio fingerprinting technology, eliminating the need for manual user input. The audio document serves itself by providing the information needed for recognition through its own content analysis.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual input of basic information is required, then some recognition can be performed, but recognition accuracy decreases when input information is incorrect

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser effort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces the mechanical process of manual information entry with an automated acoustic analysis system. The system uses audio fingerprinting and content analysis to automatically extract and verify basic information from the audio signal itself, substituting human manual input with automated signal processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If conventional audio recognition methods are used with manual input, then the system structure remains simple, but the system cannot recognize audio documents when basic information is unavailable

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary extraction and storage of basic information from audio documents during the indexing phase, before recognition is needed. Audio fingerprints and metadata are pre-computed and stored in a database, enabling rapid automated recognition without requiring manual input at the time of use.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9373336B2Method and device for audio recognition
Publication Date: 2016.06.21 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US9373336B2 patent drawing
  • US9373336B2 patent drawing
  • US9373336B2 patent drawing

AI summary

A method and device for performing audio recognition, including: collecting a first audio document to be recognized; initiating calculation of first characteristic information of the first audio document, including: conducting time-frequency analysis for the first audio document to generate a first preset number of phase channels; and extracting at least one peak value characteristic point from each phase channel of the first preset number of phrase channels, where the at least one peak value characteristic point of each phase channel constitutes the peak value characteristic point sequence of said each phase channel; and obtaining a recognition result for the first audio document, wherein the recognition result is identified based on the first characteristic information, and wherein the first characteristic information is calculated based on the respective peak value characteristic point sequences of the preset number of phase channels.