Song Identification via Matched Audio Frame Units

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining the song corresponding to a video segment are inaccurate due to low extraction accuracy of video segment fragments and simple matching techniques, leading to a poor user experience as users need to manually identify songs across different applications.

Innovation Solution

A method and device that extracts an audio file from a video, acquires candidate song identifications, matches audio frames, and forms a matched audio frame unit to determine the target song identification, allowing for accurate display of supplemental content alongside the video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If simple matching manner is used for song matching, then the operation is simple and fast, but the accuracy for determining the song corresponding to the video segment is low

Engineering Contradiction:
Improveaccuracy for determining songVSAvoidcomplexity of matching process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio content into multiple audio frames and extracts key features from each frame. By dividing the audio signal into discrete segments and analyzing their characteristics separately, the system achieves more accurate song identification without requiring overly complex global analysis of the entire audio track.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the audio signal into different parameter representations, including frequency domain parameters and time-domain features. By changing the parameter space in which matching occurs, the system improves identification accuracy while maintaining computational efficiency through standardized feature extraction pipelines.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If low extraction accuracy of video segment fragment is used, then the extraction process is simple and fast, but the accuracy for determining the song is low

Engineering Contradiction:
Improveextraction accuracyVSAvoidtime for extraction and matching
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction and preprocessing of audio features from video segments before the actual matching process. By preparing audio frames, extracting features, and organizing data in advance, the system reduces the computational burden during matching, achieving both high extraction accuracy and reasonable processing speed.

Inventive Principle:
Principle #10Preliminary action

3Extent of automation

If manual identification of song is required, then the process can be simple without complex algorithms, but the user experience is poor and time consuming

Engineering Contradiction:
Improveautomation of song identificationVSAvoidcomplexity of identification system
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent implements an automated system that performs song identification without requiring user intervention. The system extracts audio features, compares them against a database, and automatically identifies and displays song information, enabling the system to serve itself in the identification process rather than relying on manual user input.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical identification processes with automated computational methods. Instead of requiring users to manually search and identify songs, the system uses algorithmic audio analysis and database matching to automatically determine song identity, substituting human effort with automated information processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10719551B2Song determining method and device and storage medium
Publication Date: 2020.07.21 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US10719551B2 patent drawing
  • US10719551B2 patent drawing
  • US10719551B2 patent drawing

AI summary

A song determining method and device are provided. According to the embodiment of the present disclosure, by extracting the audio file in the video and acquiring the candidate song identification of the candidate song, to which the segment belongs, in the audio file, the candidate song identification set is obtained; then by acquiring the candidate song file corresponding to the candidate song identification and acquiring a matched audio frame, in which the candidate song file is matched with the audio file, the matched audio frame unit is obtained, wherein the matched audio frame unit includes multiple continuous matched audio frames; the target song identification of the target song, to which the segment belongs, is acquired from the candidate song identification set according to the matched audio frame unit corresponding to the candidate song identification, and the target song, to which the segment belongs, is determined according to the target song identification.