Highlight Section Extraction from Audio Files

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing music retrieval systems fail to efficiently extract highlight sections that accurately represent music, leading to slow search speeds and low accuracy, as they cannot convey the emotional essence of music and require professional skills or significant time to find preferred music.

Innovation Solution

An apparatus and method that analyze audio files by dividing them into frames, calculating average energy signals, and selecting highlight sections based on low-frequency signals, with secondary and tertiary candidate group selection to determine the most representative frames, allowing for rapid identification of preferred music.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If highlight sections are extracted using existing methods, then music search convenience is improved, but extraction speed is slow and accuracy is low

Engineering Contradiction:
Improvemusic search convenienceVSAvoidextraction speed
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The audio signal is divided into multiple short-time frames, and each frame is analyzed independently to calculate energy characteristics. This segmentation allows parallel processing and reduces computational complexity, enabling faster extraction speed while maintaining accuracy in identifying highlight sections.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If highlight sections are extracted using existing methods, then music search convenience is improved, but extraction accuracy is low

Engineering Contradiction:
Improvemusic search convenienceVSAvoidextraction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Different weightings are assigned to different frequency bands when calculating energy characteristics. High-frequency components are emphasized for identifying rhythmic elements, while low-frequency components are weighted for identifying melodic and harmonic features. This local quality differentiation enables accurate identification of highlight sections that represent the emotional essence of music.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts energy calculation parameters based on the musical context and detected features. By changing the weighting factors and threshold values adaptively, the system achieves high extraction accuracy while maintaining fast processing speed, resolving the contradiction between convenience and precision.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If meta information is provided about music, then music information access is improved, but emotional feeling and climax cannot be conveyed

Engineering Contradiction:
Improvemusic information accessVSAvoidemotional feeling conveyance
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system extracts and isolates the most representative highlight sections from the complete audio file, capturing the emotional essence and climax moments. These extracted segments are then presented to users, providing both rapid access to music information and authentic emotional experience, simultaneously addressing both requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9262521B2Apparatus and method for extracting highlight section of music
Publication Date: 2016.02.16 ELECTRONICS & TELECOMM RES INST
  • US9262521B2 patent drawing
  • US9262521B2 patent drawing
  • US9262521B2 patent drawing

AI summary

Disclosed are an apparatus and a method for extracting a highlight section of music. The apparatus for extracting a highlight section of music in accordance with the embodiment of the present invention includes a frame divider that divides an audio file into a plurality of frames having a predetermined sample length; an average energy signal calculator that calculates a signal representing the average magnitude of audio energy for a plurality of samples belonging to each frame for each frame of the plurality of frames; and a highlight section selector that extracts a low frequency signal from the signal representing the average audio energy magnitude for each frame and determines the highlight section from frame sections including maximum points of the low frequency signal.