Audio Data Harmonic Structure Extraction Using Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting harmonic structure information from human speech audio data are prone to large errors, especially when compared to audio data generated by musical instruments, and cannot be jointly trained with neural network models.

Innovation Solution

A method involving a neural network model that processes spectral data to obtain fundamental frequency indication information, which is then used to extract harmonic structure information by combining it with global and harmonic energy distribution information, reducing errors through a cascaded detection network approach.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods are used to extract harmonic structure information from human speech audio data, then the extraction process is simple, but large errors occur in the extraction

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio data processing into distinct stages: spectral data processing to obtain first feature information, fundamental frequency indication information extraction, and harmonic structure information extraction. This segmentation allows each stage to be optimized independently, improving overall extraction accuracy while managing complexity systematically

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces fundamental frequency indication information as an intermediary element that bridges spectral data and harmonic structure information. This intermediary provides a reference framework that guides the extraction process, significantly improving accuracy by establishing a structured approach to identifying harmonic components

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If methods designed for musical instrument audio data are applied to human speech, then the extraction process works well for instruments, but large errors occur for human speech

Engineering Contradiction:
Improvemethod adaptabilityVSAvoidextraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by extracting fundamental frequency indication information specifically tailored for human speech characteristics from spectral data, rather than using a universal method for all audio types. This localized approach accounts for the unique properties of human speech signals, improving extraction accuracy for speech-specific applications

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the processing parameters and approach based on the audio data type. By adapting the spectral data processing and fundamental frequency extraction parameters specifically for human speech characteristics, the system achieves high accuracy for speech while maintaining the ability to handle different audio types

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If harmonic structure information is extracted without fundamental frequency indication, then the process is straightforward, but the extraction contains large errors

Engineering Contradiction:
Improveprocess simplicityVSAvoidextraction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent performs preliminary extraction of fundamental frequency indication information from spectral data before proceeding to harmonic structure information extraction. This preliminary action establishes a reference framework that guides subsequent extraction, improving accuracy by preparing the necessary foundation in advance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses fundamental frequency indication information as a feedback mechanism that guides the harmonic structure extraction process. The extracted fundamental frequency serves as a reference that continuously informs and adjusts the identification of harmonic components, improving overall extraction accuracy through this feedback loop

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4120265B1Method and apparatus of processing audio data, electronic device, storage medium and program product
Publication Date: 2024.09.18 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP4120265B1 patent drawingFigure 1
  • EP4120265B1 patent drawingFigure 2A
  • EP4120265B1 patent drawingFigure 2B

AI summary

A method and an apparatus of processing audio data, an electronic device, a storage medium, and a program product are provided, which relates to a field of artificial intelligence, in particular to a field of speech processing technology. The method includes: processing spectral data of the audio data to obtain a first feature information; obtaining a fundamental frequency indication information according to the first feature information, wherein the fundamental frequency indication information indicates valid audio data of the first feature information and invalid audio data of the first feature information; obtaining a fundamental frequency information and a spectral energy information according to the first feature information and the fundamental frequency indication information; and obtaining a harmonic structure information of the audio data according to the fundamental frequency information and the spectral energy information.