Audio Data Harmonic Structure Extraction Using Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting harmonic structure information from human speech audio data are prone to large errors, especially when compared to audio data generated by musical instruments, and cannot be jointly trained with neural network models.
Innovation Solution
A method involving a neural network model that processes spectral data to obtain fundamental frequency indication information, which is then used to extract harmonic structure information by combining it with global and harmonic energy distribution information, reducing errors through a cascaded detection network approach.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods are used to extract harmonic structure information from human speech audio data, then the extraction process is simple, but large errors occur in the extraction
Solution Approach 1:
The patent segments the audio data processing into distinct stages: spectral data processing to obtain first feature information, fundamental frequency indication information extraction, and harmonic structure information extraction. This segmentation allows each stage to be optimized independently, improving overall extraction accuracy while managing complexity systematically
Solution Approach 2:
The patent introduces fundamental frequency indication information as an intermediary element that bridges spectral data and harmonic structure information. This intermediary provides a reference framework that guides the extraction process, significantly improving accuracy by establishing a structured approach to identifying harmonic components
2Adaptability or versatility
If methods designed for musical instrument audio data are applied to human speech, then the extraction process works well for instruments, but large errors occur for human speech
Solution Approach 1:
The patent applies local quality by extracting fundamental frequency indication information specifically tailored for human speech characteristics from spectral data, rather than using a universal method for all audio types. This localized approach accounts for the unique properties of human speech signals, improving extraction accuracy for speech-specific applications
Solution Approach 2:
The patent changes the processing parameters and approach based on the audio data type. By adapting the spectral data processing and fundamental frequency extraction parameters specifically for human speech characteristics, the system achieves high accuracy for speech while maintaining the ability to handle different audio types
3Ease of operation
If harmonic structure information is extracted without fundamental frequency indication, then the process is straightforward, but the extraction contains large errors
Solution Approach 1:
The patent performs preliminary extraction of fundamental frequency indication information from spectral data before proceeding to harmonic structure information extraction. This preliminary action establishes a reference framework that guides subsequent extraction, improving accuracy by preparing the necessary foundation in advance
Solution Approach 2:
The patent uses fundamental frequency indication information as a feedback mechanism that guides the harmonic structure extraction process. The extracted fundamental frequency serves as a reference that continuously informs and adjusts the identification of harmonic components, improving overall extraction accuracy through this feedback loop
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
A method and an apparatus of processing audio data, an electronic device, a storage medium, and a program product are provided, which relates to a field of artificial intelligence, in particular to a field of speech processing technology. The method includes: processing spectral data of the audio data to obtain a first feature information; obtaining a fundamental frequency indication information according to the first feature information, wherein the fundamental frequency indication information indicates valid audio data of the first feature information and invalid audio data of the first feature information; obtaining a fundamental frequency information and a spectral energy information according to the first feature information and the fundamental frequency indication information; and obtaining a harmonic structure information of the audio data according to the fundamental frequency information and the spectral energy information.