Audio Digital Signal Processing Method and Device Based on SOC Chip

By implementing audio digital signal processing methods such as multi-channel sensor acquisition, multi-resolution spectrum decomposition, complex domain transformation, subband energy analysis and reverse Fourier reconstruction on the SOC chip, the audio data processing efficiency and accuracy problems are solved, and the processing and enhancement of high-quality audio signals are achieved.

CN119785805BActive Publication Date: 2025-06-13RAYLEIGH LABS TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510273602.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

With the increase in the dimensions of audio data, how to efficiently process these data without losing key information has become an urgent problem to be solved. Traditional methods face the problems of insufficient resolution, high computational complexity, and inaccurate spectrum feature extraction.

Method used

The target site is collected through a preset multi-channel sensor, and multi-resolution spectrum decomposition is performed to obtain the hierarchical spectrum characteristics, and the complex domain transformation is performed. The frequency domain reconstruction is carried out through subband energy analysis technology to obtain the acoustic parameter sequence. Finally, the reverse Fourier reconstruction is performed based on the segmented fast Fourier algorithm in the SOC chip to obtain the enhanced audio digital signal.

Benefits of technology

It realizes efficient processing of audio data, improves processing accuracy and efficiency, reduces noise interference, enhances useful audio signals, and ensures the quality and clarity of the output signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785805B_ABST
    Figure CN119785805B_ABST
Patent Text Reader

Abstract

The present invention relates to an audio digital signal processing method and device based on an SOC chip, comprising the following steps: collecting audio digital signals of a target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream; performing multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features; performing a complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data; performing frequency domain reconstruction on the complex spectrum representation data through a sub-band energy analysis technique to obtain an acoustic parameter sequence; performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence; and performing inverse Fourier reconstruction on the calibrated spectrum sequence to obtain an enhanced audio digital signal, thereby solving the technical problem of how to efficiently process such data without losing key information as the dimension of audio data increases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of SOC chips, and particularly to an audio digital signal processing method and device based on an SOC chip. Background Art

[0002] In the development process of modern audio processing technology, the demand for high-quality audio signals is increasing day by day, especially in the fields of professional audio recording, speech recognition, and communication. The audio digital signal processing method based on SOC (System on Chip) chips has emerged, aiming to improve the processing efficiency and reduce power consumption by integrating multiple processing functions on a single chip. However, traditional audio processing methods often face problems such as insufficient resolution, high computational complexity, and inaccurate spectral feature extraction, which limit the further improvement of audio signal processing quality.

[0003] Specifically, in the audio acquisition stage, how to effectively separate useful sound information from complex environmental noise is a major challenge. Traditional methods mostly rely on single or a few fixed parameters for spectral analysis, which makes it difficult to capture the subtle differences of audio signals when facing a changing acoustic environment. In addition, with the increase in the dimension of audio data, how to efficiently process these data without losing key information has become an urgent problem to be solved. The existence of these problems not only affects the clarity and accuracy of audio signals, but also restricts the innovation and development of related application fields to a certain extent.

[0004] To solve the above problems, researchers have been continuously exploring new audio signal processing technologies and algorithms. For example, the method of combining hierarchical spectral feature decomposition and complex domain transformation can more accurately represent the characteristics of audio signals and effectively improve the accuracy and efficiency of processing. Nevertheless, the existing technologies still face challenges such as large consumption of computing resources and limited real-time processing capabilities. Therefore, it is particularly important to develop an audio digital signal processing method based on SOC chips that can not only meet the requirements of high-quality audio processing, but also achieve low energy consumption and high efficiency. This research direction is of great significance for promoting the development of audio processing technology. Summary of the Invention

[0005] The main object of the present invention is to provide an audio digital signal processing method and device based on an SOC chip, which solves the technical problem of how to efficiently process audio data without losing key information as the dimension of audio data increases.

[0006] To achieve the above object, the present invention provides an audio digital signal processing method based on an SOC chip, including the following steps:

[0007] Collect audio digital signals from the target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream;

[0008] Perform multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features;

[0009] Perform a complex-domain transformation on the hierarchical spectral features to obtain complex spectral representation data;

[0010] Perform frequency-domain reconstruction on the complex spectral representation data through sub-band energy analysis technology to obtain an acoustic parameter sequence;

[0011] Perform group delay compensation on the acoustic parameter sequence to obtain a calibrated spectral sequence;

[0012] Based on the segmented fast Fourier transform algorithm in a preset SOC chip, perform inverse Fourier reconstruction on the calibrated spectral sequence to obtain an enhanced audio digital signal.

[0013] Further, the collecting audio digital signals from the target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream includes:

[0014] Perform spatial sampling on the target site through multi-channel sensors in a microphone array to obtain multi-channel sound field data, and perform sound source localization calculation based on the multi-channel sound field data to obtain sound source spatial distribution data; wherein, the sound source spatial distribution data includes a sound source azimuth angle, a sound source distance parameter, and a sound source intensity distribution;

[0015] Extract acoustic feature vectors from the sound source spatial distribution data, and perform sound field reconstruction based on the acoustic feature vectors to obtain a three-dimensional sound field data stream;

[0016] Perform modal analysis on the three-dimensional sound field data stream through sound field decomposition to obtain sound field modal coefficients, and perform sound field synthesis based on the sound field modal coefficients to obtain a multi-dimensional audio data stream; wherein, the multi-dimensional audio data stream includes sound field time-varying features, sound field spatial features, and sound field spectral features.

[0017] Further, the performing multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features includes:

[0018] Perform frequency band division on the multi-dimensional audio data stream through wavelet packet decomposition technology to obtain multi-level sub-band coefficients, and perform energy density calculation based on the multi-level sub-band coefficients to obtain frequency band energy distribution data; wherein, the frequency band energy distribution data includes critical band coefficients and a frequency band energy matrix;

[0019] Extract the spectral envelope of the band energy distribution data to obtain an envelope feature sequence, and perform frequency-domain modulation analysis based on the envelope feature sequence to obtain a modulation spectrum matrix;

[0020] Perform harmonic structure analysis on the multi-dimensional audio data stream based on the modulation spectrum matrix to obtain a harmonic component matrix, and perform grouping and synthesis on the harmonic component matrix to obtain the hierarchical spectral features; wherein, the hierarchical spectral features include a fundamental frequency trajectory and a harmonic energy ratio.

[0021] Further, performing a complex-domain transformation on the hierarchical spectral features to obtain complex spectral representation data includes:

[0022] Perform time-frequency domain mapping on the hierarchical spectral features to obtain a time-frequency representation vector, and perform phase demodulation on the time-frequency representation vector to obtain phase modulation data; wherein, the phase modulation data includes amplitude envelope data and instantaneous phase data;

[0023] Perform orthogonal decomposition on the phase modulation data through a preset Hilbert filtering technique to obtain an orthogonal component sequence, and perform polar coordinate mapping on the orthogonal component sequence to obtain a polar coordinate parameter set; wherein, the polar coordinate parameter set includes an amplitude radius and a phase angle;

[0024] Perform phase unwrapping on the polar coordinate parameter set to obtain a phase unwrapping coefficient matrix, and perform harmonic reconstruction based on the phase unwrapping coefficient matrix to obtain a harmonic decomposition matrix;

[0025] Perform frequency-domain mapping on the harmonic decomposition matrix through a preset complex exponential transformation technique to obtain a complex exponential sequence, and perform complex-domain synthesis based on the complex exponential sequence to obtain complex frequency-domain coefficients; wherein, the complex frequency-domain coefficients include an amplitude modulation component and a phase modulation component;

[0026] Perform frequency response analysis on the complex frequency-domain coefficients to obtain a frequency response feature matrix, and perform phase compensation on the frequency response feature matrix to obtain a compensation parameter set;

[0027] Perform complex-domain reconstruction on the complex frequency-domain coefficients based on the compensation parameter set to obtain the complex spectral representation data; wherein, the complex spectral representation data includes a frequency-domain amplitude spectrum and a frequency-domain phase spectrum.

[0028] Further, performing frequency-domain reconstruction on the complex spectral representation data through a sub-band energy analysis technique to obtain an acoustic parameter sequence includes:

[0029] Perform energy integration on the complex spectral representation data through a preset critical sub-band analysis technique to obtain a sub-band energy spectrum, and perform spectral peak detection based on the sub-band energy spectrum to obtain a band energy distribution matrix;

[0030] Calculate the spectral centroid of the band energy distribution matrix to obtain a band centroid sequence, and perform sub-band merging based on the band centroid sequence to obtain a reconstructed band matrix;

[0031] Perform sub-band separation on the reconstructed band matrix by a preset cepstrum analysis method to obtain an independent sub-band sequence, and perform harmonic enhancement on the independent sub-band sequence to obtain enhanced spectral coefficients;

[0032] Perform phase reconstruction on the enhanced spectral coefficients to obtain a complex frequency domain parameter set, and perform frequency band recombination based on the complex frequency domain parameter set to obtain a spectral reconstruction matrix; wherein, the complex frequency domain parameter set includes a harmonic gain coefficient and a phase offset coefficient;

[0033] Perform feature mapping on the spectral reconstruction matrix through a preset acoustic parameter extraction model to obtain an acoustic feature set, and perform parameter sorting on the acoustic feature set to obtain the acoustic parameter sequence; wherein, the acoustic parameter sequence includes an acoustic feature vector and an acoustic parameter matrix.

[0034] Further, the performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectral sequence includes:

[0035] Perform time domain expansion on the acoustic parameter sequence to obtain a parameter time series matrix, and perform group delay estimation based on the parameter time series matrix to obtain a delay feature vector; wherein, the delay feature vector includes a phase delay coefficient and a group delay gradient;

[0036] Perform delay decomposition on the delay feature vector through a preset all-pole analysis technique to obtain a delay component set, and perform phase expansion on the delay component set to obtain a phase compensation matrix;

[0037] Perform polynomial fitting on the phase compensation matrix to obtain a compensation coefficient sequence, and perform group delay correction on the compensation coefficient sequence to obtain a correction parameter set;

[0038] Perform frequency band mapping on the correction parameter set through sub-band decomposition to obtain a frequency band correction matrix, and perform delay compensation on the frequency band correction matrix to obtain a compensated spectral set;

[0039] Perform harmonic reconstruction on the compensated spectral set to obtain reconstructed spectral coefficients, and perform spectral synthesis based on the reconstructed spectral coefficients to obtain a synthesized spectral matrix;

[0040] Perform spectral domain calibration on the synthesized spectral matrix through a spectral shaping technique to obtain a calibrated spectral sequence, and perform phase alignment based on the calibrated spectral sequence to obtain the calibrated spectral sequence; wherein, the calibrated spectral sequence includes a calibrated amplitude spectrum and a calibrated phase spectrum.

[0041] Further, performing inverse Fourier reconstruction on the calibrated spectrum sequence based on the segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal, including:

[0042] Performing complex spectrum separation on the calibrated spectrum sequence to obtain a complex spectrum feature matrix, and performing time-domain mapping on the complex spectrum feature matrix to obtain a preliminarily reconstructed time-domain signal;

[0043] Filtering the preliminarily reconstructed time-domain signal through a preset time-domain filter to obtain a denoised signal matrix, and performing amplitude equalization on the denoised signal matrix to obtain an equalized time-domain signal;

[0044] Performing dynamic range correction on the equalized time-domain signal to obtain a dynamically corrected signal, and performing peak constraint processing on the dynamically corrected signal to obtain a constrained time-domain signal; wherein, the dynamically corrected signal includes a dynamic range coefficient and a signal peak distribution;

[0045] Performing time sequence alignment on the constrained time-domain signal through a time sequence adjustment technique to obtain an aligned signal sequence, and performing short-time energy modulation on the aligned signal sequence to obtain a modulation signal matrix;

[0046] Performing harmonic enhancement on the modulation signal matrix to obtain an enhanced signal set, and performing spectrum interpolation on the enhanced signal set to obtain an interpolation signal matrix;

[0047] Performing inverse Fourier transform on the calibrated spectrum sequence based on the interpolation signal matrix through the segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal.

[0048] The present invention also provides an audio digital signal processing device based on an SOC chip, including:

[0049] An acquisition module, configured to acquire an audio digital signal of a target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream;

[0050] A decomposition module, configured to perform multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features;

[0051] A transformation module, configured to perform complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data;

[0052] A reconstruction module, configured to perform frequency domain reconstruction on the complex spectrum representation data through a sub-band energy analysis technique to obtain an acoustic parameter sequence;

[0053] A compensation module, configured to perform group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence;

[0054] A reconstruction module, configured to perform inverse Fourier reconstruction on the calibrated spectrum sequence based on a segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal.

[0055] The present invention further provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0056] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0057] A method for processing an audio digital signal based on an SOC chip provided by the present invention includes the following steps: collecting an audio digital signal of a target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream; performing multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features; performing a complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data; performing frequency domain reconstruction on the complex spectrum representation data through a sub-band energy analysis technique to obtain an acoustic parameter sequence; performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence; performing inverse Fourier reconstruction on the calibrated spectrum sequence based on a segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal, solving the technical problem of how to efficiently process these data without losing key information as the dimension of audio data increases, and implementing the processing of complex spectrum representation data by using a sub-band energy analysis technique and obtaining an acoustic parameter sequence through frequency domain reconstruction. This process can effectively reduce noise interference and enhance useful audio signals at the same time, ensuring the technical effects of the quality and clarity of the output signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a schematic diagram of the steps of a method for processing an audio digital signal based on an SOC chip in an embodiment of the present invention;

[0059] Figure 2 is a block diagram of the structure of a device for processing an audio digital signal based on an SOC chip in an embodiment of the present invention;

[0060] Figure 3 is a schematic block diagram of the structure of a computer device in an embodiment of the present invention.

[0061] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.

[0063] As Figure 1 shown, Figure 1 is a schematic diagram of the steps of an audio digital signal processing method based on an SOC chip in an embodiment of the present invention;

[0064] An embodiment of the present invention provides an audio digital signal processing method based on an SOC chip, including the following steps:

[0065] Step S1, collect audio digital signals from the target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream.

[0066] Specifically, the process of collecting audio digital signals from the target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream is the first step of the audio digital signal processing method based on the SOC chip. The core lies in using multiple sensors to simultaneously capture sound information from different directions and frequencies. Specifically, the multi-channel sensors are arranged at key positions in the target site, and each sensor is responsible for monitoring the sound environment in a specific area. These sensors can work synchronously to ensure that the data obtained from all angles has a high degree of temporal consistency and spatial resolution. For example, in a professional music recording scenario, to capture the sound details of each instrument during a band performance, multiple microphones (i.e., multi-channel sensors) can be placed in different corners of the venue. This can not only record the lead singer's voice but also capture the sounds of various instruments in the background music, including drums, guitars, basses, etc. When these sensors start working, they will convert the captured sound waves into electrical signals in real-time and further digitally process them into audio digital signals. Since each sensor is located and oriented differently, the collected audio data also has its own characteristics, forming the so-called "multi-dimensional audio data stream". This multi-dimensional data not only contains the intensity information of the sound but also covers rich features such as its propagation direction, time delay, and frequency components. For example, in the above music recording example, by analyzing the time difference of arrival of the sounds recorded by different microphones, the relative positions of each instrument in space can be accurately determined, which is crucial for the mixing process in post-production. In addition, considering the complexity of the actual application scenario, such as the presence of background noise, the design of the multi-channel sensor also needs to consider how to effectively filter out unnecessary noise. This usually involves using advanced filtering techniques and algorithms to optimize the signal quality. For example, during the recording of an outdoor concert, surrounding traffic noise or other interference sources may affect the recording. Through a carefully designed sensor layout and built-in noise reduction algorithms, the interference of these external factors can be minimized to ensure that the final obtained audio data stream is both rich and pure. In this way, subsequent multi-resolution spectral decomposition and other processing steps can more accurately extract useful information, laying a solid foundation for high-quality audio signal processing. In short, this initial audio acquisition step is crucial for the entire audio processing flow, directly determining the effectiveness of subsequent processing steps and the quality of the final audio output.

[0067] Step S2, perform multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features.

[0068] Specifically, the process of performing multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features is one of the key steps in the audio digital signal processing method based on the SOC chip, aiming to extract spectral features with different frequency resolutions from the collected multi-dimensional audio data stream. First of all, the multi-resolution spectral decomposition technology allows the system to analyze the audio signal at different resolutions within different frequency ranges, so as to capture richer acoustic information. Specifically, this process is achieved by splitting the audio data stream into multiple sub-bands and applying a spectral analysis algorithm suitable for its frequency range to each sub-band. For example, in a professional music recording scenario, in order to accurately capture the unique timbre and frequency components of each instrument during a band performance, the system needs to apply high-resolution spectral analysis methods in the low-frequency, mid-frequency, and high-frequency ranges respectively. In this process, the system will first preprocess the multi-dimensional audio data stream, such as noise reduction and smoothing processing, to ensure the quality of the input data. Then, advanced spectral analysis tools and technologies, such as the short-time Fourier transform (STFT) or wavelet transform, are used to decompose the audio data stream. These tools can provide a detailed view in the time-frequency plane, enabling different frequency components to be clearly separated. For example, in the above music recording example, by carefully analyzing the low-frequency band of the drum sound, its unique low-frequency resonance can be identified; while analyzing the mid-high frequency band of the guitar and bass sounds helps to capture their harmonic overtone structures. This hierarchical spectral analysis not only improves the frequency resolution but also provides a more accurate spectral feature representation for subsequent processing. In addition, the extraction of hierarchical spectral features is not just a simple spectral decomposition, but also includes further analysis and integration of the signal characteristics within each frequency band. For example, by comparing the energy distribution and phase relationship between different frequency bands, the system can identify which frequency components are most important for specific sound characteristics. This step is crucial for improving the accuracy and efficiency of audio signal processing. In practical applications, considering that there may be complex background noises at a concert venue, such as the cheers of the audience or other environmental noises, the multi-resolution spectral decomposition technology can effectively enhance the useful audio signal while suppressing unnecessary noises by distinguishing the frequency bands where the main sound source and interference noise are located. Therefore, by performing multi-resolution spectral decomposition on the multi-dimensional audio data stream, the system can generate spectral features with rich hierarchical structures, which lay a solid foundation for subsequent complex domain transformation and other advanced processing steps, and ultimately achieve the reconstruction and optimization of high-quality audio signals.

[0069] Step S3: Perform a complex domain transformation on the hierarchical spectral features to obtain complex spectral representation data.

[0070] Specifically, the process of performing a complex domain transformation on the hierarchical spectral features to obtain complex spectral representation data is one of the key steps in the audio digital signal processing method based on an SOC chip. The aim is to further mathematically transform the already extracted hierarchical spectral features to obtain a richer representation of audio information. Specifically, the complex domain transformation is a method of converting data from the real domain to the complex domain. Through this transformation, the phase and amplitude information of the audio signal can be captured and represented more precisely. In actual operation, the hierarchical spectral features first need to be decomposed into their real and imaginary parts for subsequent complex number operations. For example, in a professional music recording scenario, assuming that we have obtained the spectral features of different instrument sounds through multi-resolution spectral decomposition, we then need to use the complex domain transformation to enhance the expressive power of these features. In this process, algorithms such as the discrete Fourier transform (DFT) or its improved version such as the fast Fourier transform (FFT) are applied to map the spectral features in each frequency band from the real domain to the complex domain. This transformation not only retains the original amplitude information but also increases the dimension of the phase information, making the representation of the audio signal more comprehensive and accurate. For example, in the above music recording example, by performing a complex domain transformation on the sounds of the guitar and bass, not only can their respective frequency components be identified, but also the relative phase relationship between each frequency component can be accurately captured, which is crucial for subsequent sound quality optimization. Phase information is particularly important in reconstructing high-quality audio signals because it directly affects the naturalness and clarity of the sound. In addition, the complex domain transformation allows for more complex mathematical analysis and processing of the audio signal. For example, in frequency domain reconstruction, the complex spectral representation data can provide more flexibility and precision. In actual application scenarios, considering that there may be complex background noise interference at a concert venue, the complex domain transformation can effectively separate the pure audio signal by analyzing the phase difference between the noise and the useful audio signal. This step is crucial for improving the accuracy and efficiency of audio signal processing. For example, when processing the sound of drums, due to its relatively large and complex low-frequency components, the complex domain transformation can more precisely extract its unique resonance mode while suppressing other irrelevant frequency components. Therefore, by performing a complex domain transformation on the hierarchical spectral features, the system can generate complex spectral representation data containing rich phase and amplitude information. These data provide a solid foundation for subsequent advanced processing steps such as sub-band energy analysis, group delay compensation, and inverse Fourier reconstruction, ultimately achieving the enhancement and optimization of high-quality audio signals. In summary, the complex domain transformation is not only an important link in the audio signal processing flow, but also the complex spectral representation data obtained through this process greatly improves the accuracy and effect of subsequent processing steps, thus ensuring the quality and clarity of the final output audio signal. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value.

[0071] Step S4, perform frequency-domain reconstruction on the complex spectrum characterization data through sub-band energy analysis technology to obtain an acoustic parameter sequence.

[0072] Specifically, the process of performing frequency-domain reconstruction on the complex spectrum characterization data through sub-band energy analysis technology to obtain an acoustic parameter sequence is one of the important steps in the audio digital signal processing method based on the SOC chip. The aim is to further analyze and optimize the data after complex-domain transformation, so as to extract more accurate acoustic features. Specifically, the sub-band energy analysis technology first divides the complex spectrum characterization data into different frequency sub-bands, and each sub-band represents the signal components within a specific frequency range. In actual operation, the system will perform detailed energy calculation and analysis on the complex spectrum data within each sub-band to identify the energy distribution of each sub-band. For example, in a professional music recording scenario, assuming that we have obtained the complex spectrum characterization data of musical instrument sounds such as guitars, basses, and drums through complex-domain transformation, we then need to use the sub-band energy analysis technology to further optimize this data. In this process, the system will calculate the energy within each sub-band and perform frequency-domain reconstruction based on the energy distribution. This reconstruction not only takes into account the energy levels of each sub-band but also combines phase information, so that the finally generated acoustic parameter sequence can more accurately reflect the characteristics of the original audio signal. For example, in the above music recording example, by performing sub-band energy analysis on the sounds of guitars and basses, it is possible to identify their respective energy distributions within different frequency sub-bands and adjust the energy ratios of each sub-band accordingly to optimize the sound quality. This step is crucial for improving the clarity and naturalness of the audio signal because the sounds of different musical instruments often have unique energy characteristics within specific frequency ranges, and these characteristics directly affect the sound quality and audibility. In addition, the sub-band energy analysis technology can also help remove unnecessary noise interference. In actual application scenarios, there may be complex background noises at a concert venue, such as the cheers of the audience or other environmental noises. By analyzing the energy distribution of each sub-band, the system can identify which sub-bands mainly contain useful audio signals and which sub-bands are mainly noise. For example, when processing the sound of drums, due to its relatively large and complex low-frequency components, sub-band energy analysis can more accurately extract its unique resonance mode while suppressing other irrelevant frequency components. In this way, by adjusting the energy ratios of each sub-band and removing noise, the system can generate high-quality acoustic parameter sequences, which provide a solid foundation for subsequent group delay compensation and inverse Fourier reconstruction. In the specific implementation process, the system will apply advanced algorithms and technologies, such as adaptive filtering and dynamic range compression, to optimize the energy distribution of each sub-band. For example, in the case of processing a mixed recording of multiple musical instruments, the system can dynamically adjust the energy ratios of each sub-band according to the sound characteristics of different musical instruments to ensure that the sound of each musical instrument can be clearly captured and reproduced.Therefore, by performing sub-band energy analysis on the complex spectrum representation data and carrying out frequency-domain reconstruction on this basis, the system can generate acoustic parameter sequences with rich details and high precision. These sequences not only retain the main features of the original audio signal but also significantly improve the clarity and quality of the audio signal. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology has demonstrated its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals.

[0073] Step S5: Perform group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence.

[0074] Specifically, compensating the group delay of the acoustic parameter sequence to obtain the calibrated spectral sequence is one of the key steps in the audio digital signal processing method based on the SOC chip, aiming to ensure the time consistency of the final output audio signal by correcting the time delay differences of different frequency components. Specifically, the group delay compensation technique first identifies the relative time delays between the various frequency components in the acoustic parameter sequence. This delay is usually caused by filters or other non-linear elements in the signal processing process. In actual operation, the system analyzes the phase information within each frequency sub-band and calculates the corresponding group delay value based on this information. For example, in a professional music recording scenario, assuming that we have obtained the acoustic parameter sequences of the sounds of instruments such as guitars, basses, and drums through sub-band energy analysis, we then need to use the group delay compensation technique to further optimize these data. In this process, the system conducts a detailed phase analysis of the acoustic parameter sequence within each frequency sub-band to determine the actual arrival times of the various frequency components. Since different frequency components may experience different delays during transmission and processing, this can lead to distortion of the time structure of the audio signal. To restore the time consistency of the original audio signal, the system applies specific algorithms to calculate and compensate for these group delays. For example, in the above music recording example, by compensating the group delays of the guitar and bass sounds, the time delays of the various frequency components can be precisely adjusted so that the sounds of different instruments can reach the listener's ears synchronously, thereby improving the clarity and naturalness of the overall sound quality. This step is crucial for ensuring the high-fidelity reproduction of the audio signal because any time deviation will affect the harmony and audibility of the sound. In addition, group delay compensation also involves adjusting the phase relationships across the entire frequency spectrum. In actual application scenarios, there may be complex background noise or reverberation effects at a concert venue, and these factors may increase the time delay differences between different frequency components. Through precise group delay compensation, the system can effectively eliminate these delay differences and ensure that each frequency component accurately reflects its time position in the original audio signal. For example, when processing the sound of drums, since it has more complex low-frequency components, group delay compensation can be used to more precisely adjust the time delays of its various frequency components to keep them in sync with the sounds of other instruments. In this way, by adjusting the time delays of the various frequency components and eliminating phase errors, the system can generate a high-quality calibrated spectral sequence. In the specific implementation process, the system applies advanced algorithms and technologies, such as minimum-phase filtering and adaptive delay compensation, to optimize the time delays of each frequency sub-band. For example, in the case of processing a mixed recording of multiple instruments, the system can dynamically adjust the time delays of the various frequency components according to the sound characteristics of different instruments to ensure that the sound of each instrument can be clearly captured and reproduced.Therefore, by performing group delay compensation on the acoustic parameter sequence and generating a calibrated spectral sequence on this basis, the system can significantly improve the temporal consistency and overall quality of the audio signal. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology has demonstrated its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only is the clarity and naturalness of the audio signal improved, but also a more accurate data basis is provided for subsequent inverse Fourier reconstruction.

[0075] Step S6: Based on the segmented fast Fourier transform algorithm in the preset SOC chip, perform inverse Fourier reconstruction on the calibrated spectral sequence to obtain an enhanced audio digital signal.

[0076] Specifically, the process of performing inverse Fourier reconstruction on the calibrated spectral sequence based on the segmented fast Fourier algorithm in the preset SOC chip to obtain the enhanced audio digital signal is the last step in the entire audio processing flow, aiming to convert the frequency-domain data that has been optimized and adjusted multiple times back into a high-quality time-domain signal. Specifically, the system first performs complex spectral separation on the calibrated spectral sequence to generate a complex spectral feature matrix, and then performs time-domain mapping on these matrices to obtain a preliminarily reconstructed time-domain signal. For example, in a professional music recording scenario, assuming that we have obtained the calibrated spectral sequences of the sounds of instruments such as guitars, basses, and drums through group delay compensation, we then need to use inverse Fourier reconstruction technology to further optimize this data. During this process, the system will analyze the calibrated spectral sequence in detail to identify each frequency component and its corresponding phase information. For example, in the above music recording example, assuming that we divide the audio signal into multiple frequency subbands, by performing complex spectral separation on the calibrated spectral sequence of each subband, it can be found that the low-frequency band (such as the 1st to 4th subbands) mainly contains the sounds of drums and basses, while the high-frequency band (such as the 13th to 16th subbands) mainly captures the high-pitched parts of the guitar. These complex spectral feature matrices provide an important basis for subsequent time-domain mapping. Suppose in a 5-second segment, the amplitude of a certain frequency component of the guitar is 0.8 in the low-frequency band and 0.9 in the high-frequency band, indicating that its high-frequency component is more prominent. Then, the system performs time-domain mapping based on the complex spectral feature matrix to generate a preliminarily reconstructed time-domain signal. Time-domain mapping can convert frequency-domain data back into a time-domain signal and restore the time waveform of the original audio signal. For example, when processing the sound of a guitar, assuming the fundamental frequency is 440 Hz, its time waveform can be reconstructed through time-domain mapping. Suppose in a 10-second segment, the time waveform of the fundamental frequency of the guitar shows an upward trend from 0.5 seconds to 2 seconds, then remains stable from 2 seconds to 8 seconds, and finally gradually decreases from 8 seconds to 10 seconds. These time waveforms provide a reference for subsequent filtering processing. Finally, the system performs an inverse Fourier transform on the preliminarily reconstructed time-domain signal based on the segmented fast Fourier algorithm in the preset SOC chip to generate the enhanced audio digital signal. The segmented fast Fourier algorithm can efficiently convert frequency-domain data back into a time-domain signal, ensuring the high quality of the final output audio signal. For example, in the above music recording example, assuming that the segmented fast Fourier algorithm of the system can identify the fundamental frequency trajectory and harmonic energy ratio of the guitar and use them as the enhanced audio digital signal. Suppose in a 15-second segment, the fundamental frequency trajectory of the guitar shows a change process from 440 Hz gradually rising to 494 Hz, and at the same time its harmonic energy ratio indicates that the energy of the second harmonic is about 70% of the fundamental frequency, and the energy of the third harmonic is about 50% of the fundamental frequency. These detailed enhanced audio digital signals not only improve the resolution of the audio signal but also provide more accurate data support for subsequent complex processing.In this way, the system can more accurately capture the sound characteristics of each musical instrument, enabling the listeners to experience a more realistic and immersive auditory experience.

[0077] In a specific embodiment, the audio digital signal of the target scene is collected through a preset multi-channel sensor to obtain a multi-dimensional audio data stream, including:

[0078] Spatial sampling of the target scene is performed through multi-channel sensors in a microphone array to obtain multi-channel sound field data, and source localization calculation is performed based on the multi-channel sound field data to obtain source space distribution data; wherein, the source space distribution data includes source azimuth angle, source distance parameter, and source intensity distribution;

[0079] Acoustic feature extraction is performed on the source space distribution data to obtain an acoustic feature vector, and sound field reconstruction is performed based on the acoustic feature vector to obtain a three-dimensional sound field data stream;

[0080] Modal analysis is performed on the three-dimensional sound field data stream through sound field decomposition to obtain sound field modal coefficients, and sound field synthesis is performed based on the sound field modal coefficients to obtain a multi-dimensional audio data stream; wherein, the multi-dimensional audio data stream includes sound field time-varying characteristics, sound field spatial characteristics, and sound field spectrum characteristics.

[0081] Specifically, the process of collecting audio digital signals from the target site through a preset multi-channel sensor to obtain a multi-dimensional audio data stream includes multiple complex but closely related steps, aiming to capture as rich acoustic information as possible from the physical space and convert it into high-fidelity digital signals. First, the system performs spatial sampling on the target site through the multi-channel sensors in the microphone array. This process not only collects the intensity information of the sound but also records details such as its propagation direction and time delay. For example, in a professional music recording scenario, assume we use an array composed of multiple microphones to capture the sound of a band performing. These microphones are placed at different positions in the venue, enabling the recording of the sound of each instrument from different angles and distances. Next, based on the collected multi-channel sound field data, the system performs sound source localization calculations to determine the spatial distribution of each sound source. This step involves complex mathematical models and algorithms, such as beamforming technology and time difference of arrival (TDOA) methods, for estimating the azimuth angle, distance parameters, and sound source intensity distribution of the sound source. For example, in the above music recording example, by analyzing the sound signals received by different microphones, the specific positions and relative intensities of instruments such as guitars, basses, and drums can be accurately identified, which is crucial for subsequent mixing processing. These sound source spatial distribution data not only provide the basic position information of the sound source but also lay the foundation for subsequent acoustic feature extraction. Subsequently, the system extracts acoustic features from the sound source spatial distribution data to generate acoustic feature vectors. In this process, the system applies a variety of advanced signal processing techniques, such as spectral analysis and pattern recognition, to extract the most representative features from the original data. For example, when processing the sounds of guitars and basses, the system not only focuses on their frequency components but also analyzes their phase relationships and energy distributions to generate detailed acoustic feature vectors. Based on these feature vectors, the system further performs sound field reconstruction to generate a three-dimensional sound field data stream. This process simulates the actual sound field environment, enabling the accurate reproduction of the position and characteristics of each sound source. For example, at a concert venue, through the three-dimensional sound field data stream, the specific positions of each instrument in space can be simulated, allowing the audience to experience a more realistic and immersive auditory experience. Immediately afterwards, the system performs modal analysis on the three-dimensional sound field data stream through sound field decomposition to obtain sound field modal coefficients. This process utilizes modal analysis technology to decompose the complex three-dimensional sound field into multiple independent modes, each mode representing a specific vibration mode or frequency component. For example, when processing the sound of drums, due to its relatively large and complex low-frequency components, modal analysis can more accurately separate each resonance mode and assign corresponding modal coefficients to it. Based on these modal coefficients, the system further performs sound field synthesis to generate a multi-dimensional audio data stream. This data stream not only contains the time-varying characteristics of the sound field but also covers its spatial and spectral characteristics, thus providing comprehensive and detailed information for the final audio processing.In a specific implementation process, the system applies a series of advanced algorithms and technologies, such as adaptive filtering, dynamic range compression, and minimum-phase filtering, etc., to optimize the effects of each processing step. For example, in the case of processing a mixed recording of multiple musical instruments, the system can dynamically adjust the proportion of each modal coefficient according to the sound characteristics of different musical instruments to ensure that the sound of each musical instrument can be clearly captured and reproduced. Therefore, through a series of complex operations such as spatial sampling of multi-channel sound field data, sound source localization calculation, acoustic feature extraction, sound field reconstruction, modal analysis, and sound field synthesis, the system can generate high-quality multi-dimensional audio data streams. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology has demonstrated its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also a solid data foundation is provided for subsequent multi-resolution spectral decomposition and other advanced processing steps. The entire process is closely linked, and each step contributes an indispensable force to the final high-quality audio output.

[0082] In a specific embodiment, the multi-resolution spectral decomposition of the multi-dimensional audio data stream to obtain hierarchical spectral features includes:

[0083] The frequency band of the multi-dimensional audio data stream is divided by wavelet packet decomposition technology to obtain multi-level sub-band coefficients, and the energy density is calculated based on the multi-level sub-band coefficients to obtain frequency band energy distribution data; wherein, the frequency band energy distribution data includes critical band coefficients and a frequency band energy matrix;

[0084] The spectral envelope of the frequency band energy distribution data is extracted to obtain an envelope feature sequence, and the frequency domain modulation analysis is performed based on the envelope feature sequence to obtain a modulation spectral matrix;

[0085] Based on the modulation spectral matrix, the harmonic structure analysis of the multi-dimensional audio data stream is performed to obtain a harmonic component matrix, and the harmonic component matrix is grouped and synthesized to obtain the hierarchical spectral features; wherein, the hierarchical spectral features include a fundamental frequency trajectory and a harmonic energy ratio.

[0086] Specifically, the process of performing multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features is one of the key steps in the audio digital signal processing method based on the SOC chip. The aim is to extract spectral features with different frequency resolutions from the multi-dimensional audio data stream through a series of complex mathematical transformation and analysis techniques. First, the system divides the frequency bands of the multi-dimensional audio data stream through wavelet packet decomposition technology, splits the original audio signal into multiple sub-bands, and calculates the energy density of each sub-band. For example, in a professional music recording scenario, assume that we have collected the sound data of instruments such as guitars, basses, and drums through a microphone array, and these data contain rich frequency information. During the wavelet packet decomposition process, the system generates multi-level sub-band coefficients, which represent the signal components within different frequency bands. To further analyze these coefficients, the system calculates their energy density to obtain the frequency band energy distribution data. Specifically, the frequency band energy distribution data includes critical band coefficients and a frequency band energy matrix. Taking a typical music recording as an example, assume that we divide the audio signal into 16 sub-bands. By calculating the energy density of each sub-band, it can be found that the low-frequency band (such as the first to fourth sub-bands) mainly contains the sounds of drums and basses, while the high-frequency band (such as the 13th to 16th sub-bands) mainly captures the high-pitched parts of the guitar. This frequency band energy distribution data provides an important basis for subsequent spectral envelope extraction. Next, the system extracts the spectral envelope from the frequency band energy distribution data to obtain an envelope feature sequence. This step not only captures the energy change trend of each frequency band but also lays the foundation for further frequency domain modulation analysis. For example, in the above music recording example, by analyzing the envelope feature sequence, it is possible to identify the volume change pattern of the guitar during the performance and the intensity fluctuations of the drums under different rhythms. Based on these envelope feature sequences, the system performs frequency domain modulation analysis to obtain a modulation spectrum matrix. The modulation spectrum matrix can reveal the hidden modulation characteristics in the audio signal. For example, in a 5-second music segment, the modulation spectrum matrix may show that there are 3 obvious modulation phenomena per second, which is crucial for understanding the overall dynamic changes of the music. Subsequently, the system performs harmonic structure analysis on the multi-dimensional audio data stream based on the modulation spectrum matrix to obtain a harmonic component matrix. This process uses harmonic analysis technology to separate the fundamental frequency and its harmonic components from the complex audio signal. For example, when processing the sound of a guitar, assume that the fundamental frequency is 440 Hz. Through harmonic structure analysis, its second harmonic (880 Hz), third harmonic (1320 Hz), etc. can be identified, and the energy ratio of each harmonic can be calculated. These harmonic component matrices not only reveal the harmonic structure of the audio signal but also provide key information for the final extraction of hierarchical spectral features. Finally, the system performs grouping and synthesis on the harmonic component matrix to obtain hierarchical spectral features. The hierarchical spectral features include the fundamental frequency trajectory and the harmonic energy ratio.For example, in the above example of music recording, assume that in a 10-second segment, the fundamental frequency trajectory of the guitar shows a gradual change from 440 Hz to 494 Hz, and at the same time, its harmonic energy ratio indicates that the energy of the second harmonic is about 70% of the fundamental frequency, and the energy of the third harmonic is about 50% of the fundamental frequency. Such detailed spectral features not only improve the resolution of the audio signal but also provide a solid foundation for subsequent complex domain transformations and other advanced processing steps. Through the multi-resolution spectral decomposition technology, every detail of the audio signal can be accurately captured and reproduced. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also more accurate data support is provided for subsequent complex processing.

[0087] In a specific embodiment, the complex domain transformation of the hierarchical spectral features to obtain complex spectral representation data includes:

[0088] Performing time-frequency domain mapping on the hierarchical spectral features to obtain a time-frequency representation vector, and performing phase demodulation on the time-frequency representation vector to obtain phase modulation data; wherein, the phase modulation data includes amplitude envelope data and instantaneous phase data;

[0089] Performing orthogonal decomposition on the phase modulation data through a preset Hilbert filtering technique to obtain an orthogonal component sequence, and performing polar coordinate mapping on the orthogonal component sequence to obtain a polar coordinate parameter set; wherein, the polar coordinate parameter set includes an amplitude radius and a phase angle;

[0090] Performing phase unwrapping on the polar coordinate parameter set to obtain a phase unwrapping coefficient matrix, and performing harmonic reconstruction based on the phase unwrapping coefficient matrix to obtain a harmonic decomposition matrix;

[0091] Performing frequency domain mapping on the harmonic decomposition matrix through a preset complex exponential transformation technique to obtain a complex exponential sequence, and performing complex domain synthesis based on the complex exponential sequence to obtain complex frequency domain coefficients; wherein, the complex frequency domain coefficients include an amplitude modulation component and a phase modulation component;

[0092] Performing frequency response analysis on the complex frequency domain coefficients to obtain a frequency response characteristic matrix, and performing phase compensation on the frequency response characteristic matrix to obtain a compensation parameter set;

[0093] Performing complex domain reconstruction on the complex frequency domain coefficients based on the compensation parameter set to obtain the complex spectral representation data; wherein, the complex spectral representation data includes a frequency domain amplitude spectrum and a frequency domain phase spectrum.

[0094] Specifically, the process of performing a complex-domain transformation on the hierarchical spectral features to obtain complex spectral representation data is one of the key steps in the audio digital signal processing method based on the SOC chip. The aim is to extract richer audio information representations from the hierarchical spectral features through a series of complex mathematical transformation and analysis techniques. First, the system performs a time-frequency domain mapping on the hierarchical spectral features to generate time-frequency representation vectors, and demodulates the phases of these vectors to obtain phase modulation data. For example, in a professional music recording scenario, assuming that we have obtained the hierarchical spectral features of the sounds of instruments such as guitars, basses, and drums through multi-resolution spectral decomposition, we then need to use time-frequency domain mapping techniques to convert these features into time-frequency representation vectors. During this process, the system generates phase modulation data containing amplitude envelope data and instantaneous phase data. Specifically, assume that in a 5-second music segment, the amplitude variation of a certain frequency component of the guitar over the time axis is recorded to form amplitude envelope data; at the same time, the instantaneous phase of this frequency component is accurately captured to form instantaneous phase data. These phase modulation data provide an important basis for subsequent orthogonal decomposition. Subsequently, the system performs orthogonal decomposition on the phase modulation data through a preset Hilbert filtering technique to generate an orthogonal component sequence. The Hilbert filtering technique can effectively separate the real and imaginary parts of the signal, thus preparing for polar coordinate mapping. For example, in the above music recording example, by performing Hilbert filtering on a certain frequency component of the guitar, its corresponding orthogonal component sequence can be obtained. Then, the system performs polar coordinate mapping on these orthogonal component sequences to generate a polar coordinate parameter set, including amplitude radius and phase angle. Assume that in a 10-second segment, the average amplitude radius of a certain frequency component of the guitar is 0.8 and the average phase angle is 45 degrees. These parameters provide detailed references for subsequent phase unwrapping. Next, the system performs phase unwrapping on the polar coordinate parameter set to obtain a phase unwrapping coefficient matrix, and based on these coefficients, performs harmonic reconstruction to generate a harmonic decomposition matrix. For example, when processing the sound of the bass, assuming the fundamental frequency is 55 Hz, through phase unwrapping, its second harmonic (110 Hz), third harmonic (165 Hz), etc. can be identified, and the energy ratio of each harmonic can be calculated. These harmonic decomposition matrices not only reveal the harmonic structure of the audio signal but also provide key information for the final complex exponential transformation. Then, the system performs frequency domain mapping on the harmonic decomposition matrix through a preset complex exponential transformation technique to generate a complex exponential sequence, and based on these sequences, performs complex-domain synthesis to obtain complex frequency domain coefficients. The complex frequency domain coefficients include amplitude modulation components and phase modulation components, and these coefficients lay the foundation for subsequent frequency response analysis. For example, when processing the sound of the drums, assuming that the amplitude modulation component of a certain frequency component is 0.7 and the phase modulation component is 30 degrees, these detailed coefficients can better reflect the true characteristics of the audio signal.Furthermore, the system performs frequency response analysis on the complex frequency domain coefficients to generate a frequency response characteristic matrix, and compensates the phases of these matrices to obtain a set of compensation parameters. For example, in a 20-second music clip, assume that after frequency response analysis of a certain frequency component of a guitar, it is found that a phase compensation of 3 degrees is required to ensure the consistency of the phase relationship with other frequency components. These sets of compensation parameters provide the necessary adjustment basis for the final complex domain reconstruction. Finally, the system performs complex domain reconstruction on the complex frequency domain coefficients based on the set of compensation parameters to generate complex spectrum characterization data. The complex spectrum characterization data includes the frequency domain amplitude spectrum and the frequency domain phase spectrum. These data not only retain the main features of the original audio signal but also significantly improve the clarity and quality of the audio signal. For example, in the above example of music recording, assume that in a 30-second clip, the frequency domain amplitude spectrum of the guitar shows a significant energy enhancement in the frequency band from 440 Hz to 880 Hz, while the frequency domain phase spectrum shows that the phase relationship between each frequency component has been precisely calibrated. Through the complex domain transformation technology, every detail of the audio signal can be accurately captured and reproduced. In both professional music recording and other high-fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also more accurate data support is provided for subsequent complex processing.

[0095] In a specific embodiment, the frequency domain reconstruction of the complex spectrum characterization data by the sub-band energy analysis technology to obtain an acoustic parameter sequence includes:

[0096] Performing energy integration on the complex spectrum characterization data through a preset critical sub-band analysis technology to obtain a sub-band energy spectrum, and performing spectral peak detection based on the sub-band energy spectrum to obtain a frequency band energy distribution matrix;

[0097] Calculating the spectral centroid of the frequency band energy distribution matrix to obtain a frequency band centroid sequence, and performing sub-band merging based on the frequency band centroid sequence to obtain a reconstructed frequency band matrix;

[0098] Performing sub-band separation on the reconstructed frequency band matrix through a preset cepstrum analysis method to obtain an independent sub-band sequence, and performing harmonic enhancement on the independent sub-band sequence to obtain enhanced spectral coefficients;

[0099] Performing phase reconstruction on the enhanced spectral coefficients to obtain a set of complex frequency domain parameters, and performing frequency band recombination based on the set of complex frequency domain parameters to obtain a spectrum reconstruction matrix; wherein, the set of complex frequency domain parameters includes a harmonic gain coefficient and a phase offset coefficient;

[0100] Performing feature mapping on the spectrum reconstruction matrix through a preset acoustic parameter extraction model to obtain an acoustic feature set, and performing parameter sorting on the acoustic feature set to obtain the acoustic parameter sequence; wherein, the acoustic parameter sequence includes an acoustic feature vector and an acoustic parameter matrix.

[0101] Specifically, the process of performing frequency-domain reconstruction on the complex spectrum characterization data through sub-band energy analysis technology to obtain an acoustic parameter sequence is one of the important steps in the audio digital signal processing method based on the SOC chip. The aim is to generate a high-quality acoustic parameter sequence through detailed energy analysis and frequency-domain reconstruction of the complex spectrum characterization data. First, the system performs energy integration on the complex spectrum characterization data through a preset critical sub-band analysis technology to obtain the sub-band energy spectrum, and performs spectral peak detection based on these energy spectra, thereby generating a frequency-band energy distribution matrix. For example, in a professional music recording scenario, assuming that we have obtained the complex spectrum characterization data of the sounds of instruments such as guitars, basses, and drums through complex-domain transformation, we then need to use sub-band energy analysis technology to further optimize these data. In this process, the system will perform detailed energy integration calculations on the complex spectrum characterization data to identify the energy distribution of each frequency sub-band. For example, in the above music recording example, assuming that we divide the audio signal into 16 sub-bands, by integrating the energy of each sub-band, it can be found that the low-frequency band (such as the first to fourth sub-bands) mainly contains the sounds of drums and basses, while the high-frequency band (such as the 13th to 16th sub-bands) mainly captures the high-pitched parts of the guitar. These sub-band energy spectra provide an important basis for subsequent spectral peak detection. Assuming that in a 5-second segment, the energy ratio of a certain frequency component of the guitar in the low-frequency band is 20%, while the energy ratio in the high-frequency band is as high as 80%, this indicates that its high-frequency component is more prominent. Then, the system performs spectral peak detection based on the sub-band energy spectrum to generate a frequency-band energy distribution matrix. Spectral peak detection can help the system identify the main energy peaks within each sub-band, thereby better understanding the frequency characteristics of the audio signal. For example, when processing the sound of a guitar, assuming that the fundamental frequency is 440 Hz, the energy peaks of its second harmonic (880 Hz), third harmonic (1320 Hz), etc. can be identified through spectral peak detection. Assuming that in a 10-second segment, the energy peak of the fundamental frequency of the guitar is 0.7, the energy peak of the second harmonic is 0.5, and the energy peak of the third harmonic is 0.3, these energy peaks reflect the relative intensities of different frequency components. Subsequently, the system performs spectral centroid calculation on the frequency-band energy distribution matrix to obtain the frequency-band centroid sequence, and performs sub-band merging based on these sequences to generate a reconstructed frequency-band matrix. Spectral centroid calculation can provide the central frequency position of each sub-band, which is crucial for understanding the overall frequency distribution of the audio signal. For example, in the above music recording example, assuming that the spectral centroid of a certain frequency component of the guitar in the low-frequency band is 50 Hz, while the spectral centroid in the high-frequency band is 1000 Hz, these spectral centroid values provide a reference for subsequent sub-band merging. Assuming that in a 15-second segment, multiple frequency components of the guitar form a new reconstructed frequency-band matrix after sub-band merging, which contains the main frequency components from 50 Hz to 1000 Hz.Then, the system performs sub-band separation on the reconstructed frequency band matrix through a preset cepstrum analysis method to generate independent sub-band sequences, and enhances the harmonics of these sequences to generate enhanced spectral coefficients. Cepstrum analysis can effectively separate the independent frequency components within each sub-band, thus providing a basis for harmonic enhancement. For example, when processing the sound of a bass, assuming the fundamental frequency is 55 Hz, cepstrum analysis can separate components such as its second harmonic (110 Hz) and third harmonic (165 Hz), and enhance their harmonics. Suppose in a 20-second segment, the fundamental frequency energy of the bass is enhanced by 20%, the second harmonic energy is enhanced by 15%, and the third harmonic energy is enhanced by 10%. These enhanced spectral coefficients not only improve the quality of the audio signal but also lay a foundation for subsequent phase reconstruction. Next, the system performs phase reconstruction on the enhanced spectral coefficients to generate a set of complex frequency domain parameters, and based on these parameter sets, performs frequency band recombination to generate a spectral reconstruction matrix. Phase reconstruction can restore the time structure of the audio signal, ensuring its continuity and naturalness in the time domain. For example, when processing the sound of drums, assuming the amplitude modulation component of a certain frequency component is 0.7 and the phase modulation component is 30 degrees, phase reconstruction can precisely adjust its phase relationship to keep it synchronized with other frequency components. Suppose in a 25-second segment, multiple frequency components of the drums form a new spectral reconstruction matrix after phase reconstruction, which contains the main frequency components from 20 Hz to 500 Hz and their corresponding phase information. Further, the system performs feature mapping on the spectral reconstruction matrix through a preset acoustic parameter extraction model to generate a set of acoustic features, and sorts these sets of parameters to finally generate an acoustic parameter sequence. The acoustic parameter extraction model can extract the most representative acoustic features from the spectral reconstruction matrix, and these features are crucial for subsequent audio processing. For example, in the above example of music recording, assume that the system's acoustic parameter extraction model can identify the fundamental frequency trajectory and harmonic energy ratio of the guitar and use them as an acoustic feature vector. Suppose in a 30-second segment, the fundamental frequency trajectory of the guitar shows a change process from 440 Hz to 494 Hz, and its harmonic energy ratio indicates that the energy of the second harmonic is about 70% of the fundamental frequency, and the energy of the third harmonic is about 50% of the fundamental frequency. These detailed acoustic features not only improve the resolution of the audio signal but also provide more accurate data support for subsequent complex processing. Throughout the process, the system extracts a rich sequence of acoustic parameters from the complex spectral representation data through a series of complex mathematical transformation and analysis techniques. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also more accurate data support is provided for subsequent complex processing.For example, at a concert venue, through these detailed sequences of acoustic parameters, the system can more precisely capture the sound characteristics of each musical instrument, enabling the audience to experience a more vivid and immersive auditory experience. In addition, these acoustic parameters can also be used for post-mixing processing to ensure that the sound of each musical instrument can be clearly captured and reproduced. The entire process uses subband energy analysis technology to precisely capture and reproduce every detail of the audio signal, ultimately achieving the goal of enhancing and optimizing high-quality audio signals.

[0102] In a specific embodiment, the group delay compensation for the acoustic parameter sequence to obtain a calibrated spectrum sequence includes:

[0103] Performing time-domain expansion on the acoustic parameter sequence to obtain a parameter time-series matrix, and performing group delay estimation based on the parameter time-series matrix to obtain a delay feature vector; wherein, the delay feature vector includes a phase delay coefficient and a group delay gradient;

[0104] Performing delay decomposition on the delay feature vector through a preset all-pole analysis technique to obtain a set of delay components, and performing phase expansion on the set of delay components to obtain a phase compensation matrix;

[0105] Performing polynomial fitting on the phase compensation matrix to obtain a compensation coefficient sequence, and performing group delay correction on the compensation coefficient sequence to obtain a set of correction parameters;

[0106] Performing frequency band mapping on the set of correction parameters through subband decomposition to obtain a frequency band correction matrix, and performing delay compensation on the frequency band correction matrix to obtain a set of compensated spectra;

[0107] Performing harmonic reconstruction on the set of compensated spectra to obtain reconstructed spectral coefficients, and performing spectrum synthesis based on the reconstructed spectral coefficients to obtain a synthesized spectrum matrix;

[0108] Performing spectral domain calibration on the synthesized spectrum matrix through spectral shaping technology to obtain a calibrated spectrum sequence, and performing phase alignment based on the calibrated spectrum sequence to obtain the calibrated spectrum sequence; wherein, the calibrated spectrum sequence includes a calibrated amplitude spectrum and a calibrated phase spectrum.

[0109] Specifically, the process of performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectral sequence is one of the key steps in the audio digital signal processing method based on the SOC chip. It aims to eliminate the time delay differences between different frequency components through a series of complex mathematical transformation and analysis techniques, thereby ensuring the time consistency of the final output audio signal. First, the system expands the acoustic parameter sequence in the time domain to generate a parameter time series matrix, and based on these matrices, performs group delay estimation to obtain a delay eigenvector. For example, in a professional music recording scenario, assume that we have obtained the acoustic parameter sequences of the sounds of instruments such as guitars, basses, and drums through sub-band energy analysis. Next, we need to further optimize these data using group delay compensation technology. In this process, the system will perform detailed time-domain expansion calculations on the acoustic parameter sequence to identify the relative time delays between each frequency component. For example, in the above music recording example, assume that we divide the audio signal into multiple frequency sub-bands. By expanding the acoustic parameter sequence of each sub-band in the time domain, it can be found that the low-frequency band (such as the 1st to 4th sub-bands) mainly contains the sounds of drums and basses, while the high-frequency band (such as the 13th to 16th sub-bands) mainly captures the high-pitched parts of the guitars. These parameter time series matrices provide an important basis for subsequent group delay estimation. Assume that in a 5-second segment, the phase delay coefficient of a certain frequency component of the guitar in the low-frequency band is 0.2, while in the high-frequency band it is 0.8, indicating that there is a significant time delay in its high-frequency components. Then, the system performs group delay estimation based on the parameter time series matrix to generate a delay eigenvector. Group delay estimation can help the system identify the actual arrival times of each frequency component, thereby better understanding the time structure of the audio signal. For example, when processing the sound of a guitar, assume that the fundamental frequency is 440 Hz. Through group delay estimation, the actual arrival times of its second harmonic (880 Hz), third harmonic (1320 Hz), etc. can be identified. Assume that in a 10-second segment, the time delay of the fundamental frequency of the guitar is 2 milliseconds, the time delay of the second harmonic is 4 milliseconds, and the time delay of the third harmonic is 6 milliseconds. These time delays reflect the relative delay relationships between different frequency components. Subsequently, the system performs delay decomposition on the delay eigenvector through a preset all-pole analysis technique to generate a set of delay components, and performs phase unwrapping on these sets to generate a phase compensation matrix. The all-pole analysis technique can effectively separate the delay characteristics of each frequency component, thereby providing a basis for phase unwrapping. For example, when processing the sound of a bass, assume that the fundamental frequency is 55 Hz. Through all-pole analysis, the delay components of its second harmonic (110 Hz), third harmonic (165 Hz), etc. can be separated and phase unwrapped.Suppose that in a 15 - second segment, the fundamental frequency time delay of the bass is 1 millisecond, the second - harmonic time delay is 2 milliseconds, and the third - harmonic time delay is 3 milliseconds. These delay components provide a reference for subsequent polynomial fitting. Then, the system performs polynomial fitting on the phase compensation matrix to generate a sequence of compensation coefficients, and performs group - delay correction on these sequences to generate a set of correction parameters. Polynomial fitting can smoothly adjust the time delays of each frequency component to ensure its consistency across the entire frequency spectrum. For example, when processing the sound of a drum, assume that the amplitude modulation component of a certain frequency component is 0.7 and the phase modulation component is 30 degrees. Through polynomial fitting, its phase relationship can be precisely adjusted to keep it in sync with other frequency components. Suppose that in a 20 - second segment, multiple frequency components of the drum form a new set of correction parameters after polynomial fitting, which includes the main frequency components from 20 Hz to 500 Hz and their corresponding phase adjustment values. Next, the system performs frequency - band mapping on the set of correction parameters through sub - band decomposition to generate frequency - band correction matrices, and performs delay compensation on these matrices to generate a set of compensated spectra. Sub - band decomposition can divide the frequency components across the entire frequency spectrum into multiple independent sub - bands, thus providing a basis for delay compensation. For example, in the above music recording example, assume that the frequency - band correction matrix of a certain frequency component of the guitar in the low - frequency band is [[0.8, 0.9], [0.7, 0.8]], and in the high - frequency band is [[0.6, 0.7], [0.5, 0.6]]. These frequency - band correction matrices provide a reference for subsequent harmonic reconstruction. Suppose that in a 25 - second segment, multiple frequency components of the guitar form a new set of compensated spectra after delay compensation, which includes the main frequency components from 50 Hz to 1000 Hz and their corresponding delay compensation values. Further, the system performs harmonic reconstruction on the set of compensated spectra to generate reconstructed spectral coefficients, and performs spectral synthesis based on these coefficients to generate a synthesized spectral matrix. Harmonic reconstruction can restore the original frequency structure of the audio signal to ensure its continuity and naturalness in the frequency domain. For example, when processing the sound of the bass, assume that the fundamental frequency is 55 Hz. Through harmonic reconstruction, its second - harmonic (110 Hz), third - harmonic (165 Hz), etc. components can be separated and recombined. Suppose that in a 30 - second segment, the energy of the fundamental frequency of the bass is increased by 20%, the energy of the second - harmonic is increased by 15%, and the energy of the third - harmonic is increased by 10%. These enhanced reconstructed spectral coefficients not only improve the quality of the audio signal but also lay a foundation for subsequent spectral shaping. Finally, the system performs spectral - domain calibration on the synthesized spectral matrix through spectral - shaping technology to generate a calibrated spectral sequence, and performs phase alignment based on these sequences to generate a calibrated spectral sequence. Spectral - shaping technology can smoothly adjust the amplitude and phase of each frequency component to ensure its consistency across the entire frequency spectrum.For example, in the above example of music recording, assume that the system's spectral shaping technology can identify the fundamental frequency trajectory and harmonic energy ratio of the guitar and use them as the calibration spectrum sequence. Assume that in a 35-second segment, the fundamental frequency trajectory of the guitar shows a gradual increase from 440 Hz to 494 Hz, while its harmonic energy ratio indicates that the energy of the second harmonic is about 70% of the fundamental frequency and the energy of the third harmonic is about 50% of the fundamental frequency. These detailed calibration spectrum sequences not only improve the resolution of the audio signal but also provide more accurate data support for subsequent complex processing. Throughout the process, the system extracts a high-quality calibrated spectral sequence from the acoustic parameter sequence through a series of complex mathematical transformations and analysis techniques. Whether in professional music recording or other high-fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value, ultimately achieving the enhancement and optimization of high-quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also more accurate data support is provided for subsequent complex processing. For example, at a concert venue, through these detailed calibrated spectral sequences, the system can more precisely capture the sound characteristics of each instrument, enabling the audience to experience a more realistic and immersive auditory experience. In addition, these calibrated spectra can be used for post-mixing processing to ensure that the sound of each instrument can be clearly captured and reproduced. The entire process uses group delay compensation technology to precisely capture and reproduce every detail of the audio signal, ultimately achieving the goal of enhancing and optimizing high-quality audio signals.

[0110] In a specific embodiment, performing an inverse Fourier reconstruction on the calibrated spectral sequence based on the segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal, including:

[0111] Performing complex spectral separation on the calibrated spectral sequence to obtain a complex spectral feature matrix, and performing a time-domain mapping on the complex spectral feature matrix to obtain a preliminarily reconstructed time-domain signal;

[0112] Filtering the preliminarily reconstructed time-domain signal through a preset time-domain filter to obtain a denoised signal matrix, and performing amplitude equalization on the denoised signal matrix to obtain an equalized time-domain signal;

[0113] Performing dynamic range correction on the equalized time-domain signal to obtain a dynamically corrected signal, and performing peak constraint processing on the dynamically corrected signal to obtain a constrained time-domain signal; wherein, the dynamically corrected signal includes a dynamic range coefficient and a signal peak distribution;

[0114] Performing time-sequence alignment on the constrained time-domain signal through time-sequence adjustment technology to obtain an aligned signal sequence, and performing short-time energy modulation on the aligned signal sequence to obtain a modulation signal matrix;

[0115] Perform harmonic enhancement on the modulation signal matrix to obtain an enhanced signal set, and perform spectral interpolation on the enhanced signal set to obtain an interpolated signal matrix;

[0116] Based on the interpolated signal matrix, perform an inverse Fourier transform on the calibrated spectral sequence through the segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal.

[0117] Specifically, the process of performing inverse Fourier reconstruction on the calibrated spectrum sequence based on the segmented fast Fourier algorithm in the preset SOC chip to obtain the enhanced audio digital signal is the last step in the audio digital signal processing method based on the SOC chip. It aims to re-convert the frequency-domain data that has undergone multiple optimizations and adjustments into a high-quality time-domain signal through a series of complex mathematical transformation and analysis techniques. First, the system performs complex spectrum separation on the calibrated spectrum sequence to generate a complex spectrum feature matrix, and performs time-domain mapping on these matrices to obtain a preliminarily reconstructed time-domain signal. For example, in a professional music recording scenario, assume that we have obtained the calibrated spectrum sequences of the sounds of instruments such as guitars, basses, and drums through group delay compensation. Next, we need to further optimize these data using inverse Fourier reconstruction technology. In this process, the system will perform detailed complex spectrum separation calculations on the calibrated spectrum sequence to identify each frequency component and its corresponding phase information. For example, in the above music recording example, assume that we divide the audio signal into multiple frequency subbands. By performing complex spectrum separation on the calibrated spectrum sequence of each subband, it can be found that the low-frequency band (such as the 1st to 4th subbands) mainly contains the sounds of drums and basses, while the high-frequency band (such as the 13th to 16th subbands) mainly captures the high-pitched parts of the guitar. These complex spectrum feature matrices provide an important basis for subsequent time-domain mapping. Assume that in a 5-second segment, the amplitude of a certain frequency component of the guitar is 0.8 in the low-frequency band and 0.9 in the high-frequency band, indicating that its high-frequency component is more prominent. Then, the system performs time-domain mapping based on the complex spectrum feature matrix to generate a preliminarily reconstructed time-domain signal. Time-domain mapping can convert frequency-domain data back into a time-domain signal and restore the time waveform of the original audio signal. For example, when processing the sound of a guitar, assume that the fundamental frequency is 440 Hz. Through time-domain mapping, its time waveform can be reconstructed. Assume that in a 10-second segment, the time waveform of the fundamental frequency of the guitar shows an upward trend from 0.5 seconds to 2 seconds, then remains stable from 2 seconds to 8 seconds, and finally gradually decreases from 8 seconds to 10 seconds. These time waveforms provide a reference for subsequent filtering processing. Subsequently, the system filters the preliminarily reconstructed time-domain signal through a preset time-domain filter to generate a denoised signal matrix, and performs amplitude equalization on these matrices to generate an equalized time-domain signal. The time-domain filter can effectively remove background noise and other interference signals to ensure the clarity of the audio signal. For example, when processing the sound of a bass, assume that the fundamental frequency is 55 Hz. Through time-domain filtering, environmental noise can be removed and the main frequency components can be retained. Assume that in a 15-second segment, the fundamental frequency energy of the bass increases by 20%, the second harmonic energy increases by 15%, and the third harmonic energy increases by 10%. These enhanced denoised signals not only improve the quality of the audio signal but also lay the foundation for subsequent amplitude equalization.Suppose in a 20 - second segment, multiple frequency components of the bass, after amplitude equalization, form a new equalized time - domain signal, which contains the main frequency components from 50 Hz to 1000 Hz and their corresponding amplitude adjustment values. Then, the system performs dynamic range correction on the equalized time - domain signal to generate a dynamically corrected signal, and performs peak constraint processing on these signals to generate a constrained time - domain signal. Dynamic range correction can smoothly adjust the dynamic range of the audio signal to ensure its consistency over the entire time axis. For example, when processing the sound of drums, assume that the amplitude modulation component of a certain frequency component is 0.7 and the phase modulation component is 30 degrees. Through dynamic range correction, its dynamic range can be precisely adjusted to keep it in sync with other frequency components. Suppose in a 25 - second segment, multiple frequency components of the drums, after dynamic range correction, form a new dynamically corrected signal, which contains the main frequency components from 20 Hz to 500 Hz and their corresponding dynamic range coefficients. Assume that in this segment, the average value of the dynamic range coefficients is 0.8, and the signal peak distribution ranges from 0.6 to 0.9. These adjusted signals provide a reference for the subsequent peak constraint processing. Next, the system performs timing alignment on the constrained time - domain signal through timing adjustment techniques to generate an aligned signal sequence, and performs short - time energy modulation on these sequences to generate a modulated signal matrix. Timing adjustment techniques can ensure that the relative time relationships between various frequency components remain consistent, thereby improving the time consistency of the audio signal. For example, in the above example of music recording, assume that the aligned signal sequence of a certain frequency component of the guitar in the low - frequency band is [[0.8, 0.9], [0.7, 0.8]], and in the high - frequency band is [[0.6, 0.7], [0.5, 0.6]]. These aligned signal sequences provide a reference for the subsequent short - time energy modulation. Suppose in a 30 - second segment, multiple frequency components of the guitar, after short - time energy modulation, form a new modulated signal matrix, which contains the main frequency components from 50 Hz to 1000 Hz and their corresponding energy adjustment values. Further, the system performs harmonic enhancement on the modulated signal matrix to generate an enhanced signal set, and performs spectral interpolation on these sets to generate an interpolated signal matrix. Harmonic enhancement can restore the original frequency structure of the audio signal to ensure its continuity and naturalness in the frequency domain. For example, when processing the sound of the bass, assume that the fundamental frequency is 55 Hz. Through harmonic enhancement, its second harmonic (110 Hz), third harmonic (165 Hz), etc. can be separated and recombined. Suppose in a 35 - second segment, the fundamental frequency energy of the bass is enhanced by 20%, the second - harmonic energy is enhanced by 15%, and the third - harmonic energy is enhanced by 10%. These enhanced signal sets not only improve the quality of the audio signal but also lay a foundation for the subsequent spectral interpolation.Suppose in a 40 - second segment, multiple frequency components of the bass form a new interpolated signal matrix after spectral interpolation, which contains the main frequency components from 50 Hz to 1000 Hz and their corresponding interpolation adjustment values. Finally, based on the segmented fast Fourier transform algorithm in the preset SOC chip, the system performs an inverse Fourier transform on the calibrated spectral sequence based on the interpolated signal matrix to generate an enhanced audio digital signal. The segmented fast Fourier transform algorithm can efficiently convert frequency - domain data back into a time - domain signal, ensuring the high quality of the final output audio signal. For example, in the above music recording example, assume that the segmented fast Fourier transform algorithm of the system can identify the fundamental frequency trajectory and harmonic energy ratio of the guitar and use them as the enhanced audio digital signal. Suppose in a 45 - second segment, the fundamental frequency trajectory of the guitar shows a change from 440 Hz to 494 Hz gradually, and its harmonic energy ratio indicates that the energy of the second harmonic is about 70% of the fundamental frequency, and the energy of the third harmonic is about 50% of the fundamental frequency. These detailed enhanced audio digital signals not only improve the resolution of the audio signal but also provide more accurate data support for subsequent complex processing. Throughout the process, the system extracts high - quality enhanced audio digital signals from the calibrated spectral sequence through a series of complex mathematical transformations and analysis techniques. Whether in professional music recording or other high - fidelity audio processing application scenarios, this technology demonstrates its unique advantages and value, ultimately achieving the enhancement and optimization of high - quality audio signals. In this way, not only the clarity and naturalness of the audio signal are improved, but also more accurate data support is provided for subsequent complex processing. For example, at a concert venue, through these detailed enhanced audio digital signals, the system can more precisely capture the sound characteristics of each instrument, enabling the audience to experience a more vivid and immersive auditory experience. In addition, these enhanced audio digital signals can also be used for post - mixing processing to ensure that the sound of each instrument can be clearly captured and reproduced. The entire process uses inverse Fourier reconstruction technology to precisely capture and reproduce every detail of the audio signal, ultimately achieving the goal of enhancing and optimizing high - quality audio signals.

[0118] The above described the audio digital signal processing method based on the SOC chip in the embodiments of the present invention. Next, the audio digital signal processing device based on the SOC chip in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the audio digital signal processing device based on the SOC chip in the embodiments of the present invention includes:

[0119] An acquisition module 21, configured to collect audio digital signals of a target scene through a preset multi - channel sensor to obtain a multi - dimensional audio data stream;

[0120] A decomposition module 22 for performing multi-resolution spectral decomposition on the multi-dimensional audio data stream to obtain hierarchical spectral features;

[0121] A transformation module 23 for performing a complex-domain transformation on the hierarchical spectral features to obtain complex spectral representation data;

[0122] A reconstruction module 24 for performing frequency-domain reconstruction on the complex spectral representation data through sub-band energy analysis technology to obtain an acoustic parameter sequence;

[0123] A compensation module 25 for performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectral sequence;

[0124] A reconstruction module 26 for performing inverse Fourier reconstruction on the calibrated spectral sequence based on the segmented fast Fourier algorithm in a preset SOC chip to obtain an enhanced audio digital signal.

[0125] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to that described in the above method embodiment, and details are not described herein again.

[0126] Refer to Figure 3 , and this embodiment of the present invention also provides a computer device, the internal structure of which can be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.

[0127] Those skilled in the art can understand that Figure 3 the structure shown in

[0128] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0130] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method that includes such element.

[0131] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. An audio digital signal processing method based on a SOC chip, characterized in that: The following steps are involved: The audio digital signal is collected at the target scene through a preset multi-channel sensor to obtain a multi-dimensional audio data stream; Performing multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features; Performing complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data; Reconstructing the complex spectrum characterization data in the frequency domain by using a sub-band energy analysis technique to obtain an acoustic parameter sequence; Performing group delay compensation on the acoustic parameter sequence to obtain a calibrated frequency spectrum sequence; Based on the preset segmented fast Fourier algorithm in the SOC chip, the calibrated spectrum sequence is subjected to inverse Fourier reconstruction to obtain an enhanced audio digital signal; The performing complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data includes: Performing time-frequency domain mapping on the hierarchical spectrum features to obtain a time-frequency characterization vector, and performing phase demodulation on the time-frequency characterization vector to obtain phase modulation data; wherein the phase modulation data includes amplitude envelope data and instantaneous phase data; The phase modulation data is orthogonally decomposed by a preset Hilbert filtering technique to obtain an orthogonal component sequence, and the orthogonal component sequence is polar mapped to obtain a polar coordinate parameter set; wherein the polar coordinate parameter set includes an amplitude radius and a phase angle; Performing phase expansion on the polar coordinate parameter set to obtain a phase expansion coefficient matrix, and performing harmonic reconstruction based on the phase expansion coefficient matrix to obtain a harmonic decomposition matrix; The harmonic decomposition matrix is ​​mapped to the frequency domain by a preset complex exponential transformation technology to obtain a complex exponential sequence, and complex domain synthesis is performed based on the complex exponential sequence to obtain complex frequency domain coefficients; wherein the complex frequency domain coefficients include an amplitude modulation component and a phase modulation component; Performing frequency response analysis on the complex frequency domain coefficients to obtain a frequency response characteristic matrix, and performing phase compensation on the frequency response characteristic matrix to obtain a compensation parameter set; The complex frequency domain coefficients are reconstructed in a complex domain based on the compensation parameter set to obtain the complex frequency spectrum characterization data; wherein the complex frequency spectrum characterization data includes a frequency domain amplitude spectrum and a frequency domain phase spectrum.

2. The audio digital signal processing method based on SOC chip according to claim 1, characterized in that: The method collects audio digital signals from the target scene through a preset multi-channel sensor to obtain a multi-dimensional audio data stream, including: The target scene is spatially sampled by a multi-channel sensor in a microphone array to obtain multi-channel sound field data, and sound source localization calculation is performed based on the multi-channel sound field data to obtain sound source spatial distribution data; wherein the sound source spatial distribution data includes a sound source azimuth angle, a sound source distance parameter, and a sound source intensity distribution; Extracting acoustic features from the sound source spatial distribution data to obtain acoustic feature vectors, and reconstructing the sound field based on the acoustic feature vectors to obtain a three-dimensional sound field data stream; The three-dimensional sound field data stream is subjected to modal analysis by sound field decomposition to obtain sound field modal coefficients, and sound field synthesis is performed based on the sound field modal coefficients to obtain a multi-dimensional audio data stream; wherein the multi-dimensional audio data stream includes time-varying characteristics of the sound field, spatial characteristics of the sound field, and spectral characteristics of the sound field.

3. The audio digital signal processing method based on SOC chip according to claim 1, characterized in that: The performing multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features includes: The multi-dimensional audio data stream is divided into frequency bands by wavelet packet decomposition technology to obtain multi-level sub-band coefficients, and energy density is calculated based on the multi-level sub-band coefficients to obtain frequency band energy distribution data; wherein the frequency band energy distribution data includes critical frequency band coefficients and frequency band energy matrices; Performing spectrum envelope extraction on the frequency band energy distribution data to obtain an envelope feature sequence, and performing frequency domain modulation analysis based on the envelope feature sequence to obtain a modulation spectrum matrix; Based on the modulation spectrum matrix, the harmonic structure of the multi-dimensional audio data stream is analyzed to obtain a harmonic component matrix, and the harmonic component matrix is ​​grouped and synthesized to obtain the hierarchical spectrum feature; wherein the hierarchical spectrum feature includes a fundamental frequency trajectory and a harmonic energy ratio.

4. The audio digital signal processing method based on SOC chip according to claim 1, characterized in that: The method of reconstructing the complex spectrum characterization data in the frequency domain by using the sub-band energy analysis technology to obtain an acoustic parameter sequence includes: Performing energy integration on the complex spectrum characterization data by a preset critical subband analysis technique to obtain a subband energy spectrum, and performing spectrum peak detection based on the subband energy spectrum to obtain a frequency band energy distribution matrix; Calculating the spectral centroid of the frequency band energy distribution matrix to obtain a frequency band centroid sequence, and merging sub-bands based on the frequency band centroid sequence to obtain a reconstructed frequency band matrix; Performing sub-band separation on the reconstructed frequency band matrix by a preset cepstrum analysis method to obtain an independent sub-band sequence, and performing harmonic enhancement on the independent sub-band sequence to obtain enhanced spectrum coefficients; Performing phase reconstruction on the enhanced spectrum coefficients to obtain a complex frequency domain parameter set, and performing frequency band reorganization based on the complex frequency domain parameter set to obtain a spectrum reconstruction matrix; wherein the complex frequency domain parameter set includes a harmonic gain coefficient and a phase offset coefficient; The spectrum reconstruction matrix is ​​feature mapped by a preset acoustic parameter extraction model to obtain an acoustic feature set, and the acoustic feature set is parameter sorted to obtain the acoustic parameter sequence; wherein the acoustic parameter sequence includes an acoustic feature vector and an acoustic parameter matrix.

5. The audio digital signal processing method based on SOC chip according to claim 1, characterized in that: The performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence includes: Expanding the acoustic parameter sequence in the time domain to obtain a parameter time series matrix, and performing group delay estimation based on the parameter time series matrix to obtain a delay feature vector; wherein the delay feature vector includes a phase delay coefficient and a group delay gradient; Decomposing the delay characteristic vector by a preset full-pole analysis technique to obtain a delay component set, and performing phase unwrapping on the delay component set to obtain a phase compensation matrix; Performing polynomial fitting on the phase compensation matrix to obtain a compensation coefficient sequence, and performing group delay correction on the compensation coefficient sequence to obtain a correction parameter set; Performing frequency band mapping on the correction parameter set by sub-band decomposition to obtain a frequency band correction matrix, and performing delay compensation on the frequency band correction matrix to obtain a compensated spectrum set; Performing harmonic reconstruction on the compensation spectrum set to obtain reconstructed spectrum coefficients, and performing spectrum synthesis based on the reconstructed spectrum coefficients to obtain a synthesized spectrum matrix; The synthetic spectrum matrix is ​​calibrated in the spectrum domain by spectrum shaping technology to obtain a calibration spectrum sequence, and phase alignment is performed based on the calibration spectrum sequence to obtain the calibrated spectrum sequence; wherein the calibrated spectrum sequence includes a calibration amplitude spectrum and a calibration phase spectrum.

6. The audio digital signal processing method based on SOC chip according to claim 1, characterized in that: The step of performing reverse Fourier reconstruction on the calibrated spectrum sequence based on the segmented fast Fourier algorithm in the preset SOC chip to obtain an enhanced audio digital signal includes: Performing complex spectrum separation on the calibrated spectrum sequence to obtain a complex spectrum feature matrix, and performing time domain mapping on the complex spectrum feature matrix to obtain a preliminarily reconstructed time domain signal; Filtering the preliminarily reconstructed time domain signal through a preset time domain filter to obtain a denoised signal matrix, and performing amplitude equalization on the denoised signal matrix to obtain an equalized time domain signal; Performing dynamic range correction on the equalized time domain signal to obtain a dynamic correction signal, and performing peak constraint processing on the dynamic correction signal to obtain a constrained time domain signal; wherein the dynamic correction signal includes a dynamic range coefficient and a signal peak distribution; Performing timing alignment on the constrained time domain signal by using a timing adjustment technology to obtain an aligned signal sequence, and performing short-time energy modulation on the aligned signal sequence to obtain a modulated signal matrix; Performing harmonic enhancement on the modulated signal matrix to obtain an enhanced signal set, and performing spectrum interpolation on the enhanced signal set to obtain an interpolated signal matrix; Through the preset segmented fast Fourier algorithm in the SOC chip, the calibrated spectrum sequence is subjected to inverse Fourier transform based on the interpolation signal matrix to obtain an enhanced audio digital signal.

7. An audio digital signal processing device based on a SOC chip, characterized in that: include: The acquisition module is used to acquire audio digital signals from the target scene through a preset multi-channel sensor to obtain a multi-dimensional audio data stream; A decomposition module, used for performing multi-resolution spectrum decomposition on the multi-dimensional audio data stream to obtain hierarchical spectrum features; A transformation module, used for performing complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data; A reconstruction module, used for reconstructing the complex spectrum characterization data in the frequency domain by using a sub-band energy analysis technique to obtain an acoustic parameter sequence; A compensation module, used for performing group delay compensation on the acoustic parameter sequence to obtain a calibrated spectrum sequence; A reconstruction module, used for performing inverse Fourier reconstruction on the calibrated spectrum sequence based on a segmented fast Fourier algorithm preset in a SOC chip to obtain an enhanced audio digital signal; The performing complex domain transformation on the hierarchical spectrum features to obtain complex spectrum representation data includes: Performing time-frequency domain mapping on the hierarchical spectrum features to obtain a time-frequency characterization vector, and performing phase demodulation on the time-frequency characterization vector to obtain phase modulation data; wherein the phase modulation data includes amplitude envelope data and instantaneous phase data; The phase modulation data is orthogonally decomposed by a preset Hilbert filtering technique to obtain an orthogonal component sequence, and the orthogonal component sequence is polar mapped to obtain a polar coordinate parameter set; wherein the polar coordinate parameter set includes an amplitude radius and a phase angle; Performing phase expansion on the polar coordinate parameter set to obtain a phase expansion coefficient matrix, and performing harmonic reconstruction based on the phase expansion coefficient matrix to obtain a harmonic decomposition matrix; The harmonic decomposition matrix is ​​mapped to the frequency domain by a preset complex exponential transformation technology to obtain a complex exponential sequence, and complex domain synthesis is performed based on the complex exponential sequence to obtain complex frequency domain coefficients; wherein the complex frequency domain coefficients include an amplitude modulation component and a phase modulation component; Performing frequency response analysis on the complex frequency domain coefficients to obtain a frequency response characteristic matrix, and performing phase compensation on the frequency response characteristic matrix to obtain a compensation parameter set; The complex frequency domain coefficients are reconstructed in a complex domain based on the compensation parameter set to obtain the complex frequency spectrum characterization data; wherein the complex frequency spectrum characterization data includes a frequency domain amplitude spectrum and a frequency domain phase spectrum.

8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Sound field optimization method and device of multichannel audio system, equipment and storage medium

    CN119383526A

  • Audio processing method and system for speech spectrum reconstruction

    CN119541475A