Ultra-low power consumption end side AI audio signal processing method and system

By adopting ultra-low-power end-side AI audio signal processing methods on resource-constrained terminal devices, using technologies such as frame division, logarithmic compression, critical band analysis and variable Bayes fixed-point operation, the challenge of extracting second-order cyclic stationary signals is solved, efficient real-time audio processing and signal separation is achieved, and audio quality and system robustness are improved.

CN120220714AInactive Publication Date: 2025-06-27SHENZHEN ULTRA EASY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510687140.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In resource-constrained terminal devices, extracting second-order cyclic stationary signal components from noise environments is a major challenge. Traditional audio processing methods are highly complex in computing and difficult to achieve real-time processing.

Method used

The ultra-low power side AI audio signal processing method is adopted to frame, logarithmic compression and decompose the input audio signal, extract the modulation periodic signal and non-modulated signal, and use critical band analysis, variable Bayes fixed-point operation and radial basis support vector machine for processing, to construct an audio processing model to achieve accurate signal separation and processing.

Benefits of technology

It significantly reduces computing complexity, enables efficient operation on resource-constrained end-side devices, reduces storage space usage, ensures real-time processing performance, and improves audio quality and system robustness through precise signal separation and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220714A_ABST
    Figure CN120220714A_ABST
Patent Text Reader

Abstract

The invention relates to an ultra-low power consumption end side AI audio signal processing method and system. The method comprises the steps of performing framing and logarithmic compression processing on an input audio signal to obtain target spectrum data, and decomposing the target spectrum data to obtain a modulation period signal and a non-modulation signal; performing critical frequency band analysis and variational Bayesian fixed-point operation on the non-modulated signal to obtain a quantized noise parameter; extracting a main period characteristic of the modulation period signal and a frequency band characteristic of the quantization noise parameter, and constructing a target discrimination characteristic according to the main period characteristic and the frequency band characteristic; and inputting the target discrimination feature into a radial basis support vector machine, and training by adopting a piecewise linear kernel function and a sequence minimum optimization algorithm of second-order Taylor expansion to obtain an audio processing model. According to the invention, accurate separation of modulation period signals and non-modulation signals is realized, and efficient utilization of system resources and accurate control of power consumption are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio signal processing, and particularly relates to an ultra-low power end-side AI audio signal processing method and system. Background Art

[0002] With the popularization of end-side devices such as smart headphones and smart speakers, end-side audio processing technology has been increasingly widely used in daily life. These devices usually need to perform real-time processing of audio signals in resource-constrained environments, including functions such as noise cancellation and audio enhancement. However, traditional audio processing methods often have high computational complexity and are difficult to achieve real-time processing on end-side devices.

[0003] In actual application scenarios, audio signals usually contain a mixture of human voices, ambient sounds, and background noise. Among them, ambient sounds with periodic modulation characteristics can be modeled as second-order cyclostationary signals. Although a large amount of research has been devoted to extracting the human voice component from a single-channel microphone, it is still a major challenge to extract second-order cyclostationary signal components from a noisy environment on resource-constrained terminal devices. Summary of the Invention

[0004] The main object of the present invention is to provide an ultra-low power end-side AI audio signal processing method and system, which realizes the precise separation of modulated periodic signals and non-modulated signals, and realizes the efficient utilization of system resources and the precise control of power consumption.

[0005] To achieve the above object, the present invention provides an ultra-low power end-side AI audio signal processing method, including the following steps: Perform frame division and logarithmic compression processing on the input audio signal to obtain target spectral data, and decompose the target spectral data to obtain a modulated periodic signal and a non-modulated signal; Perform critical band analysis and variational Bayesian fixed-point operation on the non-modulated signal to obtain quantization noise parameters; Extract the main period feature of the modulated periodic signal and the band feature of the quantization noise parameters, and construct a target discriminant feature according to the main period feature and the band feature; Input the target discriminant feature into a radial basis support vector machine, and use a piecewise linear kernel function and a sequential minimal optimization algorithm of second-order Taylor expansion for training to obtain an audio processing model.

[0006] The present invention also provides an ultra-low power end-side AI audio signal processing system, including: A decomposition unit for performing frame division and logarithmic compression processing on the input audio signal to obtain target spectral data, and decomposing the target spectral data to obtain a modulated periodic signal and a non-modulated signal; An arithmetic unit for performing critical band analysis and variational Bayesian fixed-point arithmetic on the non-modulated signal to obtain quantization noise parameters; A construction unit for extracting the main period feature of the modulated periodic signal and the band feature of the quantization noise parameters, and constructing a target discrimination feature according to the main period feature and the band feature; A training unit for inputting the target discrimination feature into a radial basis support vector machine, and training by using a piecewise linear kernel function and a sequential minimal optimization algorithm with second-order Taylor expansion to obtain an audio processing model.

[0007] In summary, the technical solution provided by the present invention significantly reduces the computational complexity by adopting techniques such as fast Fourier transform with fixed-point arithmetic, replacing exponential operations with piecewise linear kernel functions, and simplifying calculations with second-order Taylor expansion, enabling the algorithm to operate efficiently on resource-constrained edge devices. By methods such as performing K-means clustering on the band energy spectrum matrix to simplify the model, performing fixed-point quantization on the model parameters, and implementing probability density calculation by means of look-up tables, the storage space occupancy is greatly reduced. Based on the design of a three-core pipeline architecture and a double-buffer FIFO queue, parallel processing of signal preprocessing, model inference, and signal synthesis is achieved, ensuring real-time processing performance. Through real-time monitoring of the CPU occupancy rate and memory usage rate, combined with dynamic task allocation and core frequency adjustment strategies, efficient utilization of system resources and precise control of power consumption are achieved. By using an improved linearly decaying ICEEMDAN algorithm for signal decomposition and combining variational Bayesian fixed-point arithmetic for parameter estimation, precise separation of modulated periodic signals and non-modulated signals is achieved. Through techniques such as automatic gain adjustment of the frequency band and overlapping add smoothing processing, the robustness and reliability of the system in the actual application environment are improved. Description of the Drawings

[0008] Figure 1 is a schematic diagram of the steps of an ultra-low power consumption edge AI audio signal processing method according to an embodiment of the present invention; Figure 2 is a block diagram of the structure of an ultra-low power consumption edge AI audio signal processing system according to an embodiment of the present invention.

[0009] The implementation, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0010] In order to make the object, technical solution, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0011] Refer to Figure 1, this embodiment provides an ultra-low power consumption end-side AI audio signal processing method, including the following steps: S1. Perform frame division and logarithmic compression processing on the input audio signal to obtain target spectrum data, and decompose the target spectrum data to obtain a modulated periodic signal and an unmodulated signal; Among them, when performing frame division processing on the input audio signal, the continuous audio signal is divided into multiple time windows, and a certain overlap length is set between adjacent frames to ensure smooth transition of inter-frame information, reduce information loss, and obtain a frame sequence of the audio signal. Subtract the mean value of each frame signal in the audio signal frame sequence from the frame signal to eliminate the DC component and improve the accuracy of subsequent signal processing. On this basis, in order to enhance the high-frequency signal components and suppress the low-frequency components, pre-emphasis processing is performed through a first-order finite impulse response (FIR) filter to obtain a preprocessed signal frame. Apply a Hamming window function to the preprocessed signal frame, thereby reducing the impact of spectral leakage on the analysis accuracy when performing spectral analysis. Perform a fast Fourier transform on the windowed time-domain signal to convert the signal from the time domain to the frequency domain and obtain frequency-domain spectrum data. Extract spectral information from the frequency-domain spectrum data, including calculating the amplitude distribution of the spectrum, that is, obtaining the amplitude spectrum by calculating the modulus value of the complex spectrum, to obtain the initial spectrum data. In this process, since the frequency-domain spectrum data has a large dynamic range, dynamic range compression processing is performed on it to improve the distinguishability of weak signals and reduce the problem of excessive amplitude differences. To achieve dynamic range compression, logarithmic operations are used to transform the amplitude spectrum to obtain a logarithmically compressed spectrum. Segment the logarithmically compressed spectrum and normalize the band energies of each segment to reduce the problem of uneven energy of the spectrum in different frequency regions and obtain the target spectrum data. Decompose the target spectrum data to distinguish different audio components. Adopt a linear attenuation decomposition method to decompose the target spectrum data into a modulated periodic signal and an unmodulated signal, where the modulated periodic signal corresponds to a sound signal with a certain periodicity, such as a voice or music signal, and the unmodulated signal mainly corresponds to non-periodic signals such as environmental noise and transient noise. The extraction of the modulated periodic signal is based on a periodic analysis method, such as the autocorrelation function or the Hilbert transform, in order to identify the main periodic components in the signal, and the decomposition of the unmodulated signal is based on a noise modeling method, and by analyzing the energy distribution characteristics of the signal in different frequency bands, the components related to the background noise are extracted.

[0012] Normalize the target spectral data to ensure the numerical stability of subsequent calculations, reduce the energy differences between different signal components, and improve the decomposition accuracy. Generate N Gaussian white noise sequences for constructing noise-assisted signals in the subsequent linear attenuation decomposition process. The role of the Gaussian white noise sequences is to provide a random perturbation benchmark, enabling the subsequent signal decomposition to more effectively capture the intrinsic mode components of the signal. To adjust the noise intensity, each Gaussian white noise sequence is multiplied by an initial noise intensity coefficient to obtain a group of initial noise sequences. The setting of this initial noise intensity coefficient is related to the dynamic range of the target signal to ensure that the impact of the noise on signal decomposition is within a controllable range. According to the current decomposition layer, perform a linear attenuation operation on each noise sequence in the group of initial noise sequences, that is, gradually reduce the noise intensity coefficient at each layer of decomposition to control the impact range of the noise on signal decomposition and generate linearly attenuated noise sequences. Superimpose the target spectral data and the linearly attenuated noise sequences respectively to construct a signal containing noise perturbations, and perform the Hilbert transform on it to obtain the analytical form of the signal. The Hilbert transform can expand the original signal into a complex form, extract the amplitude envelope and instantaneous frequency information of the signal, and facilitate the calculation of the local mean of the signal. Calculate the local mean of each analytical signal and extract the modal components from the signal. These modal components represent the characteristic information of the signal at different time scales. Perform zero-mean processing on the modal components, that is, zero the mean of each modal component, and iteratively perform the superposition of the linearly attenuated noise sequences and the Hilbert transform to continuously decompose the signal until the standard deviation of the modal components is less than a preset threshold. At this time, the signal is considered to have been decomposed into stable intrinsic mode functions. The intrinsic mode functions represent the inherent oscillation modes of the signal in different frequency ranges. Calculate the autocorrelation coefficient of each intrinsic mode function to analyze its periodic characteristics. The calculation of the autocorrelation coefficient reveals whether the signal has a significant periodic pattern, and signals with strong periodicity correspond to audio components with regular changes such as speech and music. Set a coefficient threshold. When the autocorrelation coefficient of a certain modal component is greater than this threshold, classify it as a modulated periodic signal. These signals contain the pitch information, formants, and other periodic acoustic characteristics of speech. For the remaining modal components, that is, those with autocorrelation coefficients below the threshold, classify them as non-modulated signals. These signals contain background noise, ambient sounds, and other non-periodic audio components.

[0013] S2. Perform critical band analysis and variational Bayesian fixed-point operation on the non-modulated signal to obtain quantization noise parameters; Specifically, for non-modulated signals, Mel-frequency scaling is performed. The human ear's perception ability for different frequencies is not linearly distributed but is more sensitive to low-frequency signals. Therefore, when performing the division, a non-uniformly distributed critical band division is carried out according to the Mel scale, so that the center frequency of each critical band conforms to the human ear's auditory characteristics and better adapts to the signal processing method of the natural auditory system. After the division, the signals in each critical band are filtered to obtain a group of critical band signals, ensuring that the signals are independently processed within their respective frequency ranges and avoiding unnecessary band interference. The short-time Fourier transform is performed on each band signal in the group of critical band signals to convert it into a frequency-domain representation, and the corresponding band energy spectrum matrix is obtained, which describes the energy change of each critical band at different time slices. The K-means clustering with fixed-point operations is performed on the band energy spectrum matrix. Through K-means clustering, the complex energy spectrum information is simplified into a group of representative band models, thereby reducing the computational complexity. The high-order statistics analysis is performed on the simplified band models, and the fourth-order moment statistics of each critical band are calculated to effectively describe the kurtosis and skewness of the signal distribution. The fourth-order moment statistics are used as the initial parameters of the gamma distribution, and the calculation results of the fourth-order moment are used to estimate the shape parameter and scale parameter of the gamma distribution. The parameter estimation process is optimized by the maximum likelihood estimation method with fixed-point operations to obtain the band distribution parameters. The band distribution parameters are input into the variational Bayesian inference model to optimize the band parameters. The method of diagonal approximation is used to reduce the computational complexity. The covariance matrix of diagonal approximation simplifies the calculation process of variational inference, enabling efficient execution of Bayesian inference on edge devices, thereby optimizing the band parameters under limited computing resources. The optimized band parameters are quantized with a fixed number of bits to meet the low-power computing requirements of edge devices. The core of fixed-point quantization is to map continuous numerical values to a finite set of discrete numerical values to reduce storage and computational overhead. On this basis, the probability density calculation of fixed-point operations is implemented by means of a look-up table. The look-up table method pre-stores the complex calculation process as a look-up table, so that only index look-up is required during actual calculation, thus significantly reducing the calculation time and finally obtaining the quantization noise parameters.

[0014] S3. Extract the main period characteristics of the modulated periodic signal and the band characteristics of the quantization noise parameters, and construct the target discriminant characteristics according to the main period characteristics and the band characteristics; It should be noted that the Hilbert transform is performed on the modulated periodic signal to obtain the analytic representation of the signal. The purpose of the Hilbert transform is to separate the instantaneous amplitude and instantaneous phase of the signal, so as to calculate the instantaneous frequency sequence. The instantaneous frequency reflects the dynamic change characteristics of the signal on the time axis. Local extreme point analysis is performed on the instantaneous frequency, that is, the local maxima in the instantaneous frequency sequence are detected, and the time intervals between adjacent maximum points are statistically counted to obtain the main period characteristics. Frequency domain analysis is performed on the main period characteristics to calculate the frequency domain harmonic coefficients. The calculation method of the harmonic coefficients is to perform a Fourier transform on the main period characteristics and extract the amplitude and phase information of the first M harmonic components from the transformation result. Since the computing resources of the edge device are limited, 16-bit fixed-point numbers are used to represent the harmonic characteristics when storing them, so as to reduce the storage overhead and improve the computing efficiency. Fixed-point number representation effectively reduces the complexity of floating-point calculations, thus adapting to the computing requirements of ultra-low power consumption. The obtained harmonic feature vector contains the energy information and phase characteristics of the harmonic components, thus describing the harmonic structure of the modulated periodic signal. At the same time, the shape parameter and scale parameter of the gamma distribution of each frequency band are processed based on the quantization noise parameter. The differential coding method is used to compress the shape parameter and scale parameter, so that the parameter changes between adjacent frequency bands are represented more compactly. Differential coding is an effective data compression method that can reduce the storage requirements and computing complexity. In order to extract the frequency band characteristics, the cross-correlation coefficient of adjacent frequency band parameters is calculated, and the statistical dependence relationship between different frequency bands is analyzed to capture the correlation characteristics of the noise signal in different frequency regions. Through the calculation of the cross-correlation coefficient, a frequency band feature vector is established to characterize the distribution pattern of the noise signal. The harmonic feature vector and the frequency band feature are concatenated, and a feature selection method is used to reduce the feature dimension and improve the computing efficiency. The Lasso algorithm of the coordinate descent method is used for feature selection, and the sparsity threshold is set to 0.85. The Lasso algorithm sparsifies the feature weights through regularization constraints, so as to screen out the most discriminative feature components and obtain the target discriminative features.

[0015] S4. Input the target discriminative features into a radial basis support vector machine, and use the piecewise linear kernel function and the sequential minimal optimization algorithm with second-order Taylor expansion for training to obtain an audio processing model.

[0016] Specifically, the target discriminant features are input into a radial basis support vector machine, and the support vector machine measures the non-linear mapping of the input features based on the radial basis kernel function. In the radial basis support vector machine, by kernelizing the input data, the data is mapped from the original space to a high-dimensional space, making the data linearly separable. To optimize the computational efficiency, a kernel matrix is constructed based on the radial basis kernel function, and the exponential operation interval is divided into three segments. Within each segment, linear functions with different slopes are used for fitting to form a piecewise linear kernel function. The Hessian matrix is constructed using the calculation results of the piecewise linear kernel function. The Hessian matrix is a core element in quadratic optimization problems, describing the second-order derivative information of the objective function and helping the optimization algorithm to judge the convergence of the optimal solution. After constructing the Hessian matrix, based on the Taylor formula, it is expanded to the second-order term at the current iteration point for more accurate approximate calculations. And by performing Cholesky decomposition on the Hessian matrix, it is transformed into an upper triangular matrix form, thus simplifying the subsequent calculation steps and avoiding the high computational cost of directly solving the matrix inverse. The Hessian matrix obtained by expansion constitutes a quadratic programming problem, and an objective function for sequential minimal optimization is constructed based on this problem, so as to gradually approximate the optimal solution during the optimization process to reduce the computational amount and improve the accuracy. At initialization, the Lagrange multipliers are reasonably set to provide an initial solution for the optimization process. The Lagrange multipliers play a role in constraining the optimization problem during the training process of the support vector machine. By reasonably initializing the Lagrange multipliers, it is ensured that the algorithm quickly converges to an effective solution space. At initialization, the update direction of the Lagrange multipliers is calculated based on the current solution, and iterative optimization is performed along this direction. By continuously adjusting the Lagrange multipliers, the optimization problem of the support vector machine gradually approaches the optimal solution, and the weight coefficients of the support vectors are obtained. The weight coefficients are sparsified to reduce unnecessary calculations. By setting the unimportant or nearly zero parts of the weight coefficients to zero, the computational complexity of the model is reduced. The sparsified weight coefficients are used to calculate the piecewise linear kernel function through a look-up table method, thereby improving the computational efficiency and reducing the power consumption, and an audio processing model is obtained.

[0017] Deploy the audio processing model onto a triple-core processor to build an efficient pipeline processing architecture. The first core is responsible for signal preprocessing and feature extraction, including frame segmentation, logarithmic compression, spectrum analysis of the audio signal, and calculation of target discriminant features. The second core performs the core model inference task, that is, classification based on the radial basis support vector machine and piecewise linear kernel function, and calculates the category of the audio signal using the optimized weight coefficients, providing a discriminant basis for subsequent signal enhancement. At the same time, the third core is responsible for signal synthesis. After obtaining the classification result, it reconstructs and enhances the signal, making the final output audio have better quality and clarity. To efficiently transfer data in the multi-core architecture, a double-buffered FIFO queue is used for data exchange between each processing core, reducing processing latency while ensuring smooth data transfer, making the operation of the entire system more real-time and stable, and forming an efficient pipeline processing architecture. In the multi-core system, the operating state of each processing core is monitored in real time to ensure that computing resources are fully utilized. When the CPU occupancy rate exceeds the set first target value, the system automatically detects whether there are idle cores and allocates part of the task load to the underutilized cores to achieve the purpose of dynamically balancing the computing load. In terms of memory management, to avoid excessive data occupying storage space, when the memory usage rate exceeds the set second target value, the intermediate data is automatically compressed and stored to reduce storage pressure and improve data access efficiency. Through an intelligent scheduling method, the entire system operates efficiently and stably under limited computing resources, forming an effective resource scheduling strategy. Based on the resource scheduling strategy, the operating frequencies of the three processing cores are dynamically adjusted to optimize the balance between energy consumption and performance. When the system detects a low task load, it reduces the core frequency to reduce power consumption; when the computing demand is high, it increases the core frequency to ensure the real-time nature of the task and computing efficiency. By dynamically adjusting the operating frequency, it automatically adapts to different computing demands, thereby improving the overall energy efficiency ratio, enabling end-side AI audio signal processing to meet real-time requirements while maintaining low-power characteristics, and being applicable to various low-power embedded platforms. During the processing of the audio signal, based on the trained audio processing model, the input audio signal is separated into modulated periodic signals and non-modulated signals. Since the modulated periodic signals contain important audio information such as speech or music, and the non-modulated signals correspond to environmental noise or background interference, noise suppression is performed on the non-modulated signals to improve the audio quality. The Wiener filter is used to process the non-modulated signals. Wiener filtering is based on the optimal linear estimation theory, effectively suppressing noise and retaining the main features of the speech signal. After filtering, the first enhanced signal is obtained. The first enhanced signal is sub-band synthesized with the modulated periodic signals to optimize the audio quality. Based on the quantization noise parameters, the gain coefficients of each frequency band are automatically adjusted so that the signal enhancement degrees of different frequency bands conform to the noise characteristics, thereby enhancing speech clarity while avoiding distortion introduced by over-enhancement.Signal reconstruction is completed by the overlap - add method, making the signal maintain a smooth connection in the time domain. At the same time, it is ensured that the enhanced audio signal has good continuity in both the time and frequency dimensions, obtaining a second enhanced signal. The second enhanced signal is smoothed through overlap - add processing, effectively reducing the boundary effect caused by frame - splitting processing, so that the synthesized signal will not produce unnecessary breaks or uneven listening sensations during continuous playback. De - emphasis filtering is performed on the second enhanced signal to restore its original frequency response characteristics, thereby ensuring that the final audio output not only has a high - quality enhancement effect but also can maintain a natural listening sensation, finally obtaining an enhanced audio signal.

[0018] In one example, the input audio signal is framed and logarithmically compressed to obtain target spectral data, and the target spectral data is decomposed to obtain modulated periodic signals and non - modulated signals, including: The input audio signal is framed, and the overlapping length of adjacent frames is set to obtain a sequence of audio signal frames; For each frame signal in the sequence of audio signal frames, the mean value of the frame signal is subtracted, and pre - emphasis processing is performed through a first - order FIR filter to obtain pre - processed signal frames; The Hamming window function is applied to the pre - processed signal frames, and the fast Fourier transform is performed to convert the time - domain signal into a frequency - domain signal, obtaining frequency - domain spectral data; Spectral information is extracted from the frequency - domain spectral data, and the magnitude spectrum is calculated to obtain initial spectral data; Dynamic range compression operation is performed on the initial spectral data to convert the magnitude spectrum into a logarithmic spectrum, obtaining logarithmically compressed spectra. The logarithmically compressed spectra are segmented and band - energy normalized to obtain target spectral data; Linear attenuation decomposition is performed on the target spectral data to obtain modulated periodic signals and non - modulated signals.

[0019] In this example, the input audio signal is framed. The audio signal is a continuous time series, which is divided into several small time periods, and each time period is called a "frame". When framing, the overlapping length between adjacent frames is set to reduce the boundary effect between frames. Setting an appropriate overlapping length ensures sufficient similarity between adjacent frames and avoids distortion in the analysis of the audio signal due to sudden changes between frames. Assume that the length of each frame is , and the overlapping length is , then the time length of each frame signal is , where is the sampling rate, the overlapping length is , then the time interval between two adjacent frames is . Through this step, the continuous audio signal is converted into a series of frames of fixed length, and these frames form a sequence of audio signal frames. Preprocessing is performed on each frame of the signal, including removing the mean value of each frame of the signal to eliminate the DC component in the signal, so that subsequent processing focuses on the changing part of the signal. In this step, assume that the signal of the th frame is (where is the time index), and the mean value of this frame of signal is , then the signal after removing the mean value of each frame of signal is: After completing the removal of the mean value, a first-order HR filter is applied to each frame of the signal for pre-emphasis processing to increase the energy of the high-frequency part in the signal, thereby improving the spectral characteristics of the signal. Especially in the processing of human voice or speech signals, it helps to enhance the high-frequency components. The transfer function of the pre-emphasis filter is , where is the pre-emphasis coefficient, and its value is 0.95 or 0.97. The signal processed by this filter is represented by the following formula: The processed signal is the pre-emphasized signal frame. After completing the pre-emphasis, the Hamming window function is applied to each frame of the signal to reduce the discontinuity of the signal at the frame boundary, thereby reducing spectral leakage. The definition of the Hamming window function is: where is the length of the window, is the index of the sample. After each frame of the signal is weighted by the Hamming window function, its weighted signal is: The fast Fourier transform is performed on the weighted signal to convert the time-domain signal into a frequency-domain signal. The fast Fourier transform is an efficient method for calculating the discrete Fourier transform, and its calculation formula is: where, is the th frequency point of the frequency-domain signal, is the frame length, is the sample index of the time-domain signal, is the frequency-point index of the frequency-domain signal, is the imaginary unit. The frequency-domain signal obtained through the fast Fourier transform is the frequency-domain spectrum data. The spectral information is extracted from the frequency-domain spectrum data, and the amplitude spectrum of the signal is concerned. The amplitude spectrum is obtained by calculating the modulus of the frequency-domain signal : where and are the real part and the imaginary part of respectively. After obtaining the amplitude spectrum, the initial spectral data is constructed, which is used for subsequent processing steps. Dynamic range compression operation is performed to avoid excessive amplitude changes of the signal. Especially in the case of strong noise, the peak value of the signal may exceed the dynamic range of the processor. The dynamic range compression adopts logarithmic transformation, and the amplitude spectrum is converted to the logarithmic spectrum to compress the dynamic range: where is a small constant to prevent the occurrence of zero values in logarithmic operations. The spectrum after logarithmic compression is the logarithmically compressed spectrum. After logarithmic compression of the spectrum, segment processing and band energy normalization are performed to improve the comparability between bands and make it suitable for subsequent signal decomposition. The purpose of band energy normalization is to adjust the energy of each band so that it is maintained on a unified scale, thereby eliminating the amplitude differences between different bands. Assuming the target spectral data is , its normalization process is carried out through the following formula: The obtained normalized spectral data is the target spectral data. Linear attenuation decomposition is performed to decompose the target spectral data into a modulated periodic signal and an unmodulated signal. This process is achieved by introducing an attenuation factor. N Gaussian white noise sequences are generated and multiplied by the target spectral data to obtain the initial noise sequences. In each decomposition, the intensity of the noise gradually decays until the difference between the decomposed signal and the original signal meets certain conditions. After linear attenuation and noise suppression, the modulated periodic signal and the unmodulated signal are obtained.

[0020] In an example, linear attenuation decomposition is performed on the target spectral data to obtain a modulated periodic signal and an unmodulated signal, including: Normalize the target spectral data, generate N Gaussian white noise sequences, and multiply each Gaussian white noise sequence by the initial noise intensity coefficient to obtain a group of initial noise sequences; For each noise sequence in the group of initial noise sequences, according to the current decomposition layer number, attenuate the noise intensity coefficient to obtain a linearly attenuated noise sequence; Superimpose the target spectral data and the linearly attenuated noise sequences respectively, perform the Hilbert transform to obtain the analytic signal, and calculate the local mean of each analytic signal to obtain the modal components; Perform zero-mean processing on the modal components, and iteratively execute the superposition of the linearly decaying noise sequence and the Hilbert transform until the standard deviation of the modal components is less than the preset threshold to obtain the intrinsic mode function; Calculate the autocorrelation coefficient of the intrinsic mode function, and classify the modal components with autocorrelation coefficients greater than the coefficient threshold as modulated periodic signals, and the remaining modal components as non-modulated signals.

[0021] In this example, normalize the target spectral data to ensure that the numerical range of the data is in a stable interval, so that subsequent calculations will not cause numerical overflow or affect the stability of the algorithm due to excessive amplitude differences. Assume that the target spectral data is , where represents the frequency index, then the normalization is expressed as: where, represents the normalized target spectral data, and represent the minimum and maximum values of the target spectral data respectively. After normalization, generate Gaussian white noise sequences to provide random perturbations, so that the subsequent decomposition algorithm can effectively separate different modal components. The Gaussian white noise is expressed as: where, represents the th Gaussian white noise sequence, with a mean of 0 and a variance of 1. These noise sequences are used in the subsequent decomposition process to simulate different modal characteristics in the signal by adjusting their intensities. After generating the noise sequences, multiply each Gaussian white noise sequence by the initial noise intensity coefficient to obtain the initial noise sequence group. The initial noise sequence is expressed as: where, represents the initial noise sequence, is the initial parameter used to control the noise intensity, which is set according to the energy size of the target signal to ensure that the noise will not cause excessive damage to the signal. According to the current decomposition level , linearly decay each noise sequence to control the noise influence degree at different levels. Set the noise attenuation coefficient to ensure that as the decomposition level increases, the noise intensity gradually decreases, and its attenuation form is expressed as: where, is the parameter controlling the attenuation speed, taking a small positive number to ensure a gentle noise attenuation. When increases, Gradually decrease, making the subsequent decomposition process more accurately extract weak signal components. Correspondingly, the linearly decaying noise sequence is expressed as: After completing the linear decay, the target spectral data is superimposed with the linearly decaying noise sequence to form a new mixed signal: Perform the Hilbert transform on the mixed signal to obtain the analytic form of the signal. The Hilbert transform is used to calculate the envelope and instantaneous frequency of the signal, and its mathematical expression is: where is the analytic signal, represents the Hilbert transform of the signal The local mean of the analytic signal is calculated by the following formula: where represents the local mean, is the window length, indicating the smoothing characteristics of the signal calculated within the local region. Perform zero-mean processing on the modal component to eliminate the global trend, making the signal more suitable for subsequent analysis: where is the total number of frequency indices. Continue to iteratively perform the superposition of the linearly decaying noise sequence and the Hilbert transform until the standard deviation of the modal component is less than the preset threshold : After meeting this condition, the obtained signal component is the intrinsic mode function (IMF). These intrinsic mode functions are used for the periodic analysis of the signal. To classify the decomposed intrinsic mode functions into modulated periodic signals or non-modulated signals, calculate the autocorrelation coefficient of each intrinsic mode function. The calculation formula of the autocorrelation coefficient is as follows: where is the autocorrelation coefficient at the lag , represents the th intrinsic mode function. If the maximum autocorrelation coefficient of a certain intrinsic mode function exceeds the set coefficient threshold , then it is considered that the signal has periodicity and is classified as a modulated periodic signal: Otherwise, the modal component is classified as an unmodulated signal.

[0022] In one example, critical band analysis and variational Bayesian fixed-point operations are performed on the unmodulated signal to obtain quantization noise parameters, including: The unmodulated signal is divided into multiple critical bands according to the Mel frequency scale. The center frequencies of each critical band are non-uniformly distributed according to the human auditory characteristics, and the signals within each critical band are filtered to obtain a group of critical band signals; Perform short-time Fourier transform on each band signal in the group of critical band signals to obtain a band energy spectrum matrix, and perform K-means clustering with fixed-point operations on the band energy spectrum matrix to obtain a simplified band model; Calculate the fourth-order moment statistic for each critical band of the simplified band model, and use the fourth-order moment statistic as the initial values of the shape parameter and scale parameter of the gamma distribution. Obtain the band distribution parameters through maximum likelihood estimation with fixed-point operations; Input the band distribution parameters into the variational Bayesian inference model, and use a covariance matrix with diagonal approximation to reduce the computational complexity to obtain optimized band parameters; Perform fixed-point quantization on the optimized band parameters, and implement the probability density calculation with fixed-point operations through a look-up table method to obtain quantization noise parameters.

[0023] In this example, the unmodulated signal is divided into multiple critical bands according to the Mel frequency scale. The Mel scale is a non-uniform frequency division method based on the human auditory characteristics, with higher resolution in the low-frequency band and lower resolution in the high-frequency band, to better conform to the human ear's perception ability of sound. The Mel frequency and the linear frequency are related by the following formula: where is the Mel frequency, is the actual frequency. By dividing the input unmodulated signal according to the Mel frequency scale, critical bands are obtained, and the center frequency of each band is distributed according to the Mel scale: where is the index of the Mel frequency. Based on this division method, a set of band-pass filters are used to decompose the signal, and the center frequency of each filter corresponds to a Mel frequency point. Assuming the original unmodulated signal is , the signals of each band are extracted through the filter bank : where The filtered signal representing the n-th frequency band, is the impulse response of the corresponding band-pass filter, and the symbol * represents the convolution operation. The short-time Fourier transform is performed on the critical band signal group to obtain the spectral information of each frequency band. The short-time Fourier transform is expressed as: where, represents the n-th frequency point in the m-th frame of the k-th frequency band, is the window function (such as Hamming window), is the window length, where, represents the n-th frequency band in the m-th frame. This energy spectral matrix contains the energy distribution information of the audio signal at different frequencies and times. To reduce the computational complexity, K-means clustering is performed on the energy spectral matrix of frequency bands, so that the frequency bands with similar energy distributions are merged into similar categories, thus forming a simplified frequency band model. The goal of K-means clustering is to minimize the within-cluster sum of squares: is the number of cluster centers, is the set of frequency bands belonging to the k-th cluster, is the center value of this cluster, is the k-th frequency band energy vector. By iteratively updating the cluster centers and reassigning the cluster members in K-means, a simplified frequency band model after clustering is obtained. The fourth-order moment statistic is calculated for the simplified frequency band model to describe the distribution characteristics of the frequency bands. The fourth-order moment includes the mean , variance , skewness and kurtosis : where, is the mean, is the variance, is skewness, representing the asymmetry of the data, is kurtosis, measuring the steepness of the data. These statistics are used to initialize the shape parameters of the gamma distribution and the scale parameters : Through maximum likelihood estimation, for and are optimized to obtain more accurate band distribution parameters. The band distribution parameters are input into the variational Bayesian inference model to optimize these parameters. Since directly calculating the complete covariance matrix has a large computational cost, a diagonal approximation method is used to reduce the computational complexity. The covariance matrix of the diagonal approximation is expressed as: Through Bayesian inference, the posterior probability is calculated and the parameters are updated to make the optimized band parameters more accurate. Fixed-point quantization is performed on the optimized band parameters for efficient calculation on low-power devices. The core of fixed-point quantization is to map continuous numerical values to a finite set of discrete numerical values. For the shape parameter and the scale parameter , the following method is used for fixed-point quantization: where is the number of decimal places of the fixed-point number. The probability density calculation of fixed-point operation is realized by looking up tables: where is the gamma function. The commonly used values are pre-stored by looking up tables to avoid real-time calculation and improve the calculation efficiency, and finally the quantization noise parameters are obtained.

[0024] In an example, the main period characteristics of the modulation period signal and the band characteristics of the quantization noise parameters are extracted, and the target discrimination characteristics are constructed based on the main period characteristics and the band characteristics, including: Perform Hilbert transform on the modulation period signal and extract the instantaneous frequency sequence, and obtain the main period characteristics based on the time interval statistics of adjacent maximum points; Calculate the frequency domain harmonic coefficients for the main period characteristics, extract the amplitude and phase information of the first M harmonic components, and represent them by 16-bit fixed-point numbers to obtain the harmonic feature vector; Differentially encode the shape parameter and scale parameter of the gamma distribution for each frequency band based on the quantization noise parameter, and calculate the cross-correlation coefficient of the adjacent frequency band parameters to obtain the frequency band feature; Concatenate the harmonic feature vector with the frequency band feature, and set the sparsity threshold of 0.85 through the Lasso algorithm of the coordinate descent method for feature selection to obtain the target discriminant feature.

[0025] In this example, perform the Hilbert transform on the modulated periodic signal to extract its instantaneous frequency information. The core of the Hilbert transform lies in converting the real signal into a complex analytic signal , and its mathematical representation is: where is the Hilbert transform of , and its calculation method is: The instantaneous frequency of the analytic signal is calculated from the differential of its phase , that is: where is the phase function of the analytic signal. By calculating the instantaneous frequency sequence and extracting the local maximum points , and statistically calculating the time interval between adjacent maximum points, its mean value is the main period feature: where is the total number of maximum points. Calculate the frequency domain harmonic coefficient for the main period feature. Perform the Fourier transform on the signal to obtain the spectrum : The harmonic component is an integer multiple of the main period frequency . Let the frequency of its th harmonic component be , then the harmonic amplitude is: The phase information is: To reduce the storage and calculation amount, the harmonic feature is represented by 16-bit fixed-point numbers. Set the fixed number of decimal places , then the fixed-point representations of the amplitude and phase are as follows: where and are the fixed-point representations of the harmonic amplitude and phase, respectively. By extracting the amplitude and phase information of the first harmonic components and storing them in the fixed-point format, the harmonic feature vector is obtained: Meanwhile, the frequency band features are calculated based on the quantization noise parameters. The gamma distribution parameters of each frequency band are set , where is the shape parameter, and is the scale parameter. The change amount of adjacent frequency band parameters is calculated using differential coding: To measure the correlation between adjacent frequency bands, the cross-correlation coefficient is calculated: where is the energy value of the th frequency band, is the mean value, and is the cross-correlation coefficient of adjacent frequency bands. The frequency band feature vector is expressed as: The harmonic feature vector is concatenated with the frequency band feature vector to obtain the complete feature vector: Due to the high feature dimension, the most representative features are selected through the feature selection method. The Lasso (Least Absolute Shrinkage and Selection Operator) algorithm is used for feature selection, and its objective function is: where is the feature weight, is the classification label, is the regularization parameter, and is the total number of features. Lasso realizes feature sparsity through regularization, thereby screening out the most discriminative features. In the optimization process, the coordinate descent method is used for solution. The basic step of the coordinate descent method is to fix other variables and only optimize a single variable : where is the soft threshold function: By setting the sparsity threshold to 0.85, that is, only retaining 15% of the feature weights with larger absolute values, the optimized target discriminant features are obtained. .

[0026] In one example, the target discriminant features are input into a radial basis support vector machine, and trained using a piecewise linear kernel function and the sequential minimal optimization algorithm with second-order Taylor expansion to obtain an audio processing model, including: Input the target discriminant features into a radial basis support vector machine, construct a kernel matrix based on the radial basis kernel function, divide the exponential operation interval into three segments, and fit with linear functions with different slopes in each segment to obtain a piecewise linear kernel function; Construct a Hessian matrix for the calculation result of the piecewise linear kernel function, expand the Hessian matrix to the second-order term at the current iteration point based on Taylor's formula, and represent the Hessian matrix in the form of an upper triangular matrix through Cholesky decomposition to obtain a quadratic programming problem; Construct an objective function for sequential minimal optimization according to the quadratic programming problem, initialize the Lagrange multipliers to obtain the initial solution of the optimization problem, and calculate the update direction of the Lagrange multipliers based on the initial solution; Iteratively optimize the Lagrange multipliers along the update direction to obtain the weight coefficients of the support vectors, sparsify the weight coefficients, and implement the calculation of the piecewise linear kernel function through a look-up table method to obtain the audio processing model.

[0027] In this example, the target discriminant features are input into a radial basis support vector machine, and a kernel matrix is constructed based on the radial basis kernel function. The form of the radial basis kernel function is: where represents the kernel function value between the input samples and , is the hyperparameter of the kernel function, controlling the distribution width of the Gaussian kernel, is the square of the Euclidean distance. Since the exponential operation involves non-linear calculations, directly calculating on low-power edge devices will result in high computational complexity. Therefore, the exponential operation is approximately optimized. The exponential operation interval is divided into three segments, and different-slope linear functions are used for fitting in each interval to obtain a piecewise linear kernel function. Assume that the input range of the exponential function is segmented as: where and are the linear fitting parameters in each segment, and It is the interval boundary point. The Hessian matrix is constructed for the calculation result of the piecewise linear kernel function and is defined as the second-order partial derivative matrix for the optimization of quadratic programming problems: Among them, and are the classification labels of the corresponding samples, and the Hessian matrix is used for the optimization calculation of the support vector machine. To optimize the calculation of the Hessian matrix, Taylor expansion is used for second-order approximation and expanded at the current iteration point : Among them, is the first-order derivative of the Hessian matrix, and is the second-order derivative of the Hessian matrix. The second-order expansion method can improve the convergence speed of optimization and reduce unnecessary computational complexity. The Hessian matrix is converted into an upper triangular matrix form through Cholesky decomposition to simplify the solution of the quadratic programming problem: Among them, is a lower triangular matrix, and the solution of the Hessian matrix is efficiently completed through back substitution. After obtaining the quadratic programming problem, the sequential minimal optimization algorithm is used to construct the optimization objective function, and the Lagrange multipliers are gradually optimized to find the optimal solution of the support vector machine, and its optimization objective is: Among them, are the Lagrange multipliers corresponding to each sample. The sequential minimal optimization algorithm makes the optimization problem gradually converge by iteratively optimizing a pair of Lagrange multipliers. Initialize the Lagrange multipliers, that is: In each iteration, update the Lagrange multipliers according to the optimization direction. Set the update direction as , then the iterative update rule is: Among them, is the learning rate. Through multiple iterations, the final support vector weight coefficients are obtained: Sparsify the weight coefficients of the support vectors to reduce the computational complexity and improve the efficiency of the model. The sparsification method is to prune the components in the weights that are close to zero, that is: Among them, is the set sparsity threshold. To improve the calculation efficiency, the piecewise linear kernel function is calculated by using a look-up table method. The look-up table method pre-computes the kernel function values at different distances and stores them in the look-up table. Assuming that the look-up table stores the kernel function value as , the kernel function calculation is completed in the following way: Thereby avoiding real-time calculation of exponential operations, improving the calculation efficiency, and finally obtaining an optimized audio processing model.

[0028] In one example, the ultra-low power edge AI audio signal processing method further includes: Deploy the audio processing model to a triple-core processor. The first core executes signal preprocessing and feature extraction, the second core executes model inference, and the third core executes signal synthesis. Inter-core data transmission is achieved through a double-buffered FIFO queue to obtain a pipelined processing architecture; Monitor the running status of each processing core. When the CPU occupancy rate exceeds the first target value, allocate the task load to the idle core. When the memory occupancy rate exceeds the second target value, compress and store the intermediate data to obtain a resource scheduling strategy; According to the resource scheduling strategy, dynamically adjust the working frequencies of the three processing cores to obtain the working parameters of the processing cores; Based on the classification result of the audio processing model, separate the modulated periodic signal and the non-modulated signal of the input audio signal, and perform noise suppression on the non-modulated signal through a Wiener filter to obtain a first enhanced signal; Perform sub-band synthesis on the first enhanced signal and the modulated periodic signal, automatically adjust the gain coefficient of each frequency band based on the quantization noise parameter, and complete signal reconstruction through the overlap-add method to obtain a second enhanced signal; Smoothly transition the second enhanced signal through overlapping and stacking processing, and perform de-emphasis filtering to compensate for the frequency response to obtain an enhanced audio signal.

[0029] In this example, the processing tasks are reasonably divided so that different computing tasks can be executed in parallel, improving the overall processing efficiency. The first core is used to execute signal preprocessing and feature extraction. This step involves frame division, pre-emphasis, Fourier transform, dynamic range compression, and feature extraction of the audio signal. The input audio signal After frame division processing, the signal length of each frame is set to , and the overlapping length of adjacent frames is set to , then the time interval between adjacent frames is expressed as: Among them, is the sampling rate. Apply the Hamming window function to each frame of the signal To reduce spectral leakage: Convert the signal to the frequency domain through the fast Fourier transform to obtain the spectrum : Perform dynamic range compression on the spectrum data to obtain the logarithmically compressed spectrum through logarithmic transformation : Among them, is a small positive number to prevent numerical overflow in logarithmic operations. The feature extraction module extracts the target discrimination features from the logarithmically compressed spectrum and stores them in a FIFO (First-In-First-Out) queue, ready to be transmitted to the second core. The second core is responsible for model inference, that is, classifying the input features based on the radial basis support vector machine. Load the target discrimination features in the FIFO queue , and calculate its kernel matrix: Among them, controls the width of the kernel function. Calculate the classification decision according to the trained SVM model: Among them, is the Lagrange multiplier, is the label of the training sample, is the bias term. According to the classification result, the audio signal is divided into modulated periodic signals and non-modulated signals, and the processed data is stored in the FIFO queue for the third core to perform signal synthesis. The third core is responsible for signal synthesis. According to the model classification result, the modulated periodic signals and non-modulated signals of the input audio signal are separated. For non-modulated signals, a Wiener filter is used for noise suppression. The optimal estimate of the Wiener filter is: Among them, is the spectrum of the clean signal, is the noise spectrum. After filtering, the first enhanced signal is obtained, and the first enhanced signal is sub-band synthesized with the modulated periodic signal to restore the complete audio signal. Calculate the gain coefficient of each frequency band based on the quantization noise parameter : Through gain adjustment, ensure the balanced distribution of signal energy in different frequency bands. Signal reconstruction uses the overlap-and-add method, that is, perform the inverse Fourier transform on the enhanced signal of each frame and perform inter-frame superposition: Among them, is the frame-level reconstructed signal, is the window function. The obtained second enhanced signal is processed by overlapping and adding to smooth the connection between different frames. To compensate for high-frequency attenuation, perform de-emphasis filtering on the second enhanced signal. The transfer function of the de-emphasis filter is: Among them, is the filtering coefficient, usually taking 0.95 or 0.97. The de-emphasized signal is the final enhanced audio signal . Under the three-core processing architecture, to improve resource utilization, monitor the operating status of the processing cores. When the CPU occupancy rate exceeds the first target value , the task load will be assigned to the idle core. Let the current CPU load be , if , then the system will reallocate the computing tasks between the cores: If is greater than a certain threshold, the task will be dynamically migrated to other cores to balance the computing load. At the same time, to optimize the memory usage, when the memory occupancy rate exceeds the second target value , compress and store the intermediate data. Let the current memory usage be , if , then use Huffman coding or quantization compression method to reduce the data storage requirements: Among them, is the probability distribution of different data blocks, and the compressed data storage requirement is , reducing the storage overhead. According to the above resource scheduling strategy, dynamically adjust the working frequencies of the three processing cores to optimize the balance between power consumption and performance. Assume that the working frequency of the th core is , when the task load is high, the system will increase the frequency: Among them, is the adjustment coefficient. When the load is low, the system will reduce the frequency to reduce power consumption.

[0030] Refer to Figure 2, this embodiment provides an ultra-low power consumption end-side AI audio signal processing system, including: Decomposition unit 1, configured to perform frame division and logarithmic compression processing on the input audio signal to obtain target spectrum data, and decompose the target spectrum data to obtain a modulation period signal and an unmodulated signal; Operation unit 2, configured to perform critical band analysis and variational Bayesian fixed-point operation on the unmodulated signal to obtain quantization noise parameters; Construction unit 3, configured to extract the main period feature of the modulation period signal and the band feature of the quantization noise parameters, and construct a target discrimination feature according to the main period feature and the band feature; Training unit 4, configured to input the target discrimination feature into a radial basis support vector machine, and perform training using a piecewise linear kernel function and a sequential minimal optimization algorithm with second-order Taylor expansion to obtain an audio processing model.

[0031] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the description in the above method embodiment, and details will not be elaborated here.

[0032] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, system, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, system, article or method. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, system, article or method including that element.

[0033] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An ultra-low power consumption end-side AI audio signal processing method, characterized in that, Including the following steps: Perform frame division and logarithmic compression processing on the input audio signal to obtain target spectrum data, and decompose the target spectrum data to obtain a modulated periodic signal and an unmodulated signal; Perform critical band analysis and variational Bayesian fixed-point operation on the unmodulated signal to obtain quantization noise parameters; Extract the main period feature of the modulated periodic signal and the band feature of the quantization noise parameters, and construct a target discrimination feature according to the main period feature and the band feature; Input the target discrimination feature into a radial basis support vector machine, and use a piecewise linear kernel function and a sequential minimal optimization algorithm with second-order Taylor expansion for training to obtain an audio processing model.

2. The ultra-low power consumption end-side AI audio signal processing method according to claim 1, characterized in that, The step of performing frame division and logarithmic compression processing on the input audio signal to obtain target spectrum data, and decomposing the target spectrum data to obtain a modulated periodic signal and an unmodulated signal includes: Perform frame division on the input audio signal and set the overlapping length of adjacent frames to obtain an audio signal frame sequence; Subtract the mean value of each frame signal from each frame signal in the audio signal frame sequence, and perform pre-emphasis processing through a first-order FIR filter to obtain a preprocessed signal frame; Apply a Hamming window function to the preprocessed signal frame and perform a fast Fourier transform to convert the time-domain signal into a frequency-domain signal to obtain frequency-domain spectrum data; Extract spectrum information from the frequency-domain spectrum data and calculate the amplitude spectrum to obtain initial spectrum data; Perform dynamic range compression operation on the initial spectrum data, convert the amplitude spectrum into a logarithmic spectrum to obtain a logarithmically compressed spectrum, perform segmented processing on the logarithmically compressed spectrum and perform band energy normalization to obtain target spectrum data; Perform linear attenuation decomposition on the target spectrum data to obtain a modulated periodic signal and an unmodulated signal.

3. The ultra-low power consumption end-side AI audio signal processing method according to claim 2, wherein The step of performing linear attenuation decomposition on the target spectrum data to obtain a modulated periodic signal and an unmodulated signal includes: Normalize the target spectrum data, generate N Gaussian white noise sequences, and multiply each Gaussian white noise sequence by an initial noise intensity coefficient to obtain an initial noise sequence group; For each noise sequence in the initial noise sequence group, attenuate the noise intensity coefficient according to the current decomposition layer number to obtain a linearly attenuated noise sequence; Superimpose the target spectrum data and the linearly attenuated noise sequence respectively, perform a Hilbert transform to obtain an analytic signal, and calculate the local mean value of each analytic signal to obtain a modal component; Perform zero-mean processing on the modal component, and iteratively perform the superimposition of the linearly attenuated noise sequence and the Hilbert transform until the standard deviation of the modal component is less than a preset threshold to obtain an intrinsic mode function; Calculate the autocorrelation coefficient of the intrinsic mode function, and classify the modal components with autocorrelation coefficients greater than the coefficient threshold as modulated periodic signals, and the remaining modal components as unmodulated signals.

4. The ultra-low power consumption end-side AI audio signal processing method according to claim 1, wherein The step of performing critical band analysis and variational Bayesian fixed-point operation on the unmodulated signal to obtain quantization noise parameters includes: The non-modulated signal is divided into multiple critical bands according to the Mel frequency scale, and the center frequencies of the critical bands are unevenly distributed according to the human auditory characteristics, and the signals in each critical band are filtered to obtain a critical band signal group; Perform short-time Fourier transform on each band signal in the critical band signal group to obtain a band energy spectrum matrix, and perform K-means clustering with fixed-point operations on the band energy spectrum matrix to obtain a simplified band model; Calculate the fourth-order moment statistic for each critical band of the simplified band model, and use the fourth-order moment statistic as the initial values of the shape parameter and scale parameter of the gamma distribution, and obtain the band distribution parameter through maximum likelihood estimation with fixed-point operations; Input the band distribution parameter into the variational Bayesian inference model, and use the diagonal approximation covariance matrix to reduce the computational complexity to obtain the optimized band parameter; Perform fixed-point quantization on the optimized band parameter, and implement the probability density calculation of fixed-point operations through a look-up table method to obtain the quantization noise parameter.

5. The ultra-low power consumption end-side AI audio signal processing method according to claim 1, characterized in that, The main period feature of the modulation period signal and the band feature of the quantization noise parameter are extracted, and the target discrimination feature is constructed according to the main period feature and the band feature, including: Perform Hilbert transform on the modulation period signal and extract the instantaneous frequency sequence, and obtain the main period feature based on the time interval statistics of adjacent maximum points; Calculate the frequency domain harmonic coefficient for the main period feature, extract the amplitude and phase information of the first M harmonic components, and represent them by 16-bit fixed-point numbers to obtain the harmonic feature vector; Based on the quantization noise parameter, perform differential coding on the shape parameter and scale parameter of the gamma distribution of each band, and calculate the cross-correlation coefficient of adjacent band parameters to obtain the band feature; Concatenate the harmonic feature vector and the band feature, and perform feature selection by setting a sparsity threshold of 0.85 through the Lasso algorithm of the coordinate descent method to obtain the target discrimination feature.

6. The ultra-low power consumption edge AI audio signal processing method according to claim 1, wherein The target discrimination feature is input into the radial basis support vector machine, and the piecewise linear kernel function and the sequential minimal optimization algorithm with second-order Taylor expansion are used for training to obtain the audio processing model, including: Input the target discrimination feature into the radial basis support vector machine, construct a kernel matrix based on the radial basis kernel function, and divide the exponential operation interval into three segments, and use linear functions with different slopes to fit within each segment interval to obtain the piecewise linear kernel function; Construct a Hessian matrix for the calculation result of the piecewise linear kernel function, expand the Hessian matrix to the second-order term at the current iteration point based on the Taylor formula, and represent the Hessian matrix in the form of an upper triangular matrix through Cholesky decomposition to obtain the quadratic programming problem; Construct the objective function of sequential minimal optimization according to the quadratic programming problem, initialize the Lagrange multiplier to obtain the initial solution of the optimization problem, and calculate the update direction of the Lagrange multiplier based on the initial solution; Iteratively optimize the Lagrange multipliers along the update direction to obtain the weight coefficients of the support vectors, sparsify the weight coefficients, and calculate the piecewise linear kernel function by means of table lookup to obtain the audio processing model.

7. The ultra-low power consumption end-side AI audio signal processing method according to claim 1, characterized in that The ultra-low power edge AI audio signal processing method further includes: Deploy the audio processing model to a triple-core processor, where the first core performs signal preprocessing and feature extraction, the second core performs model inference, and the third core performs signal synthesis. Inter-core data transmission is achieved through a double-buffer FIFO queue to obtain a pipeline processing architecture; Monitor the operating status of each processing core. When the CPU occupancy rate exceeds the first target value, allocate the task load to the idle core. When the memory occupancy rate exceeds the second target value, compress and store the intermediate data to obtain a resource scheduling strategy; According to the resource scheduling strategy, dynamically adjust the operating frequencies of the three processing cores to obtain the working parameters of the processing cores; Based on the classification result of the audio processing model, separate the modulation period signal and the non-modulation signal of the input audio signal, and perform noise suppression on the non-modulation signal through a Wiener filter to obtain a first enhanced signal; Perform sub-band synthesis on the first enhanced signal and the modulation period signal, automatically adjust the gain coefficient of each frequency band based on the quantization noise parameter, and complete signal reconstruction through the overlap-add method to obtain a second enhanced signal; Perform smooth transition on the second enhanced signal through overlapping addition processing, and perform de-emphasis filtering to compensate for the frequency response to obtain an enhanced audio signal.

8. An ultra-low power end-side AI audio signal processing system, characterized in that, For implementing the steps of the method according to any one of claims 1 to 7, the system includes: A decomposition unit, configured to perform frame division and logarithmic compression processing on the input audio signal to obtain target spectrum data, and decompose the target spectrum data to obtain a modulation period signal and a non-modulation signal; An operation unit, configured to perform critical band analysis and variational Bayesian fixed-point operation on the non-modulation signal to obtain a quantization noise parameter; A construction unit, configured to extract the main period feature of the modulation period signal and the frequency band feature of the quantization noise parameter, and construct a target discrimination feature according to the main period feature and the frequency band feature; A training unit, configured to input the target discrimination feature into a radial basis support vector machine, and perform training by using a piecewise linear kernel function and a sequential minimal optimization algorithm with second-order Taylor expansion to obtain an audio processing model.

Citation Information

Cited By

  • Dynamic shift quantization modulation strong and weak mixed signal separation method and system based on parameter reconstruction

    CN121299587A