Audio processing method and device and related equipment
By constructing the target echo signal and performing adaptive linear filtering using normalized least mean square and Kalman filtering algorithms, combined with deep nonlinear enhancement processing, the problem of poor audio signal processing effect is solved, and high-accuracy audio processing is achieved in nonlinear distortion environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-03
AI Technical Summary
Current audio signal processing is inadequate, especially in nonlinear distortion environments where the performance of adaptive filtering algorithms is limited, and duplex detection mechanisms are not adaptable to complex acoustic scenarios.
By constructing the target echo signal, adaptive linear filtering is performed using normalized least mean square and Kalman filtering algorithms, combined with deep nonlinear enhancement processing, to perform duplex state recognition and deep nonlinear enhancement, thereby improving the processing accuracy of the audio signal.
The performance of the adaptive filtering algorithm is improved in nonlinear distortion environments, and its adaptability in complex acoustic scenarios is enhanced, thereby improving the accuracy of audio signal processing.
Smart Images

Figure CN121789704A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to an audio processing method, apparatus and related equipment. Background Technology
[0002] With the continuous development of communication technology, voice interaction is increasingly widely used in electronic devices. Currently, voice interaction allows users to control electronic devices to perform various operations without manual intervention, greatly improving user convenience. However, in practical use, the acquired audio signals often contain noise, which can easily interfere with the results of voice interaction. While adaptive filtering algorithms or duplex detection mechanisms are commonly used to process audio signals, adaptive filtering algorithms are prone to performance limitations in nonlinear distortion environments, and duplex detection mechanisms are not well-suited to complex acoustic scenarios. Therefore, current audio signal processing methods are relatively ineffective. Summary of the Invention
[0003] This application provides an audio processing method, apparatus, and related equipment to solve the problem of poor processing effect of current audio signals.
[0004] To solve the above problems, this application is implemented as follows:
[0005] In a first aspect, embodiments of this application provide an audio processing method applied to an electronic device, the method comprising:
[0006] Get the first audio file to be processed;
[0007] A target echo signal is constructed based on the first audio signal, and the target echo signal is constructed based on multiple echo component signals of the first audio signal;
[0008] The target echo signal is subjected to adaptive linear filtering processing according to a preset filtering algorithm to obtain the first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Squares (NLMS) algorithm and the Kalman filtering algorithm.
[0009] The first target signal is subjected to duplex state recognition processing to obtain a target result, which is used to represent the state of the electronic device when the first target signal is acquired.
[0010] If the target result meets the preset conditions, the first target signal is subjected to deep nonlinear enhancement processing to obtain the second target signal.
[0011] Secondly, embodiments of this application provide an audio processing apparatus applied to an electronic device, the audio processing apparatus comprising:
[0012] The acquisition module is used to acquire the first audio file to be processed.
[0013] A construction module is used to construct a target echo signal based on the first audio, wherein the target echo signal is constructed based on multiple echo component signals of the first audio.
[0014] The filtering module is used to perform adaptive linear filtering on the target echo signal according to a preset filtering algorithm to obtain a first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm.
[0015] The state recognition processing module is used to perform duplex state recognition processing on the first target signal to obtain a target result, wherein the target result is used to represent the state of the electronic device when the first target signal is acquired.
[0016] An enhancement processing module is used to perform deep nonlinear enhancement processing on the first target signal to obtain a second target signal when the target result meets preset conditions.
[0017] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect above.
[0018] Fourthly, embodiments of this application also provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps of the method described in the first aspect above.
[0019] Fifthly, embodiments of this application also provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method described in the first aspect above.
[0020] In this embodiment, the target echo signal can be a signal constructed from multiple echo component signals of the first audio. The types of different echo component signals can be different, so that when the target echo signal is subjected to adaptive linear filtering processing according to the preset filtering algorithm, the nonlinear echo component signals included in the target echo signal can be accurately identified, thereby enhancing the performance of the adaptive filtering algorithm in nonlinear distortion environments. At the same time, by processing the target echo signal with the preset filtering algorithm to obtain the first target signal, and performing duplex state recognition processing on the first target signal, the first audio can be combined with adaptive linear filtering processing and duplex state recognition, thereby enhancing the adaptability in complex acoustic scenes and improving the accuracy of the processing result of the first audio. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic flowchart of the audio processing method provided in the embodiments of this application;
[0023] Figure 2 This is an overall architecture diagram of the electronic device provided in the embodiments of this application;
[0024] Figure 3 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application;
[0025] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.
[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating the audio processing method provided in the embodiments of this application. Figure 1 The audio processing method shown can be executed by an electronic device, and the type of electronic device is not limited here. Optionally, the type of electronic device may include at least one of mobile devices, edge devices, and cloud devices. The network of the mobile device is a simplified network, which may include a 3-layer semantic segmentation network (U-Net) and a 1-layer long short-term memory (LSTM) network layer; the network of the edge device is a standard network, which may include a 5-layer U-Net and a 2-layer LSTM; and the network of the cloud device is an enhanced network, which may include a 7-layer U-Net, a 3-layer LSTM, and a Transformer layer.
[0029] like Figure 1 As shown, the audio processing method may include the following steps:
[0030] Step 101: Obtain the first audio file to be processed.
[0031] The method of acquiring the first audio is not limited here. Optionally, the first audio can be audio directly acquired by an electronic device. For example, the electronic device can acquire the first audio through an audio acquisition device. Alternatively, the first audio can be audio acquired by the electronic device from other electronic devices, which may include a cloud server or an electronic device that is communicatively connected to the electronic device in the embodiments of this application.
[0032] The specific content of the first audio may be related to the usage scenario of the electronic device in this application embodiment, and the specific content of the first audio is not limited here. Optionally, when the electronic device in this application embodiment is used in a financial scenario, the content of the first audio includes financial data. For example, the financial data may include audio of communication between bank staff and users regarding financial information. Alternatively, when the electronic device in this application embodiment is used in a communication scenario, the content of the first audio includes communication data.
[0033] For example, the first audio could also be the audio generated during communication with a digital human customer service system, or it could be the audio generated during communication with smart home devices.
[0034] Step 102: Construct a target echo signal based on the first audio, wherein the target echo signal is constructed based on multiple echo component signals of the first audio.
[0035] The specific method for constructing the target echo signal based on the first audio is not limited here. Optionally, the echo path corresponding to the first audio can be obtained first, and then the echo path can be split into multiple sub-paths. Each sub-path corresponds to an echo component signal. Multiple echo component signals corresponding to the multiple sub-paths can be obtained, and the multiple echo component signals can be aggregated to obtain the constructed target echo signal.
[0036] Alternatively, speech features such as time domain, frequency domain, cepstral domain, and wavelet domain of the first audio can be extracted first. Then, multiple echo component signals can be extracted for each speech feature in the time domain, frequency domain, cepstral domain, and wavelet domain. Then, the echo component signals corresponding to the same type of speech feature are aggregated. Finally, the multiple echo component signals are aggregated to obtain the target echo signal.
[0037] It should be noted that speech features in the time domain, frequency domain, cepstral domain, and wavelet domain can be found in the following descriptions:
[0038] 1. Time-domain characteristics, which may include information such as short-time energy, zero-crossing rate, and autocorrelation function;
[0039] The formula for calculating short-time energy is as follows: Where: E(n) represents the short-time energy value of the nth frame; x(ni) represents the amplitude of the speech signal at the nith sampling point; N represents the analysis window length; w(i) represents the window function weight; and i represents the index of the sampling point within the window.
[0040] The formula for calculating the zero-crossing rate is as follows: Where: ZCR(n) represents the zero-crossing rate of the nth frame; sgn() represents the sign function; x(ni) and x(ni-1) represent the speech signals of two adjacent sampling points; N represents the analysis frame length.
[0041] The formula for calculating the autocorrelation function is as follows: Where: R_xx(τ) represents the autocorrelation function value at a delay of τ; x(ni) represents the signal value at the nith sampling point; τ represents the time delay; and i represents the summation index.
[0042] 2. Frequency domain characteristics, which may include information such as power spectral density, spectral centroid, and spectral roll-off point;
[0043] The power spectral density is calculated using the following formula: PSD(ω) = |X(ω)|² / N, where: PSD(ω) represents the power spectral density at angular frequency ω; X(ω) represents the Fourier transform of the signal; N represents the number of transform points; and ω represents the angular frequency.
[0044] The formula for calculating the spectral centroid is as follows: , where: SC represents the spectral centroid value; k represents the frequency bin index; X(k) represents the spectral amplitude of the k-th frequency bin; |X(k)|² represents the power spectrum.
[0045] The formula for calculating the spectral roll-off point is as follows: Where: SR represents the frequency bin index of the spectral roll-off point; k represents the current frequency bin; X(i) represents the spectral amplitude of the i-th frequency bin; N represents the number of Fast Fourier Transform (FFT) points; 0.85 represents the 85% energy threshold.
[0046] 3. Cepstral domain characteristics, which may include information such as Mel-Frequency Cepstral Coefficients (MFCC) coefficients;
[0047] The formula for calculating the MFCC coefficient is as follows: Where: C(m) represents the m-th MFCC coefficient; S(k) represents the output of the k-th Mel filter bank; K represents the number of Mel filter banks; m represents the cepstral coefficient index; π represents pi.
[0048] 4. Wavelet domain features, which may include information such as continuous wavelet transform;
[0049] The calculation formula for continuous wavelet transform is as follows: Where: CWT(a,b) represents the wavelet coefficients at scale a and displacement b; x(t) represents the input signal; ψ* represents the conjugate of the wavelet basis function; a represents the scale parameter; b represents the displacement parameter; and t represents the time variable.
[0050] It should be noted that the adaptive feature weight allocation formulas for speech features in the time domain, frequency domain, cepstral domain, and wavelet domain can be found in the following calculation formulas:
[0051] Where: W_optimal represents the optimal weight vector; W_j represents the weight of the j-th feature domain; F_j^i represents the feature vector of the i-th sample in the j-th feature domain; y_i represents the true label of the i-th sample; f() represents the prediction function; λ represents the L1 regularization coefficient; Represents the L1 norm; This represents the L2 norm.
[0052] Step 103: Perform adaptive linear filtering on the target echo signal according to a preset filtering algorithm to obtain the first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm.
[0053] In this process, adaptive linear filtering of the target echo signal according to a preset filtering algorithm can remove noise signals from the target echo signal, thereby improving the quality of the obtained first target signal.
[0054] The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm. This increases the diversity of preset filtering algorithms, thereby making the adaptive linear filtering of the target echo signal more flexible and the processing methods more diverse.
[0055] The adaptive linear filtering process can be understood as: adaptively adjusting the processing parameters of the filtering process according to the signal parameters of the target echo signal, thereby making the processing accuracy of the target echo signal higher and further improving the processing effect of the target echo signal.
[0056] For example, when the signal parameters of the target echo signal include the signal length, and the signal length is greater than the preset length, the step size used for filtering the target echo signal can be adaptively increased, that is, the processing parameters can include the above step size.
[0057] Step 104: Perform duplex state recognition processing on the first target signal to obtain a target result. The target result is used to represent the state of the electronic device when the first target signal is acquired.
[0058] The specific form of the target result is not limited here. Optionally, the target result can be represented in at least one of the following ways: text, numbers, letters, etc.
[0059] When acquiring the first target signal, the state of the electronic device is not specifically limited here. Optionally, the above state may include at least one of the following: remote one-way talk state, near one-way talk state, and two-way talk state.
[0060] Step 105: If the target result meets the preset conditions, perform deep nonlinear enhancement processing on the first target signal to obtain the second target signal.
[0061] The specific method of deep nonlinear enhancement processing is not limited here. Optionally, the first target signal can be input into a deep neural network for deep nonlinear enhancement processing to obtain the second target signal. Alternatively, the first target signal can be subjected to deep nonlinear enhancement processing according to a deep nonlinear enhancement processing algorithm to obtain the second target signal. The specific type of the deep nonlinear enhancement processing algorithm is not limited here. For example, the deep nonlinear enhancement processing algorithm can be a self-constructed algorithm, or it can be a general enhancement processing algorithm.
[0062] The condition that the target result meets the preset conditions is not specifically limited here. Optionally, the preset conditions may include preset results. When the target result matches the preset results, it can be determined that the target result meets the preset conditions; otherwise, it can be determined that the target result does not meet the preset conditions. Alternatively, the preset conditions may include preset identifiers. When the target result is a preset identifier, it can be determined that the target result matches the preset conditions; otherwise, it can be determined that the target result does not meet the preset conditions.
[0063] In this embodiment, through steps 101 to 105, the target echo signal can be a signal constructed from multiple echo component signals of the first audio. The types of different echo component signals can be different, so that when the target echo signal is subjected to adaptive linear filtering processing according to the preset filtering algorithm, the nonlinear echo component signals included in the target echo signal can be accurately identified, thereby enhancing the performance of the adaptive filtering algorithm in nonlinear distortion environments. At the same time, by processing the target echo signal with the preset filtering algorithm to obtain the first target signal, and performing duplex state recognition processing on the first target signal, the first audio can be combined with adaptive linear filtering processing and duplex state recognition, thereby enhancing the adaptability in complex acoustic scenes and improving the accuracy of the processing result of the first audio.
[0064] It should be noted that due to the high complexity of audio echo paths, such as linear echo, nonlinear distortion, and multipath propagation effects, the audio processing effect in related technologies is usually poor. To address this problem, the following implementation method is proposed:
[0065] As an optional implementation, constructing the target echo signal based on the first audio signal includes:
[0066] Acquire multiple echo component signals of the first audio, the multiple echo component signals including: linear echo component signal, nonlinear echo component signal, multipath echo component signal and coupled echo component signal;
[0067] The target echo signal is constructed based on the linear echo component signal, the nonlinear echo component signal, the multipath echo component signal, and the coupled echo component signal.
[0068] In this embodiment, the linear echo component signal, nonlinear echo component signal, multipath echo component signal, and coupled echo component signal included in the first audio are first extracted, thereby ensuring high accuracy of the extracted linear echo component signal, nonlinear echo component signal, multipath echo component signal, and coupled echo component signal. Then, a target echo signal is constructed based on the linear echo component signal, nonlinear echo component signal, multipath echo component signal, and coupled echo component signal. In this way, since the constructed target echo signal can retain the information in the linear echo component signal, nonlinear echo component signal, multipath echo component signal, and coupled echo component signal, the diversity and accuracy of the information in the constructed target echo signal are further improved.
[0069] For example, the formula for calculating the target echo signal can be seen as follows: e(n) = e_linear(n) + e_nonlinear(n) + e_multipath(n) + e_coupling(n); where e(n) represents the total echo signal at the nth sampling point (i.e., the target echo signal); e_linear(n) represents the linear echo component (i.e., the linear echo component signal); e_nonlinear(n) represents the nonlinear echo component (i.e., the nonlinear echo component signal); e_multipath(n) represents the multipath echo component (i.e., the multipath echo component signal); e_coupling(n) represents the coupled echo component (i.e., the coupled echo component signal); and n represents the time sampling point index.
[0070] The formula for calculating the linear echo component signal is as follows: Where: h_linear(k) represents the linear echo path impulse response of the k-th tap; x(nk) represents the far-end signal delayed by k sampling points; L represents the filter length; and k represents the tap index.
[0071] The formula for calculating the nonlinear echo component signal is as follows: ,in: h_m(k) represents the m-th order nonlinear coefficient; h_m(k) represents the k-th tap coefficient of the m-th order nonlinear path; M represents the highest nonlinear order; m represents the nonlinear order.
[0072] The calculation formula for the multipath echo component signal is as follows: ;in: This represents the amplitude attenuation coefficient of the i-th path; This represents the time delay of the i-th path; Let I represent the attenuation constant of the i-th path; let I represent the total number of multipath paths; and let i represent the path index.
[0073] The formula for calculating the coupled echo component signal is as follows: ;in: This represents the coupling coefficient of the j-th coupling path; Indicates delay The remote signal; Indicates delay The near-end signal; J represents the total number of coupled paths; j represents the coupled path index.
[0074] As an optional implementation, the preset filtering algorithm includes the NLMS algorithm and the Kalman filtering algorithm. The step of performing adaptive linear filtering on the target echo signal according to the preset filtering algorithm to obtain the first target signal includes:
[0075] The target echo signal is subjected to adaptive linear filtering processing according to the NLMS algorithm to obtain a first signal, and the target echo signal is subjected to echo cancellation processing according to the Kalman filtering algorithm to obtain a second signal, wherein the step size of the NLMS algorithm is an adaptively adjusted step size based on the target echo signal;
[0076] The first signal and the second signal are aggregated to obtain the first target signal.
[0077] Since the step size of the NLMS algorithm is adaptively adjusted based on the target echo signal, the NLMS algorithm can also be called the variable step size NLMS algorithm. The calculation formula for the variable step size NLMS algorithm can be found below: Where: μ(n) represents the adaptive step size of the nth iteration; The initial step size is represented by α; the step size adjustment factor is represented by e(n); and the error signal at the nth iteration is represented by α. δ represents the variance of the error signal; w(n) represents the weight vector of the nth iteration; x(n) represents the input signal vector of the nth iteration; δ represents a small constant to prevent division by zero. This represents the 2-norm of a vector.
[0078] It should be noted that, optionally, the number of taps can be dynamically adjusted according to the following calculation formula: This allows for further improvement in the accuracy of the tap count, among which... Based on the number of taps, For reverberation time, For reference time, This represents the incremental tap count.
[0079] Alternatively, the step size of the NLMS algorithm can be optimized according to the following calculation formula. , This indicates an optimized step size, which can further improve the accuracy of the step size.
[0080] It should be noted that, optionally, the weight error vector of the step size needs to satisfy the following formula: ;
[0081] The calculation formula for the Kalman filter algorithm can be found below:
[0082] State equation: w(n+1) = w(n) + v(n);
[0083] Observation equation: d(n) = x^T(n)w(n)+η(n);
[0084] Kalman gain: K(n) = P(n-1)x(n) / (λ+ x^T(n)P(n-1)x(n));
[0085] Weight update: w(n) = w(n-1) + K(n)e(n);
[0086] Covariance update: P(n) = (I - K(n)x^T(n))P(n-1) / λ;
[0087] Where: w(n) represents the state vector (weight vector) of the nth iteration; v(n) represents the process noise; d(n) represents the desired signal; x^T(n) represents the transpose of the input vector; η(n) represents the observation noise; K(n) represents the Kalman gain matrix; P(n) represents the error covariance matrix; λ represents the forgetting factor; and I represents the identity matrix.
[0088] In this embodiment, the target echo signal is processed by adaptive linear filtering using the NLMS algorithm to obtain a first signal, and the target echo signal is processed by echo cancellation using the Kalman filter algorithm to obtain a second signal. In this way, noise signals in the target echo signal can be eliminated by the NLMS algorithm and the Kalman filter algorithm, thereby improving the signal quality of the first signal and the second signal. At the same time, since the NLMS algorithm and the Kalman filter algorithm may mistakenly eliminate some normal signals when eliminating noise signals, but the normal signals eliminated by the NLMS algorithm and the Kalman filter algorithm are different, the first signal and the second signal can be aggregated. This allows the obtained first target signal to retain the normal signals in both the first signal and the second signal, thus enhancing the integrity and accuracy of the first target signal.
[0089] It should be noted that noise signal identification can be performed using the following expression: Noise_type = arg max_kP(noise_features|model_k), and the types of noise signals can include at least one of the following: stationary white noise, colored noise, impulse noise, speech noise, and music noise. Furthermore, to enhance the noise signal identification effect, other parameters of the electronic device can also be adaptively adjusted, and the adjustment formula can be found in the following expression: .
[0090] It should be noted that the theoretical limit of echo cancellation can be found in the following formula: The theoretical limit of noise suppression can be seen in the following formula: .
[0091] As an optional implementation, the step of performing duplex state recognition processing on the first target signal to obtain the target result includes:
[0092] The first target signal is input into a duplex system for duplex state identification processing to obtain the target result;
[0093] The duplex system includes at least one of the following: an energy domain detector, a correlation domain detector, a spectrum domain detector, a cepstral domain detector, and a speech activity detector.
[0094] In this embodiment of the application, the first target signal is subjected to duplex state recognition processing by at least one of the energy domain detector, correlation domain detector, spectrum domain detector, cepstral domain detector, and voice activity detector, thereby further improving the accuracy of the obtained target result.
[0095] It should be noted that the more types of detectors a duplex system includes, the higher the accuracy of the target results obtained.
[0096] For example, a duplex system may include an energy domain detector, a correlation domain detector, a spectrum domain detector, a cepstral domain detector, and a speech activity detector. When performing duplex state recognition processing, these detectors can be processed according to the following calculation formulas:
[0097] The calculation formula for the energy domain detector is as follows: Where: P_energy represents the output probability of the energy domain detector; sigmoid() represents the sigmoid activation function; E_near represents the energy ratio coefficient; E_near represents the near-end signal energy; E_far represents the far-end signal energy. This represents the energy domain bias parameter.
[0098] The calculation formula for the correlation domain detector is as follows: ;in: This represents the output probability of the correlation domain detector; Indicates the relevant domain weight parameters; This represents the zero-delay cross-correlation value; The cross-correlation function representing the delay τ; This indicates that the maximum value is taken for all delays τ.
[0099] The calculation formula for the spectrum domain detector is as follows: Where: P_spectral represents the output probability of the spectral domain detector; The spectral domain weight parameter is represented by KL(); the (Kullback-Leibler, KL) divergence is represented by S_near; the near-end signal spectrum is represented by S_far; and the far-end signal spectrum is represented by JS(); the (Jensen-Shannon, JS) divergence is represented by JS().
[0100] The calculation formula for the cepstral domain detector is as follows: Where: P_cepstral represents the output probability of the cepstral detector; C represents the cepstral domain weighting parameter; C_near represents the cepstral characteristics of the near-end signal; C_far represents the cepstral characteristics of the far-end signal; It represents Euclidean distance.
[0101] The formula for calculating the speech activity detector is as follows: Where: P_VAD represents the output probability of the speech activity detector; The values represent the weight parameters for Voice Activity Detection (VAD); VAD_score_near represents the near-end voice activity score; and VAD_score_far represents the far-end voice activity score.
[0102] As an optional implementation, the duplex system includes: the energy domain detector, the correlation domain detector, the spectral domain detector, the cepstral domain detector, and the voice activity detector. The step of inputting the first target signal into the duplex system for duplex state recognition processing to obtain the target result includes:
[0103] The first target signal is input to the energy domain detector to obtain an energy domain detection value, the first target signal is input to the correlation domain detector to obtain a correlation domain detection value, the first target signal is input to the spectrum domain detector to obtain a spectrum domain detection value, and the first target signal is input to the speech activity detector to obtain a speech activity detection value.
[0104] Calculate the weighted sum of the energy domain detection value, the correlation domain detection value, the spectrum domain detection value, and the speech activity detection value, and determine the weighted sum as the target result.
[0105] In this embodiment, a weighted sum of the energy domain detection value, the correlation domain detection value, the spectrum domain detection value, and the speech activity detection value is calculated, and the weighted sum is determined as the target result. This can further improve the accuracy of the calculated target result, and at the same time, it can also enhance the diversity and flexibility of the target result calculation method.
[0106] It should be noted that when a duplex system includes an energy domain detector, a correlation domain detector, a spectrum domain detector, a cepstral domain detector, and a speech activity detector, the information detected by these detectors can be fused, and the target result can be determined based on the fused information. The calculation formula for fusing the above multiple types of information can be found below: ;in: This represents the final fusion probability (i.e., the target result); This represents the weight of the i-th detector (i.e., the detector numbered i among the energy domain detector, correlation domain detector, spectrum domain detector, cepstral domain detector, and speech activity detector); The output probability of the i-th detector (i.e., the information output by the detector, such as energy domain detection value, correlation domain detection value, spectrum domain detection value, or speech activity detection value); i represents the detector index (1 to 5).
[0107] As an optional implementation, the target result includes a target probability, and the step of performing deep nonlinear enhancement processing on the first target signal to obtain a second target signal when the target result meets preset conditions includes:
[0108] If the target probability is greater than the preset probability, the first target signal is corrected.
[0109] The second target signal is obtained by performing deep nonlinear enhancement processing on the corrected first target signal.
[0110] The correction process can be understood as making the first target signal after correction more suitable for deep nonlinear enhancement processing, thereby improving the accuracy of the final second target signal.
[0111] For example, by modifying the first target signal, the modified first target signal can better conform to the format of the input signal for deep nonlinear enhancement processing. For another example, deep nonlinear enhancement processing can be performed using a deep neural network model. By modifying the first target signal, the modified first target signal can have a higher degree of fit with the network layers of the deep neural network model, thereby enabling the deep neural network model to modify the first target signal more accurately.
[0112] It should be noted that the specific method of correction is not limited here. Optionally, the correction may include at least one of the following: correcting the format of the first target signal, adding vector features of the first target signal, deleting vector features of the first target signal, modifying the representation of the vector corresponding to the first target signal, etc.
[0113] In this embodiment, correcting the first target signal makes it more suitable for deep nonlinear enhancement processing, thereby improving the accuracy of the final obtained second target signal.
[0114] Optionally, the deep neural network model may include at least one of the following: a multi-scale time-frequency attention network, a temporal attention mechanism, a frequency attention mechanism, and an adaptive loss function.
[0115] 1. Multi-scale Time-Frequency Attention Network (MSTF-AttNet):
[0116] The overall network architecture is as follows: Input layer → Multi-Scale Feature Extraction layer → Temporal Attention layer → Frequency Attention layer → Cross-Domain Fusion layer → Output layer.
[0117] The calculation formula for the multi-scale feature extraction module is as follows:
[0118] Scale_1: Conv2D(kernel_size=3×3, dilation=1);
[0119] Scale_2: Conv2D(kernel_size=3×3, dilation=2);
[0120] Scale_3: Conv2D(kernel_size=3×3, dilation=4);
[0121] Scale_4: Conv2D(kernel_size=3×3, dilation=8);
[0122] The calculation formula for the time attention mechanism is as follows: Where: X represents the input feature tensor; ⊙ represents element-wise multiplication; σ represents the Sigmoid activation function; Conv1D represents one-dimensional convolution operation; AvgPool represents average pooling; MaxPool represents max pooling.
[0123] The calculation formula for the frequency attention mechanism is as follows: Frequency_Attention(X) = X ⊙ σ(MLP(GAP(X)) + MLP(GMP(X))); where: MLP represents Multi-Layer Perceptron; GAP represents Global Average Pooling; GMP represents Global Max Pooling.
[0124] The formula for calculating the adaptive loss function is as follows: Where: L_total represents the total loss function; arrive L represents the weight coefficient of each loss term; L_reconstruction represents the reconstruction loss; L_perceptual represents the perceptual loss; L_adversarial represents the adversarial loss; L_spectral represents the spectral loss; L_temporal represents the temporal loss; and L_consistency represents the consistency loss.
[0125] The formula for calculating the reconstruction loss is as follows: Where: S_enhanced represents the enhanced speech signal; S_clean represents the clean target speech signal; α represents the L1 regularization weight.
[0126] The formula for calculating the perceived loss is as follows: ;in: Let l represent the feature extraction function of the l-th layer of the pre-trained network; l represents the network layer index.
[0127] The formula for calculating the spectral loss is as follows: Where: STFT represents Short Time Fourier Transform; s_enhanced represents the enhanced time-domain signal; s_clean represents the clean time-domain signal; β represents the logarithmic spectral loss weight; and log represents the natural logarithm.
[0128] It should be noted that, in order to enhance the performance of deep neural network models, model compression and quantization, adaptive parameter adjustment and optimization, and hyperparameter adjustment and optimization can also be performed on the deep neural network models.
[0129] The model compression and quantization can be calculated using the following methods:
[0130] Weight quantization: ;
[0131] Knowledge distillation: ; where scale is used to represent the scale.
[0132] The adjustment and optimization of the adaptive parameters can be calculated as follows:
[0133] Parameter optimization based on reinforcement learning:
[0134] State space: ;
[0135] Action space: ;
[0136] Reward function: .
[0137] The description of hyperparameter tuning and optimization is as follows: Learning rate scheduling strategy: .
[0138] As an optional implementation, after performing deep nonlinear enhancement processing on the first target signal to obtain the second target signal when the target result meets preset conditions, the method further includes:
[0139] The second target signal is processed to obtain the third target signal;
[0140] The target processing includes at least one of the following: adaptive beamforming processing and three-dimensional sound source localization processing.
[0141] In this embodiment of the application, by performing target processing on the second target signal, the second target signal can be optimized, thereby further improving the accuracy of the obtained third target signal.
[0142] The adaptive beamforming process can include at least one of the following methods: generalized sidelobe canceller (GSC) beamforming and robust adaptive beamforming;
[0143] The calculation formula for the beamforming of the generalized sidelobe canceller (GSC) is as follows: , where: y(n) represents the beamformer output at time n; w_qH represents the conjugate transpose of the static beamformer weight vector; x(n) represents the multi-channel input signal vector; w_aH represents the conjugate transpose of the adaptive weight vector; B^H represents the conjugate transpose of the blocking matrix.
[0144] The calculation formula for robust adaptive beamforming is shown below: Where: w_MVDR represents the minimum variance distortionless response beamformer weights; R_nn represents the noise covariance matrix; δ represents the diagonal loading factor; I represents the identity matrix; d represents the direction vector of the desired signal; dH represents the conjugate transpose of the direction vector; {-1} represents matrix inversion.
[0145] The three-dimensional sound source localization process can be understood as processing based on a three-dimensional sound source localization algorithm. This algorithm can employ the Multiple Signal Classification Algorithm (MUSIC), and the calculation formula for the MUSIC algorithm can be found below: ;
[0146] Where: P_MUSIC(θ,φ) represents the MUSIC spectrum at azimuth angle θ and elevation angle φ; a(θ,φ) represents the array manifold vector of direction (θ,φ); E_n represents the eigenvector matrix of the noise subspace; E_n^H represents the conjugate transpose of the noise subspace matrix; θ represents the azimuth angle; φ represents the elevation angle.
[0147] It should be noted that, in the embodiments of this application, the relevant information of each evaluation indicator can be found in Table 1 below.
[0148] Table 1. Relevant Information on Evaluation Indicators
[0149]
[0150] It should be noted that, in order to improve the efficiency of obtaining the second target signal, the time delay of each step can be controlled, and the time delay control values of each step can be as follows: buffer delay: 5ms, feature extraction: 2ms, linear filtering: 3ms, duplex detection: 1ms, deep neural network (DNN) inference: 30ms, signal synthesis: 4ms. The time delay of each step can be decomposed according to the frame-level processing delay: see the following formula: T_total = T_buffer + T_feature + T_linear + T_detection + T_DNN + T_synthesis, where T_total is used to represent the total time delay, T_buffer is used to represent the buffer delay, T_feature is used to represent the feature extraction delay, T_linear is used to represent the linear filtering delay, T_detection is used to represent the duplex detection delay, T_DNN is used to represent the DNN inference delay, and T_synthesis is used to represent the signal synthesis delay.
[0151] It should be noted that, in order to more fully illustrate the embodiments of this application, as... Figure 2 As shown in the figure, this application also provides an overall architecture diagram of an electronic device. Figure 2 The full-duplex speech enhancement system in this application can be understood as the electronic device in the embodiments of this application. This electronic device may include the following structural layers: an input layer, a multi-channel processing layer, a multi-domain feature extraction layer, a hierarchical echo modeling layer, an adaptive filtering layer, an intelligent duplex detection layer, a deep neural network enhancement layer, a post-processing optimization layer, and an output layer. Through the cooperation of the above structural layers, the quality of the final output second target signal can be improved. The specific operations performed by each structural layer can be found in [reference needed]. Figure 2 As shown, the specifics will not be elaborated here.
[0152] See Figure 3 , Figure 3 This is a structural diagram of the audio processing device provided in the embodiments of this application, such as... Figure 3 As shown, the audio processing device 300 includes:
[0153] Acquisition module 301 is used to acquire the first audio to be processed;
[0154] The construction module 302 is used to construct a target echo signal based on the first audio, wherein the target echo signal is constructed based on multiple echo component signals of the first audio.
[0155] The filtering module 303 is used to perform adaptive linear filtering on the target echo signal according to a preset filtering algorithm to obtain a first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm.
[0156] The state recognition processing module 304 is used to perform duplex state recognition processing on the first target signal to obtain a target result, wherein the target result is used to represent the state of the electronic device when the first target signal is acquired.
[0157] The enhancement processing module 305 is used to perform deep nonlinear enhancement processing on the first target signal to obtain a second target signal when the target result meets the preset conditions.
[0158] As an optional implementation, construction module 302 includes:
[0159] The acquisition submodule is used to acquire multiple echo component signals of the first audio, the multiple echo component signals including: linear echo component signal, nonlinear echo component signal, multipath echo component signal and coupled echo component signal;
[0160] A submodule is constructed to construct the target echo signal based on the linear echo component signal, the nonlinear echo component signal, the multipath echo component signal, and the coupled echo component signal.
[0161] As an optional implementation, the preset filtering algorithm includes the NLMS algorithm and the Kalman filter algorithm. The filtering processing module 303 includes:
[0162] The filtering submodule is used to perform adaptive linear filtering on the target echo signal according to the NLMS algorithm to obtain a first signal, and to perform echo cancellation on the target echo signal according to the Kalman filtering algorithm to obtain a second signal, wherein the step size of the NLMS algorithm is an adaptively adjusted step size based on the target echo signal.
[0163] The aggregation processing submodule is used to aggregate the first signal and the second signal to obtain the first target signal.
[0164] As an optional implementation, the state recognition processing module 304 is further configured to input the first target signal into the duplex system for duplex state recognition processing to obtain the target result;
[0165] The duplex system includes at least one of the following: an energy domain detector, a correlation domain detector, a spectrum domain detector, a cepstral domain detector, and a speech activity detector.
[0166] As an optional implementation, the duplex system includes: the energy domain detector, the correlation domain detector, the spectrum domain detector, the cepstral domain detector, the voice activity detector, and the state recognition processing module 304, including:
[0167] The detection submodule is used to input the first target signal into the energy domain detector to obtain an energy domain detection value, input the first target signal into the correlation domain detector to obtain a correlation domain detection value, input the first target signal into the spectrum domain detector to obtain a spectrum domain detection value, and input the first target signal into the speech activity detector to obtain a speech activity detection value.
[0168] The calculation submodule is used to calculate the weighted sum of the energy domain detection value, the correlation domain detection value, the spectrum domain detection value and the speech activity detection value, and to determine the weighted sum as the target result.
[0169] As an optional implementation, the target result includes a target probability, and the enhancement processing module 305 includes:
[0170] The correction submodule is used to correct the first target signal when the target probability is greater than a preset probability.
[0171] The deep nonlinear enhancement processing submodule is used to perform deep nonlinear enhancement processing on the corrected first target signal to obtain the second target signal.
[0172] As an optional implementation, the audio processing device 300 further includes:
[0173] The target processing module is used to process the second target signal to obtain the third target signal;
[0174] The target processing includes at least one of the following: adaptive beamforming processing and three-dimensional sound source localization processing.
[0175] The audio processing device 300 is capable of implementing the embodiments of this application. Figure 1 The various processes in the method embodiments, and the ways to achieve the same beneficial effects, will not be repeated here to avoid repetition.
[0176] This application also provides an electronic device. Please refer to [link to relevant documentation]. Figure 4 The electronic device may include a processor 401, a memory 402, and a program 4021 stored in the memory 402 and executable on the processor 401. When the program 4021 is executed by the processor 401, it can achieve... Figure 1 Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.
[0177] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium. This application also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the above-described methods. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0178] The storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0179] This application also provides a computer program product, including computer instructions, which, when executed by a processor, can perform the above-described functions. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0180] The above description represents the preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An audio processing method, characterized in that, Applied to electronic devices, the method includes: Get the first audio file to be processed; A target echo signal is constructed based on the first audio signal, and the target echo signal is constructed based on multiple echo component signals of the first audio signal; The target echo signal is subjected to adaptive linear filtering processing according to a preset filtering algorithm to obtain a first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm. The first target signal is subjected to duplex state recognition processing to obtain a target result, which is used to represent the state of the electronic device when the first target signal is acquired. If the target result meets the preset conditions, the first target signal is subjected to deep nonlinear enhancement processing to obtain the second target signal.
2. The method according to claim 1, characterized in that, The step of constructing the target echo signal based on the first audio includes: Acquire multiple echo component signals of the first audio, the multiple echo component signals including: linear echo component signal, nonlinear echo component signal, multipath echo component signal and coupled echo component signal; The target echo signal is constructed based on the linear echo component signal, the nonlinear echo component signal, the multipath echo component signal, and the coupled echo component signal.
3. The method according to claim 1, characterized in that, The preset filtering algorithm includes the NLMS algorithm and the Kalman filtering algorithm. The step of performing adaptive linear filtering on the target echo signal according to the preset filtering algorithm to obtain the first target signal includes: The target echo signal is subjected to adaptive linear filtering processing according to the NLMS algorithm to obtain a first signal, and the target echo signal is subjected to echo cancellation processing according to the Kalman filtering algorithm to obtain a second signal, wherein the step size of the NLMS algorithm is an adaptively adjusted step size based on the target echo signal; The first signal and the second signal are aggregated to obtain the first target signal.
4. The method according to claim 1, characterized in that, The process of performing duplex state recognition processing on the first target signal to obtain the target result includes: The first target signal is input into a duplex system for duplex state identification processing to obtain the target result; The duplex system includes at least one of the following: an energy domain detector, a correlation domain detector, a spectrum domain detector, a cepstral domain detector, and a speech activity detector.
5. The method according to claim 4, characterized in that, The duplex system includes: the energy domain detector, the correlation domain detector, the spectrum domain detector, the cepstral domain detector, and the speech activity detector. The step of inputting the first target signal into the duplex system for duplex state recognition processing to obtain the target result includes: The first target signal is input to the energy domain detector to obtain an energy domain detection value, the first target signal is input to the correlation domain detector to obtain a correlation domain detection value, the first target signal is input to the spectrum domain detector to obtain a spectrum domain detection value, and the first target signal is input to the speech activity detector to obtain a speech activity detection value. Calculate the weighted sum of the energy domain detection value, the correlation domain detection value, the spectrum domain detection value, and the speech activity detection value, and determine the weighted sum as the target result.
6. The method according to claim 4, characterized in that, The target result includes a target probability. The step of performing deep nonlinear enhancement processing on the first target signal to obtain a second target signal when the target result meets preset conditions includes: If the target probability is greater than the preset probability, the first target signal is corrected. The second target signal is obtained by performing deep nonlinear enhancement processing on the corrected first target signal.
7. The method according to any one of claims 1 to 6, characterized in that, After performing deep nonlinear enhancement processing on the first target signal to obtain the second target signal when the target result meets preset conditions, the method further includes: The second target signal is processed to obtain the third target signal; The target processing includes at least one of the following: adaptive beamforming processing and three-dimensional sound source localization processing.
8. An audio processing apparatus, characterized in that, The audio processing device, applied to an electronic device, includes: The acquisition module is used to acquire the first audio file to be processed. A construction module is used to construct a target echo signal based on the first audio, wherein the target echo signal is constructed based on multiple echo component signals of the first audio. The filtering module is used to perform adaptive linear filtering on the target echo signal according to a preset filtering algorithm to obtain a first target signal. The preset filtering algorithm includes at least one of the Normalized Least Mean Square (NLMS) algorithm and the Kalman filter algorithm. The state recognition processing module is used to perform duplex state recognition processing on the first target signal to obtain a target result, wherein the target result is used to represent the state of the electronic device when the first target signal is acquired. An enhancement processing module is used to perform deep nonlinear enhancement processing on the first target signal to obtain a second target signal when the target result meets preset conditions.
9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps of the audio processing method as described in any one of claims 1 to 7.
10. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps of the audio processing method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the audio processing method as described in any one of claims 1 to 7.