Loudspeaker voice enhancement method and device, electronic equipment and storage medium

By using a feedback cancellation network model and an adaptive feedback control algorithm, combined with real and complex signal masking, the speech quality problems caused by noise and feedback during loudspeaker amplification are solved, resulting in improved clarity, reduced latency, suppression of howling, and enhanced user experience.

CN120564744BActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511054655.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-18
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

During the amplification process, the speaker is affected by environmental noise, room reverberation and speaker feedback, resulting in poor voice quality and large delay, which affects the user's listening experience. In addition, the existing directional microphones cannot effectively deal with the feedback phenomenon.

Method used

A feedback cancellation network model, including a first sub-model and a second sub-model, is adopted. The amplitude and phase of the initial spectral signal are masked by real signal mask and complex signal mask. Combined with an adaptive feedback control algorithm, the loudspeaker feedback signal is eliminated, and frequency shifting, phase modulation and half-wave rectification are performed to generate the target frequency domain signal.

Benefits of technology

It effectively suppresses environmental noise and speaker feedback signals, improves voice clarity, reduces link latency, avoids echo problems, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564744B_ABST
    Figure CN120564744B_ABST
Patent Text Reader

Abstract

The application provides a loudspeaker voice enhancement method and device, electronic equipment and storage medium, and belongs to the technical field of voice signal processing, and comprises the following steps: obtaining an initial frequency spectrum signal and an initial power spectrum signal corresponding to an initial audio signal; inputting the initial power spectrum signal into a feedback cancellation network model to obtain a target frequency domain signal; and obtaining a target audio signal based on the target frequency domain signal. The application firstly adjusts the frequency domain amplitude of the initial frequency spectrum signal by using a real signal mask to generate a frequency domain estimation signal, then adjusts the frequency domain amplitude and phase of the initial frequency spectrum signal by using a complex signal mask to obtain a frequency domain enhancement signal, and finally determines the target frequency domain signal based on the frequency domain estimation signal and the frequency domain enhancement signal, so that the output target frequency domain signal has higher fidelity in amplitude and phase, and can remove environmental noise and loudspeaker feedback signals to the greatest extent and effectively suppress the howling phenomenon.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, and in particular to a loudspeaker speech enhancement method and device, electronic equipment and storage medium. BACKGROUND

[0002] During the amplification process, the loudspeaker is often affected by environmental noise, room reverberation and loudspeaker feedback speech, which may cause distortion of the played speech, unnatural listening experience and even howling phenomenon, thereby seriously affecting the speech quality.

[0003] Currently, directional microphones are often used to receive sound from a certain direction, thereby suppressing noise and interference from other directions. However, directional microphone pickup is essentially a passive defense measure, but it lacks effective intervention on the feedback signal that occurs during the amplification process, which makes it unable to deal with the howling phenomenon that has already occurred. SUMMARY

[0004] The present application provides a loudspeaker speech enhancement method, device, electronic equipment and storage medium to solve the defects of poor speech quality, large amplification delay and impact on user listening experience in the prior art.

[0005] The present application provides a loudspeaker speech enhancement method, comprising the following steps:

[0006] Obtain an initial frequency spectrum signal corresponding to an initial audio signal, and determine an initial power spectrum signal corresponding to the initial frequency spectrum signal;

[0007] Input the initial power spectrum signal into a feedback cancellation network model to obtain a target frequency domain signal output by the feedback cancellation network model, wherein the feedback cancellation network model comprises a first sub-model and a second sub-model;

[0008] Obtain a target audio signal based on the target frequency domain signal;

[0009] The first sub-model outputs a real signal mask according to the input initial power spectrum signal; the real signal mask is used to make mask adjustment on the frequency domain amplitude of the initial frequency spectrum signal to obtain a frequency domain estimation signal;

[0010] The second sub-model receives a three-dimensional frequency spectrum signal obtained by splicing the frequency domain estimation signal and the initial frequency spectrum signal, and outputs a complex signal mask; the complex signal mask is used to make mask adjustment on the frequency domain amplitude and phase of the initial frequency spectrum signal to obtain a frequency domain enhancement signal;

[0011] The target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal.

[0012] According to a speaker speech enhancement method provided by the present invention, the first sub-model includes a first input module, a first GRU module, a second GRU module, a splicing module, a third GRU module, and a first output module;

[0013] The first input module receives the initial power spectrum signal and inputs the initial power spectrum signal in parallel to the first GRU module and the second GRU module;

[0014] The first GRU module performs time-series modeling on the initial power spectrum signal to obtain a first time-series feature; the second GRU module performs time-series modeling on the initial power spectrum signal to obtain a second time-series feature.

[0015] The splicing module splices the first temporal feature and the second temporal feature along the feature dimension to obtain the fused temporal feature;

[0016] The third GRU module receives the fused timing features, performs timing dependency modeling on the fused timing features, and obtains the final timing features;

[0017] The first output module is used to map the final timing features to the real number signal mask.

[0018] According to a speaker speech enhancement method provided by the present invention, the second sub-model includes a second input module, an encoder module, a bottleneck processing module, a decoder module, and a second output module;

[0019] The second input module inputs the received three-dimensional spectrum signal to the encoder module;

[0020] The encoder module includes multiple cascaded upsampling convolutional layers, which are used to perform layer-by-layer convolution and upsampling on the three-dimensional spectral signal and output intermediate coding features at each level.

[0021] The bottleneck processing module performs temporal context modeling on the intermediate encoded features output by the last upsampled convolutional layer to perform filtering and weighting processing on adjacent time frames and output the temporally enhanced bottleneck features.

[0022] The decoder module includes multiple cascaded downsampling convolutional layers, which are used to downsample and convolve the bottleneck features layer by layer to output a frequency domain decoded signal.

[0023] The second output module performs convolution mapping on the frequency domain decoded signal and outputs the complex signal mask.

[0024] According to a speaker voice enhancement method provided by the present invention, the bottleneck processing module includes a gated loop unit, a bidirectional gated loop unit, and a linear mapping layer;

[0025] The gated loop unit receives the intermediate encoded features output by the last upsampled convolutional layer of the encoder module, performs temporal modeling on them along the time frame dimension, and outputs the first temporal feature;

[0026] The bidirectional gated loop unit receives the first time-series feature, performs bidirectional time-series modeling to execute filtering and weighting processing of adjacent time frames, and outputs the second time-series feature;

[0027] The linear mapping layer performs a mapping transformation on the second temporal feature and outputs the bottleneck feature after temporal enhancement.

[0028] According to a speaker speech enhancement method provided by the present invention, the downsampling convolutional layer in the decoder module corresponds one-to-one with the upsampling convolutional layer in the encoder module;

[0029] During the process of the decoder module performing layer-by-layer downsampling and convolution on the bottleneck features to output the frequency domain decoded signal, the intermediate encoded features with the same resolution output by the corresponding upsampling convolutional layer are fused through skip connections at each downsampling convolutional layer.

[0030] According to a speaker speech enhancement method provided by the present invention, the feedback cancellation network model further includes a weighted fusion module;

[0031] The weighted fusion module performs weighted fusion on the frequency domain estimation signal output by the first sub-model and the frequency domain enhancement signal output by the second sub-model to generate the target frequency domain signal.

[0032] According to a speaker voice enhancement method provided by the present invention, the initial audio signal is obtained based on the following method:

[0033] Acquire the microphone signal;

[0034] The microphone received signal is sampled according to the preset synthesis window length and preset frame shift to obtain the initial audio signal obtained from each sampling.

[0035] According to a speaker voice enhancement method provided by the present invention, the microphone receiving signal includes at least a target voice signal and a speaker feedback signal;

[0036] After acquiring the microphone received signal, the speaker feedback signal is eliminated based on an adaptive feedback control algorithm, specifically including:

[0037] Initialize an adaptive filter, wherein the error signal of the adaptive filter is set as the difference between the microphone received signal and the estimated signal used to predict the speaker feedback signal;

[0038] The coefficients of the adaptive filter are updated based on a preset adaptive algorithm to determine the minimized error signal;

[0039] The minimized error signal is used as the new microphone reception signal after eliminating the speaker feedback signal.

[0040] According to a speaker voice enhancement method provided by the present invention, before eliminating the speaker feedback signal based on an adaptive feedback control algorithm, the method further includes:

[0041] The microphone received signal is input to a hardware equalizer to adjust the gain of abnormal frequency bands in the microphone received signal.

[0042] According to a speaker voice enhancement method provided by the present invention, after obtaining the target audio signal based on the target frequency domain signal, the method further includes:

[0043] The target audio signal is subjected to frequency shifting and phase modulation processing and / or half-wave rectification processing;

[0044] The frequency shifting and phase modulation process includes increasing or decreasing the overall spectrum of the target audio signal by a preset Hertz, and periodically changing the phase of the target audio signal;

[0045] The half-wave rectification process includes converting all sample values ​​less than zero in the time-domain waveform of the target audio signal into sample values ​​greater than zero by taking their absolute values.

[0046] According to a speaker voice enhancement method provided by the present invention, after performing frequency shifting and phase modulation processing and / or half-wave rectification processing on the target audio signal, the method further includes:

[0047] The amplitude of the target audio signal is dynamically adjusted to control the dynamic range of the speaker output signal.

[0048] The present invention also provides a speaker voice enhancement device, comprising the following modules:

[0049] The signal acquisition unit is used to acquire the initial spectrum signal corresponding to the initial audio signal and determine the initial power spectrum signal corresponding to the initial spectrum signal.

[0050] A signal processing unit is used to input the initial power spectrum signal into a feedback cancellation network model and obtain the target frequency domain signal output by the feedback cancellation network model. The feedback cancellation network model includes a first sub-model and a second sub-model.

[0051] A signal generation unit is used to obtain a target audio signal based on the target frequency domain signal.

[0052] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speaker voice enhancement method as described above.

[0053] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speaker speech enhancement method as described above.

[0054] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speaker voice enhancement method as described above.

[0055] The speaker voice enhancement method, apparatus, electronic device, and storage medium provided by this invention first use a real number signal mask to adjust the frequency domain amplitude of the initial spectrum signal to generate a frequency domain estimation signal. Then, a complex number signal mask is used to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhancement signal. Finally, the target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal, thereby enabling the output target frequency domain signal to have higher fidelity in both amplitude and phase, maximizing the removal of environmental noise and speaker feedback signals, and effectively suppressing howling. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of the speaker and microphone provided by the present invention receiving sound signals.

[0058] Figure 2 This is a schematic flowchart of the speaker voice enhancement method provided by the present invention.

[0059] Figure 3 This is a schematic diagram of the model architecture of the first sub-model provided by the present invention.

[0060] Figure 4 This is a schematic diagram of the model architecture of the second sub-model provided by the present invention.

[0061] Figure 5 This is a schematic diagram of the architecture of the upsampling convolutional layer provided by the present invention.

[0062] Figure 6This is a schematic diagram of the architecture of the downsampling convolutional layer provided by the present invention.

[0063] Figure 7 This is a schematic diagram of the speaker voice enhancement device provided by the present invention.

[0064] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0066] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0067] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more.

[0068] The following is combined Figures 1-8 This invention describes the speaker voice enhancement method, apparatus, electronic device, and storage medium provided by the present invention.

[0069] Figure 1 This is a schematic diagram of the speaker and microphone receiving sound signals provided by the present invention, as shown below. Figure 1 As shown, the microphone receives signals in a loudspeaker scenario. It can be divided into three components: 1) target speech 1) Direct speech from the speaker; 2) Feedback signal 3) Background noise: The speaker plays a direct signal; room reverberation The noise in the environment, along with the reflection and delay of the speaker's voice and speaker signals within the room, causes sound overlap and distortion.

[0070] like Figure 1 As shown, the speaker plays a reference signal of... The feedback signal transmitted back to the microphone is then played through the speaker. As shown in formula (1):

[0071] (1)

[0072] in, Indicates a feedback signal. This indicates the nonlinear transformation of the loudspeaker. This indicates that the speaker is playing a reference signal. Indicates the propagation path from the speaker to the microphone. This represents linear convolution.

[0073] Feedback signal Signal is received by microphone The signal is obtained after delay and amplification, and is a feedback signal. It will be continuously picked up by the microphone at any time. Corresponding microphone receiving signal As shown in formula (2):

[0074] (2)

[0075] in, This indicates that the microphone is receiving signals. Indicates the target speech. Indicates background noise. Indicates room reverberation. This indicates the nonlinear transformation of the loudspeaker. This indicates the transmission delay from the microphone to the speaker. This indicates the total time for the feedback signal to be amplified. Indicates the speaker amplification gain value. This indicates the propagation path from the speaker to the microphone.

[0076] From formula (2), it can be seen that the feedback signal Repeated pickups can create positive feedback, causing the signal to be amplified cyclically at certain frequencies, which in turn produces a howling phenomenon.

[0077] Therefore, the present invention provides a speaker speech enhancement method, the main objective of which is:

[0078] On the one hand, it suppresses feedback signals and reverberation components in microphone pickup to improve the clarity of the target speech; on the other hand, it reduces link latency to ensure that the time difference between the speaker's direct speech and the speaker's output signal is minimized, thus avoiding echo problems during the amplification process.

[0079] Figure 2 This is a flowchart illustrating the speaker voice enhancement method provided by the present invention, as shown below. Figure 2 As shown, the entity executing the speaker voice enhancement method provided by this invention can be a speaker, amplifier, voice interaction device, smart speaker, electronic whiteboard with audio enhancement function, vehicle voice system, or other electronic device with voice playback and processing function. Unless otherwise specified, the following embodiments will use a speaker as an example for description.

[0080] As an optional embodiment, the implementation steps of this loudspeaker voice enhancement method mainly include, but are not limited to:

[0081] First, obtain the initial spectrum signal corresponding to the initial audio signal, and determine the initial power spectrum signal corresponding to the initial spectrum signal.

[0082] Specifically, the initial audio signal can be a time-domain speech signal captured by a microphone, which inevitably includes the target speech, background noise, room reverberation, and feedback signals from the speakers.

[0083] It should be noted that time-domain speech signals refer to the waveform curve of sound changing over time, recording how the sound fluctuates from the first second to the last second. Considering that time-domain speech signals cannot be directly analyzed in the frequency domain or separated from noise, this invention converts them into frequency-domain signals to facilitate further processing, such as separating the target speech from background noise and suppressing feedback signals.

[0084] Specifically, the initial spectrum signal refers to the frequency domain signal obtained by processing the initial audio signal through a Short-Time Fourier Transform (STFT). The initial spectrum signal can be obtained through the following steps:

[0085] (1) The continuous initial audio signal is divided into frames, that is, it is divided into multiple frames of audio signal according to the preset frame length and frame shift.

[0086] (2) Window each frame of audio signal (e.g., using Hamming or Hanning window) to reduce spectral leakage.

[0087] (3) Perform a Fast Fourier Transform (FFT) on each frame of the audio signal after windowing to obtain the frequency domain representation of each frame.

[0088] (4) Finally, the frequency domain of all frames is spliced ​​together to form the initial spectrum signal.

[0089] The initial spectral signal obtained after the above processing is a complex signal, which completely preserves the amplitude and phase information of the initial audio signal. For example, the initial spectral signal can be used... To indicate, among which This indicates two channels: the real part and the imaginary part. Indicates the index of the time frame. Indicates the index of the frequency point.

[0090] By reasonably setting the window length, frame shift, sampling rate, and processing delay, this invention can ensure that there is almost no obvious time difference between the playback and reception of the voice signal while meeting the requirements of real-time audio processing, thus avoiding echo or delay.

[0091] For example, in a preferred embodiment, the sampling rate can be set to 16000Hz, the synthesis window length to 256 sampling points, and the frame shift to 80 sampling points. Based on these parameters, the processing latency is (window length + frame shift) / sampling rate = (256 + 80) / 16000 = 0.021 seconds, or 21 milliseconds. Such low latency makes the time difference between the speaker's direct sound and the sound played through the speaker almost imperceptible to the human ear, thus effectively avoiding the echo and delay commonly found in amplification, and significantly improving the user experience.

[0092] Specifically, determining the initial power spectrum signal corresponding to the initial spectrum signal means calculating based on the acquired initial spectrum signal to extract its energy distribution characteristics at different time frames and frequency points, which are then used as inputs for the subsequent first sub-model.

[0093] Specifically, the obtained initial power spectrum signal is a real number signal, which is calculated based on the amplitude information of the initial spectrum signal. It reflects the energy intensity of the initial audio signal at each time frame and frequency point, but does not contain phase information. Its calculation method can be shown in formula (3):

[0094] (3)

[0095] in, Represents the initial power spectrum signal. and These represent the initial spectral signal at the 1st... The time frame, the Two channels, one for the real part and one for the imaginary part, at each frequency point. Indicates the first The time frame, the The total power of the two channels, real and imaginary, at each frequency point This means converting the calculated signal power into decibels (dB) in logarithmic form to match the human ear's perception of loudness changes, which also helps in the stable training and learning of subsequent neural network models.

[0096] Then, the initial power spectrum signal is input into the feedback cancellation network model to obtain the target frequency domain signal output by the feedback cancellation network model. The feedback cancellation network model mainly includes the first sub-model and the second sub-model.

[0097] The feedback cancellation network model is the core processing unit of this invention, and can be a deep neural network model trained on a large amount of data. The main goal of the feedback cancellation network model is to intelligently estimate and suppress unwanted signal components (such as ambient noise, room reverberation, and speaker feedback signals) from the input initial audio signal, which includes the target speech, background noise, room reverberation, and speaker feedback signals, while preserving and enhancing the desired target speech, and finally generating the target frequency domain signal.

[0098] The first sub-model outputs a real-valued signal mask based on the input initial power spectrum signal.

[0099] Specifically, the first sub-model consists of a specific neural network layer. Its design goal is to learn and extract the feature information related to the target speech by performing complex nonlinear transformations and time-series modeling on the input initial power spectrum signal, and to distinguish the feature information of interference components such as noise and reverberation.

[0100] The first sub-model ultimately outputs a real-valued signal mask. This mask is used to adjust the frequency domain amplitude of the initial spectral signal to obtain a frequency domain estimated signal. It is a matrix or tensor with the same dimensions (i.e., the same number of time frames and frequency points) as the initial power spectrum signal, where each element is a real value. For example, these real values ​​are constrained to the range of 0 to 1. In this case, values ​​close to 1 in the mask indicate that the corresponding time-frequency unit is judged by the model to mainly contain target speech energy and should be preserved or enhanced; while values ​​close to 0 indicate that the time-frequency unit is judged to mainly contain noise or reverberation interference energy and should be suppressed. Therefore, this real-valued signal mask can be understood as an energy-based time-frequency gain filter.

[0101] After obtaining the real-valued signal mask, it can be used to adjust the frequency domain amplitude of the initial spectral signal. Specifically, the mask adjustment involves element-wise multiplying the real-valued signal mask with the amplitude spectrum of the initial spectral signal. It should be noted that the initial spectral signal is a complex signal, and the real-valued signal mask only affects the amplitude portion of the initial spectral signal without altering its original phase information.

[0102] Specifically, the input of the first sub-model is the initial power spectrum signal, and the output is a real number signal mask. The specific calculation method for obtaining the frequency domain estimated signal by masking the frequency domain amplitude of the initial spectrum signal using the real number signal mask is shown in formula (4):

[0103] (4)

[0104] in, This represents the frequency domain estimated signal. The channel represents the imaginary part of the initial spectral signal. This represents a real number signal mask.

[0105] The frequency domain estimated signal obtained after the above mask adjustment has the same phase information as the initial spectrum signal, but its amplitude spectrum has been corrected by the real signal mask. Thus, the frequency domain amplitude dominated by the target speech component is preserved, while the frequency domain amplitude dominated by noise and reverberation components is effectively attenuated.

[0106] The second sub-model receives a three-dimensional spectrum signal obtained by splicing the frequency domain estimation signal and the initial spectrum signal, and outputs a complex signal mask.

[0107] Considering that the original initial spectrum signal contains complete, unprocessed mixed speech information, including the target speech and all types of interference, and that the frequency domain estimation signal is the result of preliminary noise reduction and dereverberation processing by the first sub-model, it can be regarded as a preliminary estimate of the clean target speech. Therefore, this invention obtains a three-dimensional spectrum signal by merging the frequency domain estimation signal and the initial spectrum signal. The resulting three-dimensional spectrum signal can provide the second sub-model with the key information carried by the initial spectrum signal and the frequency domain estimation signal respectively.

[0108] As an optional embodiment, the process of merging the frequency domain estimation signal and the initial spectrum signal to obtain a three-dimensional spectrum signal specifically involves splicing the real and imaginary channels of the initial spectrum signal with the amplitude spectrum channel of the frequency domain estimation signal in the channel dimension, thereby obtaining an input tensor with three channels, i.e., a three-dimensional spectrum signal.

[0109] When making decisions on three-dimensional spectral signals, the second sub-model can not only analyze the original, interference-laden initial spectral signal, but also refer to a frequency domain estimation signal that has been preliminarily cleaned up. This allows it to learn more accurately how to further suppress residual interference components, especially feedback and reverberation signals that are similar in energy to the target speech but differ in phase, and output a complex signal mask.

[0110] It is important to emphasize that the complex signal mask output by the second sub-model is fundamentally different from the real signal mask output by the first sub-model. It is a complex tensor with the same dimension as the initial spectral signal, and each element of it contains both a real and an imaginary part. Therefore, it is possible to simultaneously adjust the frequency domain amplitude and phase of a complex signal.

[0111] After obtaining the complex signal mask, the next step is to adjust the frequency domain amplitude and phase of the initial spectral signal using a mask. The complex signal mask is used to adjust the frequency domain amplitude and phase of the initial spectral signal to obtain a frequency-enhanced signal. The mask adjustment is achieved through complex multiplication, that is, multiplying the complex signal mask element-wise with the original initial spectral signal.

[0112] It is important to note that the complex mask here is applied to the initial spectral signal, rather than to the frequency domain estimation signal output from the first stage. This avoids the accumulation of errors and allows the second stage to directly correct the original signal, thereby obtaining a more accurate enhancement result.

[0113] As an optional embodiment, a complex signal mask is applied to the initial spectral signal to obtain a frequency domain enhanced signal. The specific calculation method is shown in formulas (5) and (6):

[0114] (5)

[0115] (6)

[0116] in, Represents a three-dimensional spectrum signal. The channel represents the imaginary part of the initial spectral signal. This represents the frequency domain estimated signal. Indicates frequency domain enhancement signal, This represents a complex signal mask.

[0117] Compared with the frequency domain estimation signal output in the first stage, the final frequency domain enhanced signal is not only optimized in terms of frequency domain amplitude (i.e., the energy of noise, reverberation, feedback and other interferences is suppressed), but its phase is also corrected, making it closer to the phase of the clean target speech.

[0118] The following is a brief introduction to the specific training process of the feedback elimination network model.

[0119] The training samples for the feedback cancellation network model can be multiple audio signal samples collected by a microphone. These audio signal samples also include speech signals, ambient noise, room reverberation, and speaker feedback signals. Performing a short-time Fourier transform on the audio signal samples yields the corresponding spectral signal samples. Further, based on the amplitude information of the spectral signal samples, a power spectrum signal sample is calculated and used as the input sample data for the feedback cancellation network model.

[0120] Furthermore, the training labels for the feedback cancellation network model can be labeled as frequency domain signal labels corresponding to the audio signal sample. It should be noted that the frequency domain signal label is typically a theoretically clean speech spectrum that has been preprocessed or recorded under controlled conditions. It represents the frequency domain representation of the clean target speech after completely removing environmental noise, room reverberation, and feedback signals. This frequency domain signal label serves as a supervisory signal, guiding the feedback cancellation network model during training. By adjusting the relevant model parameters of the first and second sub-models, it learns effective frequency domain amplitude and phase adjustment strategies, thereby achieving precise suppression of noise and feedback signals.

[0121] As an optional implementation, the training process of the feedback cancellation network model mainly includes, but is not limited to, the following steps:

[0122] (1) Training data preparation: Acquire audio signal samples collected by the microphone and perform short-time Fourier transform on them to obtain complex-form spectral signal samples; then calculate the power based on the real and imaginary parts of the spectral signal samples to generate power spectrum signal samples, which are used as input to the network model. At the same time, prepare paired frequency domain signal labels representing the pure target speech as training labels. Furthermore, the input data can be normalized to improve data consistency and model learning stability.

[0123] (2) Data augmentation: In order to improve the robustness and generalization ability of the model, data augmentation techniques can be used to expand the diversity of training samples. For example, the audio signal samples can be randomly shifted along the time axis, the frequency scale of the spectrum signal samples can be scaled, or different types and intensities of background noise and simulated reverberation can be superimposed on the signal to enhance the model's adaptability to interference factors in various real-world application environments.

[0124] (3) Loss function design: Construct a loss function for supervising model training. This loss function is used to measure the difference between the model output and the training label (the ideal target frequency domain signal). For example, the traditional mean squared error (MSE) can be combined with one or more auditory perception-based evaluation metrics (such as loss related to speech quality perception assessment) to optimize the overall auditory quality of the final output audio while optimizing the frequency domain recovery accuracy.

[0125] (4) Backpropagation and parameter optimization: During the training iteration, the gradient is calculated based on the error calculated by the loss function through the backpropagation algorithm, and an efficient optimization algorithm (such as the Adam optimization algorithm) is used to update all learnable parameters in the feedback elimination network model.

[0126] In addition, a learning rate decay mechanism can be used to accelerate convergence in the early stages of training and to help the model converge to a better point in the later stages of training. This helps to improve the convergence speed of the training process and reduce the risk of overfitting, thereby improving the stability and performance of the model in practical applications.

[0127] The following is a brief explanation of the implementation steps for determining the target frequency domain signal based on the frequency domain estimation signal and the frequency domain enhancement signal, and for obtaining the target audio signal based on the target frequency domain signal.

[0128] First, the frequency domain estimated signal and the frequency domain enhanced signal are the products of two different processing strategies, each with its own advantages. The frequency domain estimated signal effectively suppresses most noise and reverberation while preserving the original phase through real-number masking, but it may not be able to handle feedback signals that are highly correlated with the target speech. On the other hand, the frequency domain enhanced signal jointly optimizes amplitude and phase through complex-number masking, which can handle residual interference more finely, but may also introduce unnecessary processing traces. Therefore, the process of determining the target frequency domain signal is a process of fusing the frequency domain estimated signal and the frequency domain enhanced signal.

[0129] As an alternative implementation, the final target frequency domain signal can be determined by weighted fusion of the frequency domain estimated signal and the frequency domain enhanced signal. For example, the feedback cancellation network model can output an additional time-varying weighting coefficient, which dynamically adjusts the proportion of the two signals in the final output. In certain time frames, when the output quality of the first sub-model is higher, it can be assigned a larger weight; conversely, the output of the second sub-model can be assigned a larger weight. This adaptive fusion strategy ensures that the final target frequency domain signal achieves optimal enhancement under various conditions.

[0130] As an optional embodiment, the target frequency domain signal can be calculated as shown in formula (7):

[0131] (7)

[0132] in, The target frequency domain signal after fusion. These are weighting coefficients. This represents the frequency domain estimation signal in the first stage. This represents the frequency domain enhancement signal in the second stage.

[0133] Furthermore, after determining the target frequency domain signal, the step of obtaining the target audio signal based on the target frequency domain signal is a process of converting the signal from the frequency domain back to the time domain. The purpose is to generate an enhanced time-domain waveform signal that can be directly played by a speaker. This step can be implemented using the Inverse Short-Time Fourier Transform (ISTFT).

[0134] Specifically, the implementation steps of the inverse short-time Fourier transform can be as follows:

[0135] (1) Perform an inverse fast fourier transform (IFFT) on each frame of the target frequency domain signal to convert it from the frequency domain representation back to the time domain waveform.

[0136] (2) All short-time waveforms obtained by inverse transformation are spliced ​​and synthesized by overlapping and adding to recover a continuous and complete time-domain audio signal, which is the target audio signal.

[0137] The speaker voice enhancement method provided in this embodiment first uses a real signal mask to adjust the frequency domain amplitude of the initial spectrum signal to generate a frequency domain estimation signal. Then, it uses a complex signal mask to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhancement signal. Finally, the target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal, thereby enabling the output target frequency domain signal to have higher fidelity in both amplitude and phase, which can remove environmental noise and speaker feedback signals to the greatest extent and effectively suppress howling.

[0138] Figure 3 This is a schematic diagram of the model architecture of the first sub-model provided by the present invention, as shown below. Figure 3 As shown, in one optional embodiment, the first sub-model includes a first input module, a first GRU module, a second GRU module, a splicing module, a third GRU module, and a first output module.

[0139] In practical applications, the first input module receives the initial power spectrum signal and inputs the initial power spectrum signal in parallel to the first GRU module and the second GRU module.

[0140] The first input module serves as the data entry point for the first sub-model, and its function is to receive the initial power spectrum signal calculated in the aforementioned steps. For example... Figure 3As shown, the initial power spectrum signal is simultaneously and indiscriminately distributed by the first input module to the first GRU module and the second GRU module. This parallel processing structure allows the model to analyze the same input data using different parameter sets, thereby extracting more diverse features.

[0141] The following is a brief introduction to the implementation process of the two parallel GRU modules processing the input signals.

[0142] The first GRU module performs time-series modeling on the initial power spectrum signal to obtain the first time-series feature; the second GRU module performs time-series modeling on the initial power spectrum signal to obtain the second time-series feature.

[0143] A Gated Recurrent Unit (GRU) is a special type of recurrent neural network that excels at capturing temporal dependencies in sequential data, such as time-frame sequences of audio signals. Through its internal update and reset gate mechanisms, the GRU learns which information needs to be retained and which needs to be forgotten in the temporal dimension.

[0144] Specifically, time-series modeling refers to the GRU module processing the initial power spectrum signal along the time frame dimension and analyzing the correlation between the spectral characteristics of the current time frame and the characteristics of the preceding and following time frames.

[0145] For example, the first GRU module and the second GRU module can be configured to have different internal parameters or structures, such as different numbers of hidden layer units, or one can be a unidirectional GRU and the other a bidirectional GRU. With such a configuration, they can learn temporal context information at different scales or in different directions, thereby generating two sets of complementary but distinct features, namely the first temporal feature and the second temporal feature.

[0146] After obtaining two sets of parallel temporal features, the splicing module splices the first and second temporal features along the feature dimension to obtain the fused temporal features.

[0147] Specifically, the concatenation module performs a feature fusion operation along the feature dimension. Assuming both the first and second temporal features have a dimension of 128, the dimension of the fused temporal feature after concatenation will become 256. Through feature fusion, the diverse temporal information learned from the two parallel branches can be integrated to form a more comprehensive and information-rich feature representation.

[0148] Then, the fused temporal features obtained after the above feature fusion are further processed in depth. That is, the third GRU module receives the fused temporal features and performs temporal dependency modeling on the fused temporal features to obtain the final temporal features. This can be understood as integrating and abstracting the fused temporal features that have been initially refined, thereby capturing longer-term and more complex temporal correlations.

[0149] Finally, the final timing features are mapped to a real-valued signal mask through the first output module. The first output module can consist of one or more fully connected layers and an activation function (such as the sigmoid function). It receives the final timing features output from the third GRU module and maps them to the same magnitude as the initial power spectrum signal through a linear transformation. The sigmoid activation function in the first output module constrains the output value between 0 and 1, making its physical meaning consistent with the gain mask. This yields a real-valued signal mask that can be directly used to mask the frequency domain amplitude of the initial spectrum signal.

[0150] The speaker speech enhancement method provided by this invention, through a structure that includes a parallel GRU module and a serial GRU module connected together, enables the first sub-model to effectively perform deep temporal context modeling of the initial power spectrum signal. This structural design not only extracts diverse temporal features through parallel branches, but also integrates these features through subsequent splicing and deep modeling, thereby enabling a more accurate estimation of the energy distribution differences between the target speech and interference, and finally generating a high-quality real-number signal mask, providing strong support for achieving efficient noise and reverberation suppression.

[0151] Figure 4 This is a schematic diagram of the model architecture of the second sub-model provided by the present invention, as shown below. Figure 4 As shown, the second sub-model mainly includes a second input module, an encoder module, a bottleneck processing module, a decoder module, and a second output module.

[0152] Specifically, the second input module serves as the data entry point for the second sub-model. It receives a three-dimensional spectrum signal composed of the initial spectrum signal and the frequency domain estimation signal spliced ​​along the channel dimension, and then transmits the three-dimensional spectrum signal to the subsequent encoder module for feature extraction.

[0153] Next, the encoder module performs layer-by-layer analysis on the input 3D spectral signal. The encoder module functions as a feature extractor, and its internal structure mainly consists of multiple cascaded upsampling convolutional layers. In each upsampling convolutional layer, local features are extracted through convolution operations, while pooling operations are used to reduce the resolution of the feature map in the time and frequency dimensions and increase the number of feature channels.

[0154] For example, after the input 3D spectral signal is processed by the first upsampling convolutional layer, the number of channels increases from 3 to 8; after the second upsampling convolutional layer, the number of channels can increase to 16, and so on. This process can capture signal features from concrete to abstract step by step. The features output by each upsampling convolutional layer in the encoder module are intermediate encoded features, which are saved for skip connections in subsequent decoder modules.

[0155] The bottleneck processing module is used to perform temporal context modeling on the intermediate encoded features output by the last upsampled convolutional layer in the encoder module, so as to perform filtering and weighting processing of adjacent time frames and output the temporally enhanced bottleneck features.

[0156] Specifically, this bottleneck processing module is located at the connection between the encoder and decoder modules, processing the intermediate encoded features with the lowest resolution but the richest semantic information. The core of this module is a recurrent neural network unit (RNN), whose temporal context modeling is achieved by unfolding the input feature map along the time frame dimension and feeding it into a GRU or bidirectional GRU. The filtering and weighting processing of adjacent time frames is accomplished through the gating mechanism of the GRU or bidirectional GRU. This mechanism learns the dependencies between the features of the current time frame and the features of the preceding and following time frames, and performs weighted fusion of these dependencies, ultimately outputting a bottleneck feature enhanced with temporal information.

[0157] The decoder module comprises multiple cascaded downsampling convolutional layers, its function being to reconstruct the bottleneck features output by the bottleneck processing module. In each downsampling convolutional layer, transpose convolution is used to improve the resolution of the feature map in the time and frequency dimensions and reduce the number of channels, thereby gradually recovering a feature map with the same resolution as the input signal. This decoder module uses the bottleneck features output by the bottleneck processing module as initial input, decodes them layer by layer through multiple downsampling convolutional layers, and finally outputs a frequency-domain decoded signal.

[0158] Finally, the second output module performs convolution mapping on the frequency domain decoded signal to output a complex signal mask.

[0159] Specifically, the second output module can be a 1×1 convolutional layer that receives the frequency domain decoded signal output by the decoder module. By performing the final convolution mapping, the number of channels of the frequency domain decoded signal is mapped to 2, which correspond to the real part and the imaginary part of the complex signal mask, respectively, thereby obtaining the final complex signal mask.

[0160] The speaker speech enhancement method provided by this invention utilizes an encoder module to extract intermediate coded features from a three-dimensional spectral signal, then a bottleneck processing module performs time-series dependency modeling on the intermediate coded features, outputting bottleneck features enhanced by time-series information, and finally a decoder module reconstructs the bottleneck features output by the bottleneck processing module into a high-resolution mask, providing strong model support for generating high-quality complex signal masks, thereby significantly improving the performance of speech enhancement.

[0161] Figure 5 This is a schematic diagram of the architecture of the upsampling convolutional layer provided by the present invention. This upsampling convolutional layer is the basic unit constituting the encoder module. Figure 5 As shown, the upsampling convolutional layer starts from the input layer, and its core part mainly consists of two sequentially cascaded convolutional-batch normalization layers and an activation function.

[0162] Specifically, the processing flow of the upsampling convolutional layer includes, but is not limited to, the following steps:

[0163] (1) The input layer receives the input feature map and passes it to the first cascaded convolution-batch normalization layer. The convolutional layer performs convolution operations on the input feature map through its internal convolution kernel to extract local features. Then, the batch normalization layer normalizes the feature map output by the convolution to adjust the data distribution, thereby improving the stability and convergence speed of the model training.

[0164] (2) The feature map after being processed by the first cascaded convolution-batch normalization layer is then fed into the second cascaded convolution-batch normalization layer, which also includes a convolutional layer and a batch normalization layer, aiming to perform deeper feature extraction.

[0165] (3) After the feature map is processed by the second cascaded convolution-batch normalization layer, the feature map is further transformed by the ReLU activation function to set all negative elements to zero and retain only positive elements, thereby enhancing the expressive power of the model and helping to alleviate the gradient vanishing problem.

[0166] Figure 6 This is a schematic diagram of the architecture of the downsampling convolutional layer provided by the present invention, as shown below. Figure 6 As shown, the downsampling convolutional layer mainly includes an input layer, a convolutional layer, a batch normalization layer, a transposed convolutional layer, a second batch normalization layer, a ReLU activation function layer, and an output layer.

[0167] Specifically, the processing flow of the downsampling convolutional layer includes, but is not limited to, the following steps:

[0168] (1) The input layer receives the feature map from the previous layer and passes it to the convolutional layer for preliminary feature extraction. The convolutional layer performs convolution operations on the input feature map through multiple convolutional kernels to extract local feature information.

[0169] (2) The feature map output by the convolutional layer is passed to the batch normalization layer. The batch normalization layer normalizes the feature map, adjusts the data distribution, and reduces gradient fluctuations during training, thereby improving the stability and convergence speed of model training.

[0170] (3) The normalized feature map is input into the transposed convolutional layer. The transposed convolutional layer performs spatial downsampling of the feature map, improves the temporal and frequency dimension resolution of the feature map, and provides conditions for subsequent fine feature recovery.

[0171] (4) The output of the transposed convolutional layer undergoes a second batch normalization process to further standardize the data distribution and ensure the effectiveness of subsequent nonlinear transformations.

[0172] (5) Subsequently, the feature map is transformed nonlinearly through the ReLU activation function layer, setting all negative values ​​to zero and retaining only positive elements, thereby enhancing the nonlinear expressive power of the model and effectively alleviating the gradient vanishing problem.

[0173] (6) Finally, the feature map processed as described above is output by the output layer and passed to the next processing unit of the decoder module as the processing result of the downsampling convolutional layer.

[0174] In another embodiment of the present invention, the bottleneck processing module includes a gated loop unit, a bidirectional gated loop unit, and a linear mapping layer; it receives the intermediate coding features output by the last upsampled convolutional layer of the encoder module, performs temporal modeling on the features along the time frame dimension, and outputs the first temporal feature.

[0175] Specifically, this intermediate encoded feature has the lowest time and frequency dimension resolution but the highest feature channel dimension, containing the most abstract high-level semantic information. The gated recurrent unit iteratively processes the feature map along the sequence direction of the time frame, using its internal gating mechanism to learn and capture the forward dependencies of the signal in time. After this unidirectional temporal processing, the first temporal feature is output.

[0176] Next, a deeper modeling is performed on the preliminary timing features. The bidirectional gated loop unit receives the first timing features, performs bidirectional timing modeling to execute filtering and weighting processing of adjacent time frames, and outputs the second timing features.

[0177] Specifically, the bidirectional gated recurrent unit (GRU) consists of a forward GRU and a backward GRU. After receiving the first temporal feature, the forward GRU processes the sequence from beginning to end, while the backward GRU processes it from end to beginning. This bidirectional temporal modeling allows the second sub-model to utilize both past and future information when processing any given time frame. Through this bidirectional mechanism, the second sub-model can weight and filter the features of each time frame based on complete contextual information, thereby more effectively identifying and modeling signal components with specific temporal evolution patterns. After processing, the second temporal feature output by the bidirectional gated recurrent unit contains richer and more comprehensive temporal contextual information.

[0178] Finally, the linear mapping layer performs a mapping transformation on the second temporal feature, outputting the temporally enhanced bottleneck feature.

[0179] Specifically, the linear mapping layer can be a fully connected layer that receives a second temporal feature from the output of a bidirectional gated recurrent unit. This linear mapping layer performs a linear transformation on the second temporal feature, the parameters of which are learned by the second sub-model during training.

[0180] The speaker speech enhancement method provided by this invention, through a bottleneck processing module composed of a gated loop unit, a bidirectional gated loop unit, and a linear mapping layer cascaded in series, can perform comprehensive temporal analysis on the intermediate encoded features extracted by the encoder, from shallow to deep and from unidirectional to bidirectional. This structural design enhances the model's ability to capture and understand the temporal dynamics of the signal, and has significant advantages, especially in processing feedback and reverberation signals that are strongly correlated with the target speech in time. This provides a solid foundation for the decoder module to accurately reconstruct the signal and ultimately generate a high-quality complex signal mask.

[0181] In another embodiment provided by the present invention, the downsampling convolutional layer in the decoder module corresponds one-to-one with the upsampling convolutional layer in the encoder module.

[0182] Specifically, this one-to-one correspondence is the basis for constructing a U-shaped network symmetrical structure. That is, a layer in the encoder module that is at a certain resolution level must have a corresponding layer in the decoder module that is also at that resolution level.

[0183] During the process of downsampling and convolution of bottleneck features layer by layer in the decoder module to output frequency domain decoded signals, intermediate encoded features with the same resolution output from the corresponding upsampling convolutional layer are fused through skip connections at each downsampling convolutional layer.

[0184] Specifically, the fusion process may include, but is not limited to, the following steps:

[0185] (1) A downsampling convolutional layer in the decoder module first performs a downsampling operation on the feature map received from the previous layer to improve its time and frequency resolution.

[0186] (2) By skipping connections, intermediate coding features of the corresponding level in the encoder module with the same resolution are directly passed to the downsampling convolutional layer.

[0187] (3) The downsampled feature map and the intermediate encoded features passed through the skip connection are concatenated in the channel dimension to achieve feature fusion.

[0188] The speaker speech enhancement method provided by this invention employs skip connections to fuse coded features with the same resolution, allowing the decoder module to directly utilize the precise location and detail information captured by the encoder module during upsampling when reconstructing the signal. This enables the final generated complex signal mask to more accurately recover the spectral structure of the target speech, effectively avoiding speech distortion or blurring caused by information loss, thereby significantly improving the quality of the frequency domain enhanced signal and the clarity and fidelity of the final output target audio signal.

[0189] In another embodiment of the present invention, the feedback cancellation network model further includes a weighted fusion module; the weighted fusion module performs weighted fusion on the frequency domain estimation signal output by the first sub-model and the frequency domain enhancement signal output by the second sub-model to generate the target frequency domain signal.

[0190] Considering that the frequency domain estimation signal output by the first sub-model mainly suppresses interference by adjusting the frequency domain amplitude, it retains the original phase; while the frequency domain enhancement signal output by the second sub-model jointly adjusts the frequency domain amplitude and phase through complex masking, which can more finely process residual reverberation and feedback signals, but may also introduce imperceptible processing distortion.

[0191] Therefore, the present invention provides this weighted fusion module to comprehensively utilize the processing advantages of the two sub-models.

[0192] Specifically, this weighted fusion process can be controlled by one or more learnable weighting coefficients. These weighting coefficients can be learned and output by the feedback elimination network model itself, for example, as an additional output of a second sub-model. The weighting coefficients can be dynamic values ​​that vary over time frames, used to dynamically adjust the proportion of the two signals in the final output.

[0193] This invention takes into account the frequency domain estimation signal output by the first sub-model, which mainly suppresses interference by adjusting the frequency domain amplitude. Its advantage is that it preserves the original phase. The frequency domain enhancement signal output by the second sub-model is also considered. Therefore, the frequency domain amplitude and phase are jointly adjusted by using a complex mask, which can more finely process residual reverberation and feedback signals. However, it may sometimes introduce imperceptible processing distortion.

[0194] The speaker speech enhancement method provided by this invention avoids the limitations of a single model output by fusing the outputs of two stages using a weighted fusion module. This design enables the feedback cancellation network model to dynamically combine the processing results of the two sub-models according to the signal content, thereby generating a target frequency domain signal that achieves a better balance in terms of speech clarity, noise suppression, feedback cancellation, and naturalness of sound, further enhancing the stability and robustness of the entire speaker speech enhancement system.

[0195] In another embodiment of the present invention, the initial audio signal is obtained by: acquiring the microphone received signal, and sampling the microphone received signal according to a preset synthesis window length and a preset frame shift to obtain the initial audio signal.

[0196] Specifically, the microphone received signal is the raw digital audio stream captured by the microphone hardware in the speaker system and obtained after analog-to-digital conversion. It should be noted that the microphone received signal is a continuous time-domain signal stream, which reflects all sound components of the sound field around the microphone in real time, including the target speech, background ambient noise, reverberation caused by room reflections, and feedback signals generated by the speaker playback.

[0197] Next, in order to process this continuous time-domain signal stream, the microphone received signal needs to be segmented into frames according to the preset synthesis window length and preset frame shift.

[0198] Specifically, this framing process divides a continuous signal stream into multiple overlapping data blocks, i.e., audio frames.

[0199] The preset synthesis window length defines the length of each audio frame in terms of the number of sampling points, which determines the frequency resolution for subsequent frequency domain analysis.

[0200] The preset frame shift defines the distance between the starting points of two adjacent audio frames, also expressed in the number of sample points. This preset frame shift is less than the synthesis window length, resulting in an overlapping area between adjacent frames. This overlapping design ensures that information at frame boundaries is not lost during framing and helps to obtain smoother audio during signal reconstruction.

[0201] For example, in a scenario with a sampling rate of 16000Hz, the preset synthesis window length can be set to 256 sampling points, and the preset frame shift can be set to 80 sampling points. The speaker speech enhancement method provided by this invention, by employing this method of segmenting the microphone received signal based on the synthesis window length and frame shift, successfully transforms a continuous, infinitely long audio stream into a discrete, serialized initial audio signal sequence. This processing method is the foundation for realizing real-time streaming speech enhancement, enabling instant analysis and enhancement of incoming audio data without waiting for the entire speech segment to end. This ensures the low-latency characteristics of the speaker speech enhancement method and provides a smooth, echo-free user experience for real-time sound reinforcement scenarios.

[0202] In another embodiment of the present invention, the microphone received signal includes at least the target speech signal and the speaker feedback signal; after the microphone received signal is acquired, the speaker feedback signal is eliminated based on an adaptive feedback control algorithm, specifically including: initializing an adaptive filter, wherein the error signal of the adaptive filter is set as the difference between the microphone received signal and the estimated signal for predicting the speaker feedback signal.

[0203] It should be noted that this application uses an adaptive feedback control (AFC) algorithm to process signals such as speaker feedback that are linearly related to the reference signal (i.e., the signal played by the speaker).

[0204] Specifically, the execution process of the AFC algorithm includes, but is not limited to, the following steps:

[0205] (1) Initialize the adaptive filter and set the error signal. The adaptive filter is a digital filter whose coefficients can be dynamically adjusted according to the statistical characteristics of the input signal. The coefficients can be set to zero or a small random value during initialization. The AFC algorithm requires a reference signal, i.e., the signal being sent to the speaker for playback. Filtering this reference signal through the current adaptive filter yields an estimated signal for predicting the speaker feedback signal. The error signal is set as the difference between the real-time microphone reception signal and this estimated signal.

[0206] (2) Update the filter coefficients based on a preset adaptive algorithm. The preset adaptive algorithm can be the Normalized Least Mean Squares (NLMS) algorithm. The NLMS algorithm uses the error signal calculated in the previous step to iteratively adjust the coefficients of the adaptive filter, with the goal of minimizing the energy of the subsequently calculated error signal. Over time, the coefficients of the adaptive filter will converge, making the simulated acoustic path closer and closer to the real physical path, thereby determining a minimized error signal.

[0207] (3) The minimized error signal is used as the new microphone receiving signal. When the adaptive filter converges or reaches a steady state, its output error signal is the signal after feedback cancellation. The speaker feedback signal component in this signal has been canceled to the greatest extent. At this time, the signal that has been initially cleaned, i.e. the new microphone receiving signal, will replace the original microphone receiving signal as the input for subsequent processing (such as frame sampling).

[0208] The speaker speech enhancement method provided by this invention employs an adaptive feedback control algorithm to preprocess the microphone received signal. This allows for the elimination of most linearly correlated speaker feedback signals using efficient traditional signal processing methods before they enter the complex deep learning network. This design reduces the processing burden on the subsequent feedback cancellation network model, enabling it to focus more on handling nonlinear and more complex reverberation and background noise, thereby improving the stability of the entire speech enhancement system and the final enhancement effect.

[0209] Considering the different amplification equipment (including microphones and speakers) and the usage environment, it may have excessively high gain peaks in certain frequency bands. These abnormal frequency bands are easily amplified in the feedback loop, thus generating howling.

[0210] Therefore, in another embodiment of the present invention, before eliminating the speaker feedback signal based on the adaptive feedback control algorithm, the method further includes: inputting the microphone received signal to a hardware equalizer (EQ) to adjust the gain of abnormal frequency bands in the microphone received signal using the hardware equalizer.

[0211] Specifically, this hardware equalizer is a hardware circuit or device capable of independently adjusting the gain of different frequency components in an audio signal. It is connected in series between the microphone and the subsequent adaptive feedback control processing unit.

[0212] This gain adjustment process can be a pre-configured calibration operation. For example, before the device leaves the factory or when it is used for the first time in a specific environment, the frequency response of the device can be measured by testing the signal to identify the abnormal frequency band. The hardware equalizer is then configured to specifically perform gain attenuation processing for the identified abnormal frequency band. In this way, when the original microphone received signal flows through the hardware equalizer, its energy in these high-risk frequency bands is pre-suppressed.

[0213] The speaker voice enhancement method provided by this invention actively suppresses the inherent feedback risk of the device from the very beginning of the signal entering the processing system by adding a hardware equalizer calibration step before the adaptive feedback control algorithm. This design not only greatly enhances the stability of the entire loudspeaker, but also provides a more stable input signal with better frequency response characteristics for the subsequent adaptive feedback control algorithm, which helps the AFC algorithm converge faster and work more accurately.

[0214] Considering that even after processing by the feedback cancellation network model, the output target audio signal is still highly similar to the original target speech in content, and the strong correlation between the feedback signal after it is played by the speaker and the new target speech signal still exists, there is a risk of establishing an acoustic feedback loop.

[0215] Therefore, in another embodiment of the present invention, after obtaining the target audio signal based on the target frequency domain signal, the method further includes: performing frequency shifting and phase modulation processing and / or half-wave rectification processing on the target audio signal. This step aims to actively disrupt the strong correlation between the feedback signal and the target speech signal by changing the physical characteristics of the signal.

[0216] Specifically, this step may include, but is not limited to, the following two processing methods:

[0217] (1) Frequency shifting and phase modulation processing: This processing fine-tunes the signal in the frequency domain and phase domain. It includes increasing or decreasing the overall spectrum of the target audio signal by a preset Hertz value (such as 5 Hertz) to disrupt the precise frequency alignment between the feedback signal and the original signal; and periodically changing the phase of the target audio signal to disrupt their phase consistency. This processing has minimal impact on human hearing, but it can effectively weaken the conditions for establishing a feedback loop.

[0218] (2) Half-wave rectification: This process performs nonlinear transformation on the signal waveform in the time domain. It involves converting all sample values ​​less than zero in the time-domain waveform of the target audio signal into sample values ​​greater than zero by taking their absolute values. This nonlinear operation greatly changes the waveform structure of the signal, thereby effectively destroying the linear correlation between the feedback signal and the original signal.

[0219] It should be noted that the frequency shifting and phase modulation processing and the half-wave rectification processing can be performed independently or in combination to achieve a balance according to the actual anti-feedback needs and sound quality requirements.

[0220] The speaker speech enhancement method provided by this invention performs decorrelation processing on the target audio signal to be played, making the sound emitted from the speaker significantly different from the original target speech in terms of physical characteristics. This design fundamentally destroys the strong correlation between the feedback signal and the target speech signal, making it difficult to establish positive feedback, thereby greatly improving the stability of the loudspeaker.

[0221] Considering that after multiple stages of signal processing, the amplitude of the target audio signal may become unpredictable or its dynamic range may be too wide, which may lead to hardware clipping distortion or unstable output volume.

[0222] Therefore, in another embodiment of the present invention, after performing frequency shifting and phase modulation processing and / or half-wave rectification processing on the target audio signal, the method further includes: dynamically adjusting the amplitude of the target audio signal to control the dynamic range of the speaker output signal.

[0223] Specifically, controlling the dynamic range of the speaker output signal can include, but is not limited to, the following operations:

[0224] (1) Compression: When the amplitude of the target audio signal exceeds a preset threshold, the dynamic range control module will automatically reduce its gain. This can effectively suppress sudden excessive volume and prevent the signal amplitude from exceeding the maximum range that the speaker hardware can withstand, thereby avoiding clipping distortion.

[0225] (2) Limitation: Set an absolute upper limit for the signal amplitude. Any signal that attempts to exceed this upper limit will be forced back to the upper limit value. This is the most effective way to prevent signal overload and protect speaker hardware.

[0226] (3) Gain compensation: After the signal is compressed, its overall loudness may decrease. The dynamic range control module can apply a fixed gain to the processed signal to compensate for the lost loudness and ensure that the final output volume is appropriate.

[0227] For example, a compression threshold of -6dB and a compression ratio of 2:1 can be set, along with a hard limiter of -0.1dB. This way, when the signal amplitude exceeds -6dB, the gain of the excess portion will be halved; simultaneously, the signal peak value will never exceed -0.1dB, ensuring absolute output safety.

[0228] The speaker voice enhancement method provided by this invention dynamically adjusts the amplitude of the target audio signal by adding dynamic range control before output. This effectively manages the dynamics of the output signal, preventing sound quality degradation and feedback risks caused by excessive volume. Simultaneously, it suppresses noise in silent segments, improving the listening experience. This ensures that the final sound played from the speaker is both clear and stable, while remaining within the safe operating range of the hardware, greatly enhancing the reliability and practicality of the entire system.

[0229] Figure 7 This is a schematic diagram of the speaker voice enhancement device provided by the present invention, as shown below. Figure 7 As shown, it mainly includes, but is not limited to:

[0230] The signal acquisition unit 710 is used to acquire the initial spectrum signal corresponding to the initial audio signal and determine the initial power spectrum signal corresponding to the initial spectrum signal.

[0231] The signal processing unit 720 is used to input the initial power spectrum signal into the feedback cancellation network model and obtain the target frequency domain signal output by the feedback cancellation network model. The feedback cancellation network model includes a first sub-model and a second sub-model.

[0232] The signal generation unit 730 is used to obtain the target audio signal based on the target frequency domain signal.

[0233] It should be noted that the speaker voice enhancement device provided by the present invention can execute the speaker voice enhancement method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0234] The speaker voice enhancement device provided by this invention first uses a real signal mask to adjust the frequency domain amplitude of the initial spectrum signal to generate a frequency domain estimation signal. Then, it uses a complex signal mask to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhancement signal. Finally, it determines the target frequency domain signal based on the frequency domain estimation signal and the frequency domain enhancement signal, thereby enabling the output target frequency domain signal to have higher fidelity in both amplitude and phase, and can remove environmental noise and speaker feedback signals to the greatest extent, effectively suppressing howling.

[0235] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can call logic instructions in the memory 830 to execute a speaker speech enhancement method. This method includes: acquiring an initial spectral signal corresponding to an initial audio signal and determining an initial power spectral signal corresponding to the initial spectral signal; inputting the initial power spectral signal to a feedback cancellation network model to acquire a target frequency domain signal output by the feedback cancellation network model, the feedback cancellation network model including a first sub-model and a second sub-model; obtaining a target audio signal based on the target frequency domain signal; the first sub-model outputting a real-valued signal mask based on the input initial power spectral signal; the real-valued signal mask being used to mask the frequency domain amplitude of the initial spectral signal to obtain a frequency domain estimated signal; the second sub-model receiving a three-dimensional spectral signal obtained by concatenating the frequency domain estimated signal and the initial spectral signal, and outputting a complex-valued signal mask; the complex-valued signal mask being used to mask the frequency domain amplitude and phase of the initial spectral signal to obtain a frequency domain enhanced signal; and the target frequency domain signal being determined based on the frequency domain estimated signal and the frequency domain enhanced signal.

[0236] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0237] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the speaker speech enhancement method provided in the above embodiments, the method comprising: acquiring an initial spectrum signal corresponding to an initial audio signal, and determining an initial power spectrum signal corresponding to the initial spectrum signal; inputting the initial power spectrum signal to a feedback cancellation network model, and acquiring a target frequency domain signal output by the feedback cancellation network model, the feedback cancellation network model comprising a first sub-model. The first sub-model, based on the target frequency domain signal, outputs a real-valued signal mask according to the input initial power spectrum signal. This real-valued signal mask is used to adjust the frequency domain amplitude of the initial spectrum signal to obtain a frequency domain estimated signal. The second sub-model receives a three-dimensional spectrum signal obtained by concatenating the frequency domain estimated signal and the initial spectrum signal, and outputs a complex-valued signal mask. This complex-valued signal mask is used to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhanced signal. The target frequency domain signal is determined based on the frequency domain estimated signal and the frequency domain enhanced signal.

[0238] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speaker voice enhancement method provided in the above embodiments. The method includes: acquiring an initial spectral signal corresponding to an initial audio signal and determining an initial power spectral signal corresponding to the initial spectral signal; inputting the initial power spectral signal to a feedback cancellation network model to acquire a target frequency domain signal output by the feedback cancellation network model, the feedback cancellation network model including a first sub-model and a second sub-model; obtaining a target audio signal based on the target frequency domain signal; the first sub-model outputting a real-valued signal mask based on the input initial power spectral signal; the real-valued signal mask being used to perform masking adjustment on the frequency domain amplitude of the initial spectral signal to obtain a frequency domain estimated signal; the second sub-model receiving a three-dimensional spectral signal obtained by concatenating the frequency domain estimated signal and the initial spectral signal, and outputting a complex-valued signal mask; the complex-valued signal mask being used to perform masking adjustment on the frequency domain amplitude and phase of the initial spectral signal to obtain a frequency domain enhanced signal; the target frequency domain signal being determined based on the frequency domain estimated signal and the frequency domain enhanced signal.

[0239] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0240] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0241] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for enhancing speech through a loudspeaker, characterized in that, include: Obtain the initial spectrum signal corresponding to the initial audio signal, and determine the initial power spectrum signal corresponding to the initial spectrum signal; The initial power spectrum signal is input into the feedback cancellation network model to obtain the target frequency domain signal output by the feedback cancellation network model. The feedback cancellation network model includes a first sub-model and a second sub-model. The target audio signal is obtained based on the target frequency domain signal; The first sub-model outputs a real-valued signal mask based on the input initial power spectrum signal; the real-valued signal mask is used to adjust the frequency domain amplitude of the initial spectrum signal to obtain a frequency domain estimated signal. The second sub-model receives a three-dimensional spectrum signal obtained by splicing the frequency domain estimated signal and the initial spectrum signal, and outputs a complex signal mask; the complex signal mask is used to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhanced signal; The target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal.

2. The loudspeaker speech enhancement method according to claim 1, characterized in that, The first sub-model includes a first input module, a first GRU module, a second GRU module, a splicing module, a third GRU module, and a first output module; The first input module receives the initial power spectrum signal and inputs the initial power spectrum signal in parallel to the first GRU module and the second GRU module; The first GRU module performs time-series modeling on the initial power spectrum signal to obtain a first time-series feature; The second GRU module performs time-series modeling on the initial power spectrum signal to obtain a second time-series feature; The splicing module splices the first temporal feature and the second temporal feature along the feature dimension to obtain the fused temporal feature; The third GRU module receives the fused timing features, performs timing dependency modeling on the fused timing features, and obtains the final timing features; The first output module is used to map the final timing features to the real number signal mask.

3. The loudspeaker speech enhancement method according to claim 1, characterized in that, The second sub-model includes a second input module, an encoder module, a bottleneck processing module, a decoder module, and a second output module; The second input module inputs the received three-dimensional spectrum signal to the encoder module; The encoder module includes multiple cascaded upsampling convolutional layers, which are used to perform layer-by-layer convolution and upsampling on the three-dimensional spectral signal and output intermediate coding features at each level. The bottleneck processing module performs temporal context modeling on the intermediate encoded features output by the last upsampled convolutional layer to perform filtering and weighting processing on adjacent time frames and output the temporally enhanced bottleneck features. The decoder module includes multiple cascaded downsampling convolutional layers, which are used to downsample and convolve the bottleneck features layer by layer to output a frequency domain decoded signal. The second output module performs convolution mapping on the frequency domain decoded signal and outputs the complex signal mask.

4. The loudspeaker speech enhancement method according to claim 3, characterized in that, The bottleneck processing module includes a gated loop unit, a bidirectional gated loop unit, and a linear mapping layer; The gated loop unit receives the intermediate encoded features output by the last upsampled convolutional layer of the encoder module, performs temporal modeling on them along the time frame dimension, and outputs the first temporal feature; The bidirectional gated loop unit receives the first time-series feature, performs bidirectional time-series modeling to execute filtering and weighting processing of adjacent time frames, and outputs the second time-series feature; The linear mapping layer performs a mapping transformation on the second temporal feature and outputs the bottleneck feature after temporal enhancement.

5. The loudspeaker speech enhancement method according to claim 3, characterized in that, The downsampling convolutional layer in the decoder module corresponds one-to-one with the upsampling convolutional layer in the encoder module; During the process of the decoder module performing layer-by-layer downsampling and convolution on the bottleneck features to output the frequency domain decoded signal, the intermediate encoded features with the same resolution output by the corresponding upsampling convolutional layer are fused through skip connections at each downsampling convolutional layer.

6. The loudspeaker speech enhancement method according to claim 3, characterized in that, The feedback elimination network model also includes a weighted fusion module; The weighted fusion module performs weighted fusion on the frequency domain estimation signal output by the first sub-model and the frequency domain enhancement signal output by the second sub-model to generate the target frequency domain signal.

7. The loudspeaker speech enhancement method according to any one of claims 1-6, characterized in that, The initial audio signal was obtained in the following way: Acquire the microphone signal; The microphone received signal is sampled according to the preset synthesis window length and preset frame shift to obtain the initial audio signal obtained from each sampling.

8. The loudspeaker speech enhancement method according to claim 7, characterized in that, The microphone receives signals including at least the target speech signal and the speaker feedback signal; After acquiring the microphone received signal, the speaker feedback signal is eliminated based on an adaptive feedback control algorithm, specifically including: Initialize an adaptive filter, wherein the error signal of the adaptive filter is set as the difference between the microphone received signal and the estimated signal used to predict the speaker feedback signal; The coefficients of the adaptive filter are updated based on a preset adaptive algorithm to determine the minimized error signal; The minimized error signal is used as the new microphone reception signal after eliminating the speaker feedback signal.

9. The loudspeaker speech enhancement method according to claim 8, characterized in that, Before eliminating the speaker feedback signal based on the adaptive feedback control algorithm, the following steps are also included: The microphone received signal is input to a hardware equalizer to adjust the gain of abnormal frequency bands in the microphone received signal.

10. The loudspeaker speech enhancement method according to any one of claims 1-6, characterized in that, After obtaining the target audio signal based on the target frequency domain signal, the process further includes: The target audio signal is subjected to frequency shifting and phase modulation processing and / or half-wave rectification processing; The frequency shifting and phase modulation process includes increasing or decreasing the overall spectrum of the target audio signal by a preset Hertz, and periodically changing the phase of the target audio signal; The half-wave rectification process includes converting all sample values ​​less than zero in the time-domain waveform of the target audio signal into sample values ​​greater than zero by taking their absolute values.

11. The loudspeaker speech enhancement method according to claim 10, characterized in that, After performing frequency shifting and phase modulation processing and / or half-wave rectification processing on the target audio signal, the method further includes: The amplitude of the target audio signal is dynamically adjusted to control the dynamic range of the speaker output signal.

12. A loudspeaker voice enhancement device, characterized in that, include: The signal acquisition unit is used to acquire the initial spectrum signal corresponding to the initial audio signal and determine the initial power spectrum signal corresponding to the initial spectrum signal. A signal processing unit is used to input the initial power spectrum signal into a feedback cancellation network model and obtain the target frequency domain signal output by the feedback cancellation network model. The feedback cancellation network model includes a first sub-model and a second sub-model. A signal generation unit is used to obtain a target audio signal based on the target frequency domain signal; The first sub-model outputs a real-valued signal mask based on the input initial power spectrum signal; the real-valued signal mask is used to adjust the frequency domain amplitude of the initial spectrum signal to obtain a frequency domain estimated signal. The second sub-model receives a three-dimensional spectrum signal obtained by splicing the frequency domain estimated signal and the initial spectrum signal, and outputs a complex signal mask; the complex signal mask is used to adjust the frequency domain amplitude and phase of the initial spectrum signal to obtain a frequency domain enhanced signal; The target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the speaker speech enhancement method as described in any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the speaker speech enhancement method as described in any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the speaker speech enhancement method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Audio enhancement method and device, electronic equipment and readable storage medium

    CN114974292A

  • Audio signal processing method and device, vehicle-mounted entertainment system, electronic equipment and storage medium

    CN118629415A