Speech enhancement method, device, equipment, medium and program product

By performing technical means such as two voice enhancement and Fourier transform on the microphone received signal, the target enhancement signal is determined and amplified, which solves the problem of high whistling rate in the amplified environment, and achieves better voice quality and system stability.

CN119943072AActive Publication Date: 2025-05-06IFLYTEK CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510086727.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In the amplified environment, due to factors such as environmental noise, room reverb and loudspeaker feedback voice, voice quality deteriorates and makes it easy to whistle. The existing technology is difficult to effectively reduce the probability of whistle.

Method used

By performing two voice enhancements on the microphone reception signal, the first is full-band voice enhancement, and the second is voice enhancement on the low-frequency band signals of the first enhancement signal. Combined with Fourier transform, power spectrum characteristics and complex spectrum characteristics to predict the time frequency domain mask, the target enhancement signal is determined and amplified.

Benefits of technology

It significantly reduces the probability of howling in the amplified environment, improves the voice quality, and reduces the occurrence of howling, while no additional hardware costs are required, and improves the stability of the amplified system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943072A_ABST
    Figure CN119943072A_ABST
Patent Text Reader

Abstract

The invention provides a voice enhancement method and device, equipment, a medium and a program product, and the method comprises the steps: carrying out the voice enhancement of a microphone receiving signal, and obtaining a first enhanced signal; performing voice enhancement on a low-frequency-band signal in the first enhanced signal to obtain a second enhanced signal; determining a target enhanced signal based on the first enhanced signal and the second enhanced signal; and amplifying the target enhanced signal to obtain an amplified signal. According to the invention, the occurrence probability of howling in a sound amplification environment can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and in particular to a speech enhancement method, device, equipment, medium and program product. Background Art

[0002] Loudspeakers are mainly used to amplify the speaker's voice. However, due to differences in the speaker's volume, environment, and equipment, the signal played through the loudspeaker is often affected by environmental noise, room reverberation, and feedback from the loudspeaker, resulting in poor voice quality and the possibility of howling. Currently, the probability of howling in a loudspeaker environment is high. Summary of the invention

[0003] The present application provides a speech enhancement method, apparatus, device, medium and program product for reducing the probability of howling in a sound reinforcement environment.

[0004] According to a first aspect of an embodiment of the present application, a speech enhancement method is provided, comprising:

[0005] Performing voice enhancement on a signal received by a microphone to obtain a first enhanced signal;

[0006] Performing speech enhancement on a signal in a low frequency band of the first enhanced signal to obtain a second enhanced signal;

[0007] Determining a target enhanced signal based on the first enhanced signal and the second enhanced signal;

[0008] The target enhanced signal is amplified to obtain an amplified signal.

[0009] Optionally, performing voice enhancement on the microphone received signal to obtain a first enhanced signal includes:

[0010] Perform Fourier transform on the microphone received signal to obtain frequency domain characteristics;

[0011] Based on the frequency domain features, a power spectrum feature and a complex spectrum feature are obtained;

[0012] Predicting a time-frequency domain mask based on the power spectrum feature and the complex spectrum feature;

[0013] Based on the frequency domain features and the time-frequency domain mask, a first enhanced signal is obtained.

[0014] Optionally, performing speech enhancement on the signal in the low frequency band of the first enhanced signal to obtain the second enhanced signal includes:

[0015] Based on the power spectrum characteristics and the complex spectrum characteristics, obtaining filter estimation complex coefficients corresponding to the low frequency band;

[0016] A second enhanced signal is obtained based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal.

[0017] Optionally, a filter order, a target frame and a frequency correspond to a filter estimation complex coefficient; in the frequency domain feature, the acquisition time of the target frame is greater than or equal to the acquisition time of the current frame;

[0018] The step of obtaining a second enhanced signal based on the filter estimation complex coefficients corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal comprises:

[0019] Determine an intermediate enhanced signal corresponding to any one of the filter orders and the low frequency band based on the estimated complex coefficients of the filter corresponding to each target frame in the estimated complex coefficients of the filter corresponding to any one of the filter orders and the low frequency band, and the signal of the low frequency band in the first enhanced signal;

[0020] A second enhanced signal is obtained based on each filter order and the intermediate enhanced signal corresponding to the low frequency band.

[0021] Optionally, determining a target enhanced signal based on the first enhanced signal and the second enhanced signal includes:

[0022] Determining a weighting coefficient based on the power spectrum characteristics and the complex spectrum characteristics;

[0023] Determine a third enhanced signal based on the first enhanced signal, the second enhanced signal and the weighting coefficient;

[0024] Performing inverse Fourier transform on the third enhanced signal to obtain a fourth enhanced signal;

[0025] Based on the fourth enhanced signal, a target enhanced signal is determined.

[0026] Optionally, determining a target enhanced signal based on the fourth enhanced signal includes:

[0027] Performing frequency shifting, phase modulation and half-wave rectification on the feedback signal in the fourth enhanced signal to remove the correlation between the feedback signal in the fourth enhanced signal and the target speech signal to obtain a fifth enhanced signal;

[0028] Determining a target enhanced signal based on the fifth enhanced signal;

[0029] The feedback signal refers to the signal played by the speaker and transmitted back to the microphone; the target voice signal refers to the user voice signal received by the microphone.

[0030] Optionally, determining a target enhanced signal based on the fifth enhanced signal includes:

[0031] determining a speaker output signal based on the fifth enhanced signal;

[0032] Mapping the dynamic range of the speaker output signal to a preset dynamic range to obtain an adjusted speaker output signal;

[0033] A target enhancement signal is determined based on the adjusted loudspeaker output signal.

[0034] According to a second aspect of an embodiment of the present application, a speech enhancement device is provided, comprising:

[0035] A first enhancement unit, configured to perform voice enhancement on a signal received by a microphone to obtain a first enhanced signal;

[0036] A second enhancement unit, configured to perform speech enhancement on a signal in a low frequency band of the first enhanced signal to obtain a second enhanced signal;

[0037] a processing unit, configured to determine a target enhanced signal based on the first enhanced signal and the second enhanced signal;

[0038] The amplification unit is used to amplify the target enhanced signal to obtain an amplified signal.

[0039] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a memory and a processor;

[0040] The memory is connected to the processor and is used to store programs;

[0041] The processor is used to implement the speech enhancement method as described in the first aspect by running the program in the memory.

[0042] According to a fourth aspect of an embodiment of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the speech enhancement method as described in the first aspect is implemented.

[0043] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer program instructions, which, when executed by a processor, enable the processor to perform the speech enhancement method as described in the first aspect.

[0044] In the present application, the microphone received signal is voice enhanced to obtain a first enhanced signal, the low-frequency signal in the first enhanced signal is voice enhanced to obtain a second enhanced signal, the target enhanced signal is determined based on the first enhanced signal and the second enhanced signal, the target enhanced signal is amplified to obtain the amplified signal. The microphone received signal is voice enhanced twice, the first voice enhancement is to perform full-band voice enhancement on the microphone received signal, which can reduce the probability of howling, the second voice enhancement is to perform voice enhancement on the low-frequency signal in the first enhanced signal after the first voice enhancement, because the signal energy in the low-frequency band is stronger, it is more likely to generate feedback howling, the low-frequency signal in the first enhanced signal is voice enhanced to obtain the second enhanced signal, which can further reduce the probability of howling, based on the first enhanced signal and the second enhanced signal, the target enhanced signal is determined, the target enhanced signal is amplified to obtain the amplified signal, which can further reduce the probability of howling in the amplified environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 A flow chart of a speech enhancement method provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of a source of a signal received by a microphone in a sound amplification scenario provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a process of step 101 provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of a process of step 102 provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of a process of step 402 provided in an embodiment of the present application;

[0051] Figure 6 A schematic diagram of a process of step 103 provided in an embodiment of the present application;

[0052] Figure 7 A flowchart of step 604 provided in an embodiment of the present application;

[0053] Figure 8A flowchart of step 702 provided in an embodiment of the present application;

[0054] Fig. 9 A schematic diagram of the structure of an exemplary adaptive feedback suppression model provided in an embodiment of the present application;

[0055] Fig.10 A schematic diagram of the structure of a speech enhancement device provided in an embodiment of the present application;

[0056] Fig.11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] Loudspeakers are mainly used to amplify the speaker's voice. However, due to differences in the speaker's volume, environment, and equipment, the signal played through the loudspeaker is often affected by environmental noise, room reverberation, and loudspeaker feedback, resulting in poor voice quality and prone to howling. At the same time, due to differences in hardware algorithms, the amplification delay is often large, and the listening experience is poor. At present, the probability of howling in a loudspeaker environment is high.

[0058] Existing solutions for speaker howling suppression mainly consider hardware, and reduce the probability of loudspeaker howling by directional microphone pickup and reducing the maximum gain of the loudspeaker. By adding a filter to attenuate the output voice low-frequency energy, or reducing the maximum gain of the loudspeaker output, the purpose of balanced amplification and reduced howling is achieved. However, this method has no effect on howling that has already occurred in the system and can only play a passive defensive role. At the same time, due to the portable nature of loudspeakers, existing noise reduction algorithms for loudspeakers have problems such as slow convergence and large delay.

[0059] In order to reduce the probability of howling in a sound reinforcement environment, the present application provides a speech enhancement method, apparatus, device, medium and program product.

[0060] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0061] Exemplary Implementation Environment

[0062] The speech enhancement method according to the embodiment of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user device, a mobile device, a computing device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling a computer-readable program instruction stored in a memory.

[0063] Exemplary Methods

[0064] See also Figure 1 In an exemplary embodiment, a speech enhancement method is provided. Figure 1 As shown, the process of the speech enhancement method mainly includes:

[0065] Step 101: Perform voice enhancement on a signal received by a microphone to obtain a first enhanced signal.

[0066] In some embodiments, Figure 2 As shown, it is a schematic diagram of the source of the signal received by the microphone in the sound reinforcement scenario. Figure 2 In the figure, the microphone received signal is represented by y(t), the target speech signal is represented by s(t), the feedback signal is represented by d(t), the background noise is represented by n(t), the room reverberation is represented by r(t), the speaker playback reference signal is represented by x(t), G is the speaker amplification gain value, and h(t) represents the attenuation during the propagation process from the speaker to the microphone. Among them, the target speech signal s(t) refers to the user speech signal received by the microphone, that is, the speaker's direct speech signal; the feedback signal d(t) refers to the signal played back by the speaker to the microphone, that is, the speaker plays the direct signal; the room reverberation r(t) refers to the signal of the speaker's speech and the speaker signal propagated to the microphone through the room reflection. In the sound reinforcement scenario, the microphone received signal y(t) can be mainly divided into three components: (1) target speech signal s(t); (2) feedback signal d(t); (3) background noise n(t) and room reverberation r(t).

[0067] In an exemplary embodiment, the specific formula of the feedback signal d(t) is as follows:

[0068] d(t)=NL(x(t))*h(t)

[0069] Among them, NL() represents the nonlinear transformation of the speaker, x(t) represents the reference signal played by the speaker, h(t) represents the attenuation during the propagation process from the speaker to the microphone, and * represents linear convolution.

[0070] Theoretically, without any other processing, the speaker playback reference signal x(t) is obtained by delaying and amplifying the microphone reception signal y(t), and this signal will be repeatedly picked up by the microphone. The formula of the microphone reception signal y(t) corresponding to any time t is as follows:

[0071]

[0072] Among them, s(t) represents the target speech signal, n(t) represents the background noise, r(t) represents the room reverberation, Δt represents the system transmission delay from microphone to speaker, NL() represents the speaker nonlinear transformation, G is the speaker amplification gain value, I is the total time for the feedback signal to be amplified, and h(t) represents the attenuation during the propagation process from speaker to microphone.

[0073] It can be seen from the above formula that repeated picking up of the feedback signal will form positive feedback and lead to cyclic amplification at certain frequencies, forming howling.

[0074] In some embodiments, Figure 3 As shown, step 101 includes:

[0075] Step 301: Perform Fourier transform on the microphone received signal to obtain frequency domain features.

[0076] In some embodiments, before step 301, the speech enhancement method further includes: filtering out signals below 150 Hz in the microphone received signal through a high-pass filter.

[0077] Among them, the signals below 150Hz are mainly background noise and have no effect on the target speech signal.

[0078] In an exemplary embodiment, the Fourier transform may be a STFT (Short-Time Fourier Transform) or other types of Fourier transform, and the present application is not limited thereto.

[0079] In an exemplary embodiment, the microphone receiving signal is represented by y(t), and the frequency domain feature is represented by X(k, f), wherein k represents the frame number and f represents the frequency point.

[0080] In an exemplary embodiment, the length of each frame of frequency domain features is 256 points, that is, the synthesis window length is 256 points, the difference between the starting points of adjacent frame frequency domain features is 80 points, that is, the frame shift is 80 points, the sampling rate is 16000 times per second, and the processing delay is (256+80) / 16000=21ms, which meets the low latency requirement of the sound reinforcement system, reduces the link delay, and makes there is no obvious time difference between the speaker's direct signal and the sound reinforcement output signal. The human ear cannot hear the delay between the direct sound of the person and the sound of the speaker, thus avoiding the "echo" problem in the sound reinforcement.

[0081] Step 302: Obtain power spectrum features and complex spectrum features based on frequency domain features.

[0082] In an exemplary embodiment, the power spectrum feature may be a log power spectrum (LPS) feature, or may be other types of power spectrum features, which is not limited in the present application.

[0083] In the exemplary embodiment, the frequency domain feature is represented by X(k,f); the power spectrum feature is represented by X LPS (k,f); complex spectrum features are represented by X DF (k,f′) represents; k represents the frame number; f represents the frequency point; f′ represents the frequency point of the low frequency band, f′=f / 2.

[0084] Power spectrum feature X LPS The calculation formula for (k,f) is as follows:

[0085] X LPS (k,f)=20log 10 (Re(X(k,f)) 2 +Im(X(k,f)) 2 )

[0086] Among them, Re represents the real part of the complex number, and Im represents the imaginary part of the complex number.

[0087] Complex spectral feature X DF The calculation formula of (k,f′) is as follows:

[0088] X DF (k,f′)=(Re(X(k,f / 2)),Im(X(k,f / 2)))

[0089] Wherein, Re represents the real part of the complex number, Im represents the imaginary part of the complex number, f′ represents the frequency point of the low frequency band, and f′=f / 2.

[0090] Step 303: predicting a time-frequency domain mask based on the power spectrum features and the complex spectrum features.

[0091] In an exemplary embodiment, the power spectrum feature is represented by X LPS(k,f); complex spectrum features are represented by X DF (k,f′); the time-frequency domain mask is represented by G(k,f); k represents the number of frames (time domain dimension); f represents the frequency point (frequency domain dimension).

[0092] In an exemplary embodiment, step 303 may include: inputting the power spectrum features and the complex spectrum features into the codec network to obtain the time-frequency domain mask output by the codec network. Specifically, the power spectrum features and the complex spectrum features may be input into the encoder to obtain the output features of the encoder; the output features of the encoder may be input into the decoder to obtain the time-frequency domain mask output by the decoder.

[0093] Step 304: Obtain a first enhanced signal based on the frequency domain features and the time-frequency domain mask.

[0094] In the exemplary embodiment, the frequency domain feature is represented by X(k, f); the time-frequency domain mask is represented by G(k, f); the first enhanced signal is represented by Y G (k,f) represents.

[0095] The first enhanced signal Y G The specific formula of (k,f) is as follows:

[0096] Y G (k,f)=X(k,f)·G(k,f)

[0097] Perform Fourier transform on the microphone received signal to obtain frequency domain features. Based on the frequency domain features, obtain power spectrum features and complex spectrum features. Based on the power spectrum features and complex spectrum features, predict the time-frequency domain mask. Based on the frequency domain features and the time-frequency domain mask, obtain the first enhanced signal. By estimating the real target signal mask, that is, the time-frequency domain mask, and acting on the original microphone signal, obtain the first stage enhanced spectrum envelope, that is, the first enhanced signal, and estimate and enhance the full-band signal. For the sound reinforcement system, it mainly suppresses the noise and reverberation components in the microphone received signal and improves the signal-to-noise ratio of the microphone received signal.

[0098] In some other embodiments, step 101 includes: performing speech enhancement on a signal received by a microphone through a frequency domain filter to obtain a first enhanced signal.

[0099] Step 102: Perform speech enhancement on the signal in the low frequency band of the first enhanced signal to obtain a second enhanced signal.

[0100] In some embodiments, the selection method of the low-frequency signal in the first enhanced signal includes but is not limited to the following methods:

[0101] Method 1

[0102] The frequency of the low frequency band is less than or equal to the middle frequency of the first enhanced signal.

[0103] In an exemplary embodiment, the middle frequency of the first enhanced signal may refer to a frequency located in the middle of the frequency band of the first enhanced signal. For example, the middle frequency of the first enhanced signal may refer to half of the maximum frequency of the first enhanced signal; the middle frequency of the first enhanced signal may also refer to the average of the maximum frequency and the minimum frequency of the first enhanced signal; the middle frequency of the first enhanced signal may also refer to other frequencies, and the present application does not limit this. The present application takes the middle frequency of the first enhanced signal as half of the maximum frequency of the first enhanced signal as an example for explanation.

[0104] Method 2

[0105] The frequency of the low frequency band is less than or equal to a preset proportional frequency of the first enhanced signal, wherein the preset proportional frequency is equal to a frequency obtained by multiplying a maximum frequency of the first enhanced signal by a preset proportion.

[0106] Method 3

[0107] The frequency of the low frequency band is less than or equal to the preset frequency, wherein the value of the preset frequency can be set by oneself.

[0108] The microphone receiving signal is voice enhanced twice. The first voice enhancement is to perform full-band voice enhancement on the microphone receiving signal, which can reduce the probability of howling. The second voice enhancement is to perform voice enhancement on the low-frequency signal in the first enhanced signal after the first voice enhancement. Since the signal energy in the low-frequency band is stronger, it is more likely to produce feedback howling. The low-frequency signal in the first enhanced signal is voice enhanced to obtain the second enhanced signal, which can further reduce the probability of howling. Based on the first enhanced signal and the second enhanced signal, the target enhanced signal is determined, and the target enhanced signal is amplified to obtain the amplified signal, which can further reduce the probability of howling in the amplification environment. Based on the original amplification equipment, without increasing additional hardware costs, the stability of the amplification system can be significantly improved and the occurrence of howling can be reduced.

[0109] In some embodiments, step 102 includes: performing speech enhancement on a signal in a low frequency band of the first enhanced signal through a frequency domain filter to obtain a second enhanced signal.

[0110] In other embodiments, Figure 4 As shown, step 102 includes:

[0111] Step 401: Obtain filter estimation complex coefficients corresponding to a low frequency band based on power spectrum characteristics and complex spectrum characteristics.

[0112] In an exemplary embodiment, step 401 may include: inputting power spectrum features and complex spectrum features into an encoder to obtain output features of the encoder; inputting the output features of the encoder into a DeepFilterNet network to obtain filter estimated complex coefficients corresponding to the low frequency band.

[0113] Among them, DeepFilterNet is a low-complexity speech enhancement framework designed for full-band audio (48kHz). Based on deep filtering technology, it processes audio signals through a neural network model to achieve efficient noise suppression.

[0114] Step 402: Obtain a second enhanced signal based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal.

[0115] Based on the power spectrum characteristics and the complex spectrum characteristics, the estimated complex coefficients of the filter corresponding to the low frequency band are obtained. Based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal, the second enhanced signal is obtained. By estimating the complex target signal mask, that is, the estimated complex coefficients of the filter, the reverberation and feedback components remaining in the first enhanced signal are further eliminated, the residual speech frame interference signal in the first enhanced signal is suppressed, the speech periodic component is further enhanced, and the target speech component is actively enhanced. Moreover, since the signal energy in the low frequency band is stronger, it is more likely to generate feedback howling. Based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal, the second enhanced signal is obtained, which can further reduce the probability of howling.

[0116] In some embodiments, a filter order, a target frame and a frequency correspond to a filter estimation complex coefficient; in the frequency domain features, the acquisition time of the target frame is greater than or equal to the acquisition time of the current frame.

[0117] In an exemplary embodiment, the filter estimates the complex coefficients using C i (k+j,f′) represents, where i refers to the i-th filter order, k represents the number of frames, the k-th frame is the current frame, the (k+j)-th frame refers to the j-th target frame on the right side of the current frame, and f′ refers to the frequency point representing the low frequency band.

[0118] In some embodiments, Figure 5 As shown, step 402 includes:

[0119] Step 501, based on the estimated complex coefficients of the filters corresponding to each target frame in the estimated complex coefficients of the filters corresponding to any filter order and the low frequency band, and the signal of the low frequency band in the first enhanced signal, determine the intermediate enhanced signal corresponding to any filter order and the low frequency band.

[0120] In an exemplary embodiment, the intermediate enhanced signal corresponding to the i-th filter order and the low frequency band is represented by Y i (k,f′) represents the intermediate enhanced signal Y i The calculation formula of (k,f′) is as follows:

[0121]

[0122] Where i refers to the i-th filter order, k refers to the number of frames, the k-th frame is the current frame, the (k+j)-th frame refers to the j-th target frame on the right side of the current frame, f′ refers to the frequency point representing the low frequency band, and the complex coefficient of the filter estimation is C i (k+j,f′) means that l refers to the total number of frames to the right of the kth frame, Y G (k―i,f′) refers to the first enhanced signal corresponding to the i-th filter order and the low frequency band.

[0123] Step 502: Obtain a second enhanced signal based on the intermediate enhanced signals corresponding to the respective filter orders and the low frequency bands.

[0124] In an exemplary embodiment, the second enhanced signal is Y DF′ (k, f′) represents the intermediate enhanced signal corresponding to the i-th filter order and the low frequency band, and Y i (k,f′) represents the second enhanced signal Y DF′ The calculation formula of (k,f′) is as follows:

[0125]

[0126] Where N is the total filter order.

[0127] Therefore, the second enhanced signal Y DF′ The calculation formula of (k,f′) is as follows:

[0128]

[0129] The first stage is to estimate and enhance the full-band signal to obtain the first enhanced signal Y G (k,f), the second stage enhances the low-frequency band of the signal based on the first stage. Since the energy of the low-frequency signal is strong, it is more likely to generate feedback howling. The use of complex filters can effectively restore the original signal periodic components and reduce the probability of system howling. By estimating the complex target signal mask, that is, the filter estimates the complex coefficients, and filtering adjacent frames, the reverberation and feedback components remaining in the first enhanced signal are further eliminated, the residual speech frame interference signal in the first enhanced signal is suppressed, the speech periodic components are further enhanced, and the target speech components are actively enhanced.

[0130] In some other embodiments, step 402 includes: directly applying the estimated complex coefficients of the filter corresponding to the low frequency band to the signal of the low frequency band in the first enhanced signal to obtain the second enhanced signal.

[0131] Step 103: Determine a target enhanced signal based on the first enhanced signal and the second enhanced signal.

[0132] In some embodiments, step 103 includes: performing an inverse Fourier transform on a signal obtained by adding the first enhanced signal and the second enhanced signal to obtain a target enhanced signal.

[0133] In other embodiments, Figure 6 As shown, step 103 includes:

[0134] Step 601, determining weighting coefficients based on power spectrum characteristics and complex spectrum characteristics.

[0135] In the exemplary embodiment, the weighting coefficient is represented by α(k), where k represents the number of frames.

[0136] In an exemplary embodiment, step 601 may include: inputting power spectrum features and complex spectrum features into an encoder to obtain output features of the encoder; and inputting the output features of the encoder into a DeepFilterNet network to obtain weighting coefficients.

[0137] Among them, DeepFilterNet is a low-complexity speech enhancement framework designed for full-band audio (48kHz). Based on deep filtering technology, it processes audio signals through a neural network model to achieve efficient noise suppression.

[0138] Step 602: Determine a third enhanced signal based on the first enhanced signal, the second enhanced signal and the weighting coefficient.

[0139] In an exemplary embodiment, the first enhanced signal is Y G (k,f) represents the second enhanced signal, Y DF′

[0140] (k, f′) represents the weighting coefficient, α(k) represents the frame number, f represents the frequency point, f′ represents the frequency point of the low frequency band, and the third enhanced signal is represented by Y DF (k,f) represents the third enhanced signal Y DF The calculation formula for (k,f) is as follows:

[0141] Y DF (k,f)=α(k)·Y DF′ (k,f′)+(1―α(k))·Y G (k,f)

[0142] Step 603: Perform inverse Fourier transform on the third enhanced signal to obtain a fourth enhanced signal.

[0143] In an exemplary embodiment, the inverse Fourier transform may be an ISTFT (Inverse Short-Time Fourier Transform) or other types of inverse Fourier transform, and the present application is not limited thereto.

[0144] Step 604: determine a target enhanced signal based on the fourth enhanced signal.

[0145] Based on the power spectrum features and the complex spectrum features, a weighting coefficient is determined, and based on the first enhanced signal, the second enhanced signal and the weighting coefficient, a third enhanced signal is determined. The first enhanced signal is a signal obtained by performing full-band speech enhancement on the signal received by the microphone. The second speech enhancement is to perform speech enhancement on the signal in the low-frequency band of the first enhanced signal after the first speech enhancement. Since the signal energy in the low-frequency band is stronger, it is more likely to produce feedback howling. Speech enhancement is performed on the signal in the low-frequency band of the first enhanced signal to obtain the second enhanced signal, which can further reduce the probability of howling. In the process of determining the third enhanced signal, both the first enhanced signal and the second enhanced signal are considered, and the first enhanced signal and the second enhanced signal are weighted by the weighting coefficient determined by the power spectrum features and the complex spectrum features to obtain the third enhanced signal, which can not only retain the target speech signal, but also suppress the interference signal, thereby reducing the probability of howling.

[0146] In some embodiments, step 604 includes: directly determining the fourth enhanced signal as the target enhanced signal.

[0147] In other embodiments, Figure 7 As shown, step 604 includes:

[0148] Step 701: Perform frequency shifting, phase modulation and half-wave rectification on the feedback signal in the fourth enhanced signal to remove the correlation between the feedback signal in the fourth enhanced signal and the target speech signal, so as to obtain a fifth enhanced signal.

[0149] The feedback signal refers to the signal played by the speaker and transmitted back to the microphone; the target voice signal refers to the user voice signal received by the microphone.

[0150] In an exemplary embodiment, phase modulation and frequency shifting refers to PMFS (Phase Modulation and Frequency Shifting), which involves phase modulation (PM) and frequency shifting (FS).

[0151] In an exemplary embodiment, half-wave rectification refers to Half-wave Rectification, wherein Rectification (rectification) is generally represented by RECT.

[0152] In an exemplary embodiment, y(t)=s(t)+n(t)+r(t)+d(t), wherein the microphone received signal is represented by y(t), the target speech signal is represented by s(t), the feedback signal is represented by d(t), the background noise is represented by n(t), and the room reverberation is represented by r(t). The feedback signal in the fourth enhanced signal can be subjected to frequency shifting, phase modulation and half-wave rectification to remove the correlation between the feedback signal and the target speech signal in the fourth enhanced signal, that is, to remove the correlation between d(t) and s(t). After removing the correlation, the formula of the microphone received signal y(t) is as follows:

[0153] y(t)=s(t)+z(t)

[0154] Among them, z(t) refers to the noise interference that is unrelated to the target speech signal s(t). The formula of z(t) is as follows:

[0155] z(t)=n(t)+r(t)+d(t)

[0156] Among them, d(t) and s(t) in z(t) have removed the correlation and are not relevant.

[0157] Step 702: determine a target enhanced signal based on the fifth enhanced signal.

[0158] Since the feedback signal and the target voice signal in the sound reinforcement system have a strong correlation, the fourth enhanced signal obtained by nonlinear processing may have distortion of the target voice signal. Therefore, the feedback signal in the fourth enhanced signal is frequency-shifted, phase-modulated and half-wave rectified to remove the correlation between the feedback signal and the target voice signal in the fourth enhanced signal, so as to obtain the fifth enhanced signal. Based on the fifth enhanced signal, the target enhanced signal is determined, which can reduce the correlation between the feedback signal and the target voice signal and reduce the occurrence of distortion of the target voice signal.

[0159] In some embodiments, step 702 includes: directly determining the fifth enhanced signal as the target enhanced signal.

[0160] In other embodiments, Figure 8 As shown, step 702 includes:

[0161] Step 801: determine a speaker output signal based on a fifth enhanced signal.

[0162] Step 802 : Map the dynamic range of the speaker output signal to a preset dynamic range to obtain an adjusted speaker output signal.

[0163] In an exemplary embodiment, mapping the dynamic range of the speaker output signal to a preset dynamic range may refer to DRC (Dynamic Range Control).

[0164] Step 803: Determine a target enhanced signal based on the adjusted speaker output signal.

[0165] Based on the fifth enhancement signal, the speaker output signal is determined, the dynamic range of the speaker output signal is mapped to the preset dynamic range, and an adjusted speaker output signal is obtained. Based on the adjusted speaker output signal, a target enhancement signal is determined, which can avoid howling caused by an excessively large speaker reference signal and reduce the probability of howling.

[0166] In some embodiments, step 803 includes: determining a sixth enhanced signal based on the adjusted speaker output signal; and processing the sixth enhanced signal through an equalizer to obtain a target enhanced signal.

[0167] Among them, the equalizer (EQ) can suppress some frequency bands that are prone to howling and compensate for voice distortion.

[0168] Step 104: amplify the target enhanced signal to obtain an amplified signal.

[0169] In an exemplary embodiment, step 104 may include: amplifying the target enhanced signal through a loudspeaker to obtain an amplified signal.

[0170] In the exemplary embodiment, the microphones in common loudspeakers are mainly of two types: wired directional microphones and wireless omnidirectional microphones. Since wired directional microphones pick up sound in a directional manner and the microphones are relatively farther away from the speakers, the probability of howling is lower. Therefore, the speech enhancement method in this application is mainly aimed at wireless omnidirectional microphones.

[0171] In an exemplary embodiment, Fig. 9 , which is a schematic diagram of the structure of an exemplary adaptive feedback suppression model. Fig. 9 In the paper, the Adaptive Feedback Suppression (AFS) model includes a first-stage network and a second-stage network. The first-stage network includes an encoder and a decoder, and the second-stage network includes DeepFilterNet. DeepFilterNet is a low-complexity speech enhancement framework designed for full-band audio (48kHz). Based on deep filtering technology, it processes audio signals through a neural network model to achieve efficient noise suppression. Fig. 9In the figure, Conv refers to convolution; GGRU refers to Group Gated Recurrent Unit, group gated recurrent unit; Linear refers to linear processing.

[0172] according to Fig. 9 The exemplary adaptive feedback suppression model in the example speech enhancement method mainly includes: performing Fourier transform on the microphone received signal to obtain frequency domain features; based on the frequency domain features, obtaining power spectrum features and complex spectrum features; inputting the power spectrum features and complex spectrum features into the encoder to obtain the output features of the encoder; inputting the output features of the encoder into the decoder to obtain the time-frequency domain mask output by the decoder; based on the frequency domain features and the time-frequency domain mask, obtaining a first enhanced signal; inputting the output features of the encoder into DeepFilterNet to obtain the filter estimation complex coefficients and weighting coefficients corresponding to the low frequency band; based on the filter estimation complex coefficients corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal, obtaining a second enhanced signal; based on the first enhanced signal, the second enhanced signal and the weighting coefficient, determining a third enhanced signal; performing inverse Fourier transform on the third enhanced signal to obtain a fourth enhanced signal; based on the fourth enhanced signal, determining a target enhanced signal; amplifying the target enhanced signal to obtain an amplified signal.

[0173] Fig. 9 The exemplary adaptive feedback suppression model in is a multi-stage noise reduction compression model designed by this application with reference to DeepFilterNet, which can improve the model's suppression effect on multi-component (noise, reverberation and feedback) signals, and Fig. 9 Compared with the traditional DeepFilterNet, the exemplary adaptive feedback suppression model in has fewer parameters, faster calculation speed, and the noise reduction effect is similar to that of the traditional DeepFilterNet. As shown in Table 1, Fig. 9 A table showing the network parameters and computational complexity of the exemplary adaptive feedback suppression model and the traditional DeepFilterNet under the same comparable conditions, as well as the noise reduction effect of the model at different signal-to-noise ratios.

[0174] Table 1

[0175]

[0176] In Table 1, SNR (Signal-to-Noise Ratio), SDR (Signal-to-Distortion Ratio), and PESQ (Perceptual Evaluation of Speech Quality). The larger the SDR and PESQ, the better.

[0177] From Table 1, we can see that Fig. 9 Compared with the traditional DeepFilterNet, the exemplary adaptive feedback suppression model in has fewer parameters, faster calculation speed, and the noise reduction effect is similar to that of the traditional DeepFilterNet.

[0178] In summary, in the present application, the microphone received signal is voice enhanced to obtain a first enhanced signal, the low-frequency signal in the first enhanced signal is voice enhanced to obtain a second enhanced signal, the target enhanced signal is determined based on the first enhanced signal and the second enhanced signal, the target enhanced signal is amplified to obtain the amplified signal. The microphone received signal is voice enhanced twice, the first voice enhancement is full-band voice enhancement of the microphone received signal, which can reduce the probability of howling, the second voice enhancement is voice enhancement of the low-frequency signal in the first enhanced signal after the first voice enhancement, because the signal energy in the low-frequency band is stronger, it is more likely to generate feedback howling, the low-frequency signal in the first enhanced signal is voice enhanced to obtain the second enhanced signal, which can further reduce the probability of howling, based on the first enhanced signal and the second enhanced signal, the target enhanced signal is determined, the target enhanced signal is amplified to obtain the amplified signal, which can further reduce the probability of howling in the amplified environment.

[0179] Exemplary Devices

[0180] Accordingly, the present application also provides a speech enhancement device, such as Fig.10 As shown, the speech enhancement device comprises:

[0181] The first enhancement unit 1001 is used to perform speech enhancement on a signal received by a microphone to obtain a first enhanced signal;

[0182] The second enhancement unit 1002 is configured to perform speech enhancement on the signal in the low frequency band of the first enhanced signal to obtain a second enhanced signal;

[0183] The processing unit 1003 is configured to determine a target enhanced signal based on the first enhanced signal and the second enhanced signal;

[0184] The amplification unit 1004 is used to amplify the target enhanced signal to obtain an amplified signal.

[0185] Optionally, the first enhancement unit 1001 is specifically configured to:

[0186] Perform Fourier transform on the microphone received signal to obtain frequency domain characteristics;

[0187] Based on the frequency domain features, a power spectrum feature and a complex spectrum feature are obtained;

[0188] Predicting a time-frequency domain mask based on the power spectrum feature and the complex spectrum feature;

[0189] Based on the frequency domain features and the time-frequency domain mask, a first enhanced signal is obtained.

[0190] Optionally, the second enhancement unit 1002 includes:

[0191] A filter estimation subunit, used for obtaining filter estimation complex coefficients corresponding to a low frequency band based on the power spectrum characteristics and the complex spectrum characteristics;

[0192] The first processing subunit is used to obtain a second enhanced signal based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal.

[0193] Optionally, a filter order, a target frame and a frequency correspond to a filter estimation complex coefficient; in the frequency domain feature, the acquisition time of the target frame is greater than or equal to the acquisition time of the current frame;

[0194] The first processing subunit is specifically configured to:

[0195] Determine an intermediate enhanced signal corresponding to any one of the filter orders and the low frequency band based on the estimated complex coefficients of the filter corresponding to each target frame in the estimated complex coefficients of the filter corresponding to any one of the filter orders and the low frequency band, and the signal of the low frequency band in the first enhanced signal;

[0196] A second enhanced signal is obtained based on each filter order and the intermediate enhanced signal corresponding to the low frequency band.

[0197] Optionally, the processing unit 1003 includes:

[0198] A weighting coefficient determination subunit, used to determine a weighting coefficient based on the power spectrum characteristics and the complex spectrum characteristics;

[0199] a second processing subunit, configured to determine a third enhanced signal based on the first enhanced signal, the second enhanced signal and the weighting coefficient;

[0200] an inverse transform subunit, configured to perform inverse Fourier transform on the third enhanced signal to obtain a fourth enhanced signal;

[0201] The third processing subunit is configured to determine a target enhanced signal based on the fourth enhanced signal.

[0202] Optionally, the third processing subunit is specifically configured to:

[0203] Performing frequency shifting, phase modulation and half-wave rectification on the feedback signal in the fourth enhanced signal to remove the correlation between the feedback signal in the fourth enhanced signal and the target speech signal to obtain a fifth enhanced signal;

[0204] Determining a target enhanced signal based on the fifth enhanced signal;

[0205] The feedback signal refers to the signal played by the speaker and transmitted back to the microphone; the target voice signal refers to the user voice signal received by the microphone.

[0206] Optionally, the third processing subunit is specifically configured to:

[0207] determining a speaker output signal based on the fifth enhanced signal;

[0208] Mapping the dynamic range of the speaker output signal to a preset dynamic range to obtain an adjusted speaker output signal;

[0209] A target enhancement signal is determined based on the adjusted loudspeaker output signal.

[0210] The speech enhancement device provided in this embodiment belongs to the same application concept as the speech enhancement method provided in the above embodiments of this application, can execute the speech enhancement method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of executing the speech enhancement method. For technical details not fully described in this embodiment, please refer to the specific processing content of the speech enhancement method provided in the above embodiments of this application, which will not be repeated here.

[0211] The functions implemented by the above-mentioned first enhancement unit 1001, the second enhancement unit 1002, the processing unit 1003 and the sound amplification unit 1004 can be implemented by the same or different processors respectively, and the embodiment of the present application is not limited thereto.

[0212] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by the configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the remaining part is implemented in the form of hardware circuits.

[0213] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.

[0214] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0215] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0216] Exemplary Electronic Devices

[0217] An embodiment of the present application provides an electronic device, see Fig.11 As shown, the device includes:

[0218] Memory 200 and processor 210;

[0219] The memory 200 is connected to the processor 210 and is used to store programs;

[0220] The processor 210 is used to implement the speech enhancement method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0221] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .

[0222] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other via a bus.

[0223] A bus may include a pathway that transfers information between components of a computer system.

[0224] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0225] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0226] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0227] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0228] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0229] The communication interface 220 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0230] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any speech enhancement method provided in the above embodiments of the present application.

[0231] Exemplary computer program products and storage media

[0232] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the speech enhancement method according to various embodiments of the present application described in any of the above-mentioned embodiments of this specification.

[0233] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0234] In addition, the embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to execute the steps of the speech enhancement method according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically the following steps may be implemented:

[0235] Step 101: Perform voice enhancement on a signal received by a microphone to obtain a first enhanced signal.

[0236] Step 102: Perform speech enhancement on the signal in the low frequency band of the first enhanced signal to obtain a second enhanced signal.

[0237] Step 103: Determine a target enhanced signal based on the first enhanced signal and the second enhanced signal.

[0238] Step 104: amplify the target enhanced signal to obtain an amplified signal.

[0239] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0240] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0241] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0242] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0243] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0244] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0245] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0246] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0247] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0248] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0249] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech enhancement method, characterized in that: include: Performing voice enhancement on a signal received by a microphone to obtain a first enhanced signal; Performing speech enhancement on a signal in a low frequency band of the first enhanced signal to obtain a second enhanced signal; Determining a target enhanced signal based on the first enhanced signal and the second enhanced signal; The target enhanced signal is amplified to obtain an amplified signal.

2. The speech enhancement method according to claim 1, characterized in that: The step of performing voice enhancement on the microphone received signal to obtain a first enhanced signal includes: Perform Fourier transform on the microphone received signal to obtain frequency domain characteristics; Based on the frequency domain features, a power spectrum feature and a complex spectrum feature are obtained; Predicting a time-frequency domain mask based on the power spectrum feature and the complex spectrum feature; Based on the frequency domain features and the time-frequency domain mask, a first enhanced signal is obtained.

3. The speech enhancement method according to claim 2, characterized in that: The performing speech enhancement on the signal in the low frequency band of the first enhanced signal to obtain the second enhanced signal includes: Based on the power spectrum characteristics and the complex spectrum characteristics, obtaining filter estimation complex coefficients corresponding to the low frequency band; A second enhanced signal is obtained based on the estimated complex coefficients of the filter corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal.

4. The speech enhancement method according to claim 3, characterized in that: A filter order, a target frame and a frequency correspond to a filter estimation complex coefficient; in the frequency domain feature, the acquisition time of the target frame is greater than or equal to the acquisition time of the current frame; The step of obtaining a second enhanced signal based on the filter estimation complex coefficients corresponding to the low frequency band and the signal of the low frequency band in the first enhanced signal comprises: Determine an intermediate enhanced signal corresponding to any one of the filter orders and the low frequency band based on the estimated complex coefficients of the filter corresponding to each target frame in the estimated complex coefficients of the filter corresponding to any one of the filter orders and the low frequency band, and the signal of the low frequency band in the first enhanced signal; A second enhanced signal is obtained based on each filter order and the intermediate enhanced signal corresponding to the low frequency band.

5. The speech enhancement method according to claim 2, characterized in that: The determining a target enhanced signal based on the first enhanced signal and the second enhanced signal includes: Determining a weighting coefficient based on the power spectrum characteristics and the complex spectrum characteristics; Determine a third enhanced signal based on the first enhanced signal, the second enhanced signal and the weighting coefficient; Performing inverse Fourier transform on the third enhanced signal to obtain a fourth enhanced signal; Based on the fourth enhanced signal, a target enhanced signal is determined.

6. The speech enhancement method according to claim 5, characterized in that: The determining a target enhanced signal based on the fourth enhanced signal includes: Performing frequency shifting, phase modulation and half-wave rectification on the feedback signal in the fourth enhanced signal to remove the correlation between the feedback signal in the fourth enhanced signal and the target speech signal to obtain a fifth enhanced signal; Determining a target enhanced signal based on the fifth enhanced signal; The feedback signal refers to the signal played by the speaker and transmitted back to the microphone; the target voice signal refers to the user voice signal received by the microphone.

7. The speech enhancement method according to claim 6, characterized in that: The determining a target enhanced signal based on the fifth enhanced signal includes: determining a speaker output signal based on the fifth enhanced signal; Mapping the dynamic range of the speaker output signal to a preset dynamic range to obtain an adjusted speaker output signal; A target enhancement signal is determined based on the adjusted loudspeaker output signal.

8. A speech enhancement device, characterized in that: include: A first enhancement unit, configured to perform voice enhancement on a signal received by a microphone to obtain a first enhanced signal; A second enhancement unit, configured to perform speech enhancement on a signal in a low frequency band of the first enhanced signal to obtain a second enhanced signal; a processing unit, configured to determine a target enhanced signal based on the first enhanced signal and the second enhanced signal; The amplification unit is used to amplify the target enhanced signal to obtain an amplified signal.

9. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the speech enhancement method according to any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the speech enhancement method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, enable the processor to perform the speech enhancement method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Noise reducing apparatus and audio regeneration apparatus

    CN101304621A

  • Multimedia data processing method and device

    CN107426200A

  • Active noise reduction method and device and earphone

    CN110225429A

  • Howling suppression method and device, computer equipment and storage medium

    CN114333749A

  • Audio signal processing method, device and equipment and computer readable storage medium

    CN114822569A