Audio signal processing method, apparatus, electronic device and readable storage medium
By segmenting the band energy attenuation curve and performing linear fitting, the reverb suppression function is determined, and the problem of large linear fitting error in the prior art is solved, more accurate reverb suppression is achieved, and speech quality and recognition accuracy are improved.
Patent Information
- Application Number
- CN202211314023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-25
AI Technical Summary
When suppressing reverberation audio signals, the linear fitting error is large, resulting in poor suppression effect.
The band energy attenuation curve is divided into N-segment energy attenuation curves, and each segment is linearly fitted, and the reverb suppression function is determined based on the N linear fit curves to suppress the reverb part in the reverb audio signal.
Through segmented linear fitting, the accuracy and effect of suppressing reverberation audio signals are improved, fitting errors are reduced, and speech quality and recognition accuracy are improved.
Smart Images

Figure CN115604627B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio technology, and particularly relates to an audio signal processing method, apparatus, electronic device, and readable storage medium. Background Art
[0002] Voice dereverberation has become an important step in the audio signal processing process. An electronic device can suppress a reverberant audio signal by removing the late reverberation in the reverberant audio signal, thereby making the voice more full.
[0003] Currently, in order to obtain the late reverberation in a reverberant audio signal, an electronic device can perform linear fitting on the energy decay curve of each frequency band in the energy decay curve of the room impulse response (RIR) over the entire time axis, and obtain the slope of each sub-band energy decay curve through the least squares method. Then, the energy decay process of the RIR can be modeled and described based on the obtained slope, so that the late reverberation can be estimated.
[0004] However, according to the above method, in the first few frames with a relatively high residual energy of the direct audio, the fitting error of the linear fitting value is usually large, which will result in poor accuracy of the late reverberation estimated by the above linear fitting, thereby leading to a poor effect of suppressing the reverberant audio signal. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide an audio signal processing method, apparatus, electronic device, and readable storage medium, which can solve the problem of poor effect of suppressing reverberant audio signals.
[0006] In a first aspect, the embodiments of this application provide an audio signal processing method, which includes: dividing the frequency band energy decay curve into N energy decay curves, and performing linear fitting on each energy decay curve to obtain N linear fitting curves, where N is an integer greater than or equal to 2; determining a reverberation suppression function corresponding to a target time frame in the N energy decay curves based on the N linear fitting curves; suppressing the reverberant part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function to obtain a second audio signal; where the first audio signal is: the audio signal within the frequency band corresponding to the frequency band energy decay curve; the frequency band energy decay curve is: one of the frequency band energy decay curves in the RIR energy decay curve.
[0007] In a second aspect, an embodiment of the present application provides an audio signal processing device, which includes a processing module, a determination module, and a suppression module; the processing module is configured to divide the band energy attenuation curve into N energy attenuation curves, and perform linear fitting on each energy attenuation curve to obtain N linear fitting curves, where N is an integer greater than or equal to 2; the determination module is configured to determine a reverberation suppression function corresponding to the target time frame in the N energy attenuation curves based on the N linear fitting curves obtained by the processing module; the suppression module is configured to suppress the reverberation part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function determined by the determination module to obtain a second audio signal; wherein, the first audio signal is: the audio signal in the audio signal that is within the frequency band corresponding to the band energy attenuation curve; the band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve.
[0008] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, and a program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, and the communication interface is coupled to the processor. The processor is configured to run a program or instruction to implement the method described in the first aspect.
[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0012] In an embodiment of the present application, the band energy attenuation curve can be divided into N energy attenuation curves, and each energy attenuation curve is linearly fitted to obtain N linear fitting curves, where N is an integer greater than or equal to 2; and based on the N linear fitting curves, a reverberation suppression function corresponding to a target time frame in the energy attenuation curve is determined; and based on the reverberation suppression function, the reverberation part in the reverberant audio signal of the target time frame in the first audio signal is suppressed to obtain a second audio signal; where the first audio signal is: the audio signal in the audio signal within the frequency band corresponding to the band energy attenuation curve; the band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve. Through this solution, since the electronic device can divide one of the band energy attenuation curves in the RIR energy attenuation curve into N energy attenuation curves and perform linear fitting respectively, and can determine the reverberation suppression function corresponding to the target time frame in the N energy attenuation curves based on the obtained N linear fitting curves, so as to suppress the reverberation part in the reverberant audio signal of the target time frame in the audio signal within the frequency band corresponding to the band energy attenuation curve, it is possible to accurately suppress the reverberant audio signal of each time frame through piecewise linear fitting with a small fitting error and the reverberation suppression function corresponding to each time frame, thereby improving the effect of suppressing the reverberant audio signal. Description of the Drawings
[0013] Figure 1 It is a schematic diagram of the generation process of the reverberant audio signal;
[0014] Figure 2 It is a schematic diagram of the RIR energy attenuation curve;
[0015] Figure 3 It is a schematic diagram of the linear fitting in traditional speech dereverberation;
[0016] Figure 4 It is a flowchart of the audio signal processing method provided by an embodiment of the present application;
[0017] Figure 5 It is one of the schematic diagrams of the audio signal processing method provided by an embodiment of the present application;
[0018] Figure 6 It is another schematic diagram of the audio signal processing method provided by an embodiment of the present application;
[0019] Figure 7 It is a schematic diagram of the audio signal processing device provided by an embodiment of the present application;
[0020] Figure 8 It is a schematic diagram of the electronic device provided by an embodiment of the present application;
[0021] Figure 9It is a schematic diagram of the hardware of the electronic device provided by the embodiment of the present application. Specific Embodiments
[0022] The following will clearly describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0024] First, some nouns or terms involved in the specification and claims of the present application will be explained.
[0025] RT60 (i.e., Reverberation Time - 60dB): The time required for the sound field to decay by 60dB.
[0026] The following will describe in detail the audio signal processing method, device, electronic device, and readable storage medium provided by the embodiments of the present application with reference to the accompanying drawings and specific embodiments and their application scenarios.
[0027] Voice dereverberation is a technology widely used in audio devices, commonly found in devices such as mobile phones, speakers, and conference call devices.
[0028] In the enclosure space, a sound source continuously emits an audio signal. The emitted audio signal will continuously reflect during the propagation process due to the presence of obstacles. At the same time, the energy of the audio signal will gradually decay during this process. The audio signal after decaying energy reaches the pickup device after a certain delay, and it is collected by the pickup device together with the direct audio signal at the current moment, causing the direct audio signal at the current moment to be interfered by the reflected audio signal, forming a reverberant audio signal, and the energy of the reverberant audio signal will become stronger as the distance between the sound source and the pickup device increases.
[0029] Figure 1 A schematic diagram showing the generation process of the reverberant audio signal is shown, as Figure 1As shown in the figure, a microphone 11 and a speaker 12 are placed in the box space 10, and the propagation medium is air. Assuming that the attenuation coefficient of sound propagation in air is α and the reflection coefficient of the wall of the box space 10 is β, the audio signal emitted by the speaker 12 at time t1 is The audio signal emitted at time t2 is The audio signal at time t1 reaches the microphone 11 at time t2 after being reflected and propagated, and ignoring the propagation time of the direct audio signal, the signal received by the microphone 11 at time t1 is The signal received by the microphone 11 at time t2 is Where That is, the reverberant audio signal.
[0030] Due to the existence of the reverberant audio signal, the speech quality will be greatly reduced, affecting the subjective listening experience of users. Moreover, in some intelligent devices, it will also affect the accuracy of speech recognition. Therefore, speech dereverberation has become an important step in the field of audio signal processing.
[0031] Usually, the generation of the reverberant audio signal is to convolve the clean speech and the RIR, as shown in the following formula (1):
[0032]
[0033] Where, z(n) is the reverberant audio signal, h(n) is the RIR, and s(n) is the clean speech; after performing the Fourier transform on the above formula (1) and converting it to the time domain, it is shown in the following formula (2):
[0034]
[0035] Where, m represents the time frame and k represents the frequency point. And the reverberant audio signal is usually divided into the early reverberant audio signal and the late reverberant audio signal. After squaring the above formula (2), it can be expressed as the following formula (3):
[0036] λ z (m,k) = λ ze (m,k) + λ zl (m,k); (3)
[0037] Where, λ z (m,k) represents the energy of the reverberant audio signal at the m-th frame and the k-th frequency point, that is, the spectral variance of the reverberant audio signal, and λ ze (m,k) represents the early reverberation energy (spectral variance) at the m-th frame and the k-th frequency point, and λ zl(m, k) represents the late reverberation energy (spectral variance) at the m-th frame and the k-th frequency bin. Generally, the part that affects speech quality is the late reverberation audio signal. During the dereverberation process, only removing the late reverberation audio signal while retaining the early reverberation audio signal can make the speech more full-bodied and have a better listening experience. Generally speaking, the reflected energy within the delay range of 50 ms - 80 ms after a pulse signal is emitted belongs to the early reverberation energy, and the energy after that is the late reverberation energy. To better remove the late reverberation audio signal while retaining the early reverberation audio signal, it is necessary to accurately describe and model the RIR.
[0038] Figure 2 A schematic diagram showing the RIR energy decay curve is as Figure 2 shown. This RIR energy decay curve is an RIR energy decay curve with an RT60 of approximately 900 ms. Among them, the horizontal axis is the time frame, the vertical axis is the energy in dB, the sampling rate is 16 kHz, the frame length of the short-time Fourier transform is 512, the frame shift is 160. This RIR energy decay curve includes multiple curves, and each curve represents the trend of the energy of a sub-band changing with time. The average value of 32 frequency bins is taken for each sub-band, and the first sub-band removes the DC component.
[0039] In traditional speech dereverberation, a linear fitting is performed on the RIR energy decay curve on the entire time axis. For example, as Figure 3 shown, curve 31 is the sub-band energy decay curve from the 65th frequency bin to the 96th frequency bin, and curve 32 is the curve obtained by linearly fitting curve 31 on the entire time axis. After obtaining the linearly fitted curve, the slope of this curve can be obtained by the least squares method, and thus T60 can be obtained through the following formula (4):
[0040]
[0041] And the related parameter α(k) of the frequency is defined as the following formula (5):
[0042]
[0043] where fs is the sampling rate. Thus, the energy λ of the direct audio signal of the m-th frame can be obtained through the following formula (6): s (m, k) The energy E(i, k) after attenuation through i frames:
[0044] E(i, k) = e -2α(k)Ri λ s (m, k); (6)
[0045] where R represents the frame shift. In this way, the energy decay process of the RIR can be modeled and described, and the late reverberation energy λ can be deduced. zl(m, k).
[0046] However, according to the above method, the linear fitting of the RIR energy attenuation curve by the electronic device is based on the entire time axis, but this global linear fitting cannot achieve global optimality, and the specific manifestations are as follows:
[0047] 1. In the first few frames where the residual energy of the direct audio is relatively high, the fitting error of the linear fitting value is relatively large;
[0048] 2. According to the above formulas (5) and (6), the following formula (7) can be obtained:
[0049]
[0050] If we denote Then the above formula (7) can be expressed as the following formula (8):
[0051] E(i, k) = ε i ·λ s (m, k); (8)
[0052] It can be seen that 0 < ε < 1, so ε i will decrease as i increases, which means that the residual energy of the direct audio signal in the m-th frame after attenuation in subsequent time frames is different, and the closer to the m-th frame in time, the greater the residual energy and the higher the impact on the reverberation component. Obviously, in the first few frames, the residual energy of the direct audio signal is relatively high, while the fitting error of the linear fitting value in these time frames is relatively large, which has a huge impact on the estimation of the reverberation component, thus resulting in a poor effect of suppressing the reverberant audio signal.
[0053] To solve the above problems, in the audio signal processing method provided in the embodiments of the present application, the band energy attenuation curve can be divided into N energy attenuation curves, and each energy attenuation curve is linearly fitted to obtain N linear fitting curves, where N is an integer greater than or equal to 2; and based on the N linear fitting curves, a reverberation suppression function corresponding to the target time frame in the N energy attenuation curves is determined; and based on the reverberation suppression function, the reverberation part in the reverberant audio signal of the target time frame in the first audio signal is suppressed to obtain a second audio signal; wherein, the first audio signal is: the audio signal in the audio signal within the frequency band corresponding to the band energy attenuation curve; the band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve. Through this solution, since the electronic device can divide one of the band energy attenuation curves in the RIR energy attenuation curve into N energy attenuation curves and perform linear fitting respectively, and can determine the reverberation suppression function corresponding to the target time frame in the N energy attenuation curves based on the obtained N linear fitting curves, so as to suppress the reverberation part in the reverberant audio signal of the target time frame in the audio signal within the frequency band corresponding to the band energy attenuation curve, the reverberant audio signals of each time frame can be accurately suppressed through piecewise linear fitting with a small fitting error and the reverberation suppression function corresponding to each time frame, thereby improving the effect of suppressing the reverberant audio signal.
[0054] Embodiments of the present application provide an audio signal processing method. Figure 4 The flowchart of the audio signal processing method provided in the embodiments of the present application is shown. As Figure 4 shown, the audio signal processing method provided in the embodiments of the present application may include the following steps 401 to step 403. Hereinafter, taking an electronic device executing this method as an example, this method will be described exemplarily.
[0055] Step 401, the electronic device divides the band energy attenuation curve into N energy attenuation curves, and linearly fits each energy attenuation curve to obtain N linear fitting curves.
[0056] Among them, N is an integer greater than or equal to 2.
[0057] In the embodiments of the present application, the above band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve.
[0058] Optionally, in the embodiments of the present application, when N is equal to 3, that is, when the band energy attenuation curve is divided into 3 energy attenuation curves, the optimal piecewise linear fitting effect can be achieved; of course, in actual implementation, N can be any integer greater than or equal to 2, which is not limited in the embodiments of the present application.
[0059] For the description of the electronic device dividing the above-mentioned band energy attenuation curve into N segments of energy attenuation curves and performing piecewise linear fitting, reference can be made to the specific description of piecewise linear regression in the related art. To avoid repetition, it will not be elaborated here.
[0060] Next, the audio signal processing method provided by the embodiments of the present application will be exemplarily described with reference to the accompanying drawings.
[0061] Exemplarily, as Figure 5 shown, the electronic device divides the energy attenuation curve 50 (i.e., the above-mentioned band energy attenuation curve) into 3 segments of energy attenuation curves according to the time frames m1 and m2 on the time frame, and after performing linear fitting on each segment of the energy attenuation curve, obtains the linear fitting curve 51, the linear fitting curve 52, and the linear fitting curve 53 (i.e., the above-mentioned N linear fitting curves).
[0062] Step 402: The electronic device determines the reverberation suppression function corresponding to the target time frame in the N segments of energy attenuation curves based on the N linear fitting curves.
[0063] In the embodiments of the present application, the reverberation suppression function is used to suppress the reverberation part in the reverberant audio signal.
[0064] It should be noted that the above-mentioned reverberation part is not a separate audio signal, but the reverberation energy in the reverberant audio signal, that is, the energy generated during the propagation of the audio signal in the box; if there is no RIR or the clean audio signal in the audio signal, then the reverberant audio signal in the audio signal does not exist either.
[0065] Optionally, in the embodiments of the present application, the target time frame can be any time frame.
[0066] Optionally, in the embodiments of the present application, the target time frame can be any time frame after the 5th frame.
[0067] Optionally, in the embodiments of the present application, the above step 402 can be specifically implemented through the following steps 402a to 402c.
[0068] Step 402a: The electronic device calculates the reverberation weight corresponding to each linear fitting curve based on the slope of each linear fitting curve among the N linear fitting curves, so as to obtain N reverberation weights.
[0069] Optionally, in the embodiments of the present application, the electronic device can calculate the reverberation weight corresponding to each of the above linear fitting curves through the above formulas (4) and (5).
[0070] Exemplarily, assume that the above N linear fitting curves are linear fitting curve α, linear fitting curve β, and linear fitting curve γ. Then, the electronic device can calculate the reverberation weight α(k) corresponding to the linear fitting curve α, the reverberation weight β(k) corresponding to the linear fitting curve β, and the reverberation weight γ(k) corresponding to the linear fitting curve γ according to the slope of each linear fitting curve through the above formulas (4) and (5), as shown in the following formula (9):
[0071]
[0072]
[0073]
[0074] Step 402b: The electronic device calculates the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal based on the N reverberation weights.
[0075] In the embodiments of the present application, the above reverberant audio signal is the reverberant audio signal of the target time frame in the first audio signal.
[0076] In the embodiments of the present application, the first audio signal is the audio signal within the frequency band corresponding to the above frequency band energy attenuation curve in the audio signal.
[0077] It can be understood that the above early reverberation energy and late reverberation energy are determined by the direct audio signal of each time frame before the target time frame (i.e., the clean audio signal in the first audio signal).
[0078] Optionally, in the embodiments of the present application, the above step 402b can be specifically implemented through the following steps 402b1 and 402b2.
[0079] Step 402b1: For each time frame before the target time frame, the electronic device calculates the remaining energy of the energy of the direct audio signal of one time frame in the first audio signal at the target time frame according to one time frame and the reverberation weight corresponding to one time frame, and obtains the remaining energy corresponding to each time frame.
[0080] Optionally, in the embodiments of the present application, the electronic device can calculate the remaining energy corresponding to each time frame through the above formula (6).
[0081] Step 402b2: The electronic device calculates the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal according to the remaining energy corresponding to each time frame before the target time frame.
[0082] It should be noted that in the embodiments of the present application, the above-mentioned band energy attenuation curve is divided into 3 energy attenuation curves, and the target time frame is the m-th frame as an example. In actual implementation, the specific number of segments and the target time frame are not limited.
[0083] Optionally, in the embodiments of the present application, the electronic device may derive the above-mentioned early reverberation energy λ ze (m,k) and the expression of the late reverberation energy λ zl (m,k) according to the above formula (8), as shown in the following formula (10) and formula (11):
[0084]
[0085] Where m3 represents that the energy of the direct audio signal in the m-th frame is just greater than the preset threshold in the m3-th frame after it. For example, assuming that the energy of the direct audio signal in the first frame is -10 dB, and after 20 frames of attenuation, the energy becomes -58 dB, and the energy in the 21st frame becomes -65 dB. If the preset threshold is -60 dB, then m3 = 20. It can be understood that when the energy of the direct audio signal decays to a certain extent, its impact on the whole can be ignored. The existence of m3 also makes formula (11) a finite polynomial, which is more operable for engineering practice.
[0086] In the embodiments of the present application, since the electronic device can calculate the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal according to the remaining energy corresponding to each time frame obtained, the accuracy of the electronic device in calculating the early reverberation energy and the late reverberation energy can be improved.
[0087] Step 402c: The electronic device determines the reverberation suppression function corresponding to the target time frame in the N energy attenuation curves based on the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal.
[0088] Optionally, in the embodiments of the present application, the above step 402c may be specifically implemented by the following steps 402c1 to 402c3.
[0089] Step 402c1: The electronic device calculates the a priori signal-to-noise ratio corresponding to the target time frame according to the early reverberation energy of the reverberant audio signal, the late reverberation energy of the reverberant audio signal, and the energy of the environmental noise audio signal in the target time frame of the first audio signal.
[0090] Optionally, in the embodiments of the present application, the first audio signal may include a direct audio signal, a reverberant audio signal, and an environmental noise audio signal. Then the first audio signal may be expressed as the following formula (12):
[0091]
[0092] Among them, v(n) is the environmental noise audio signal; performing a short-time Fourier transform on the above formula (12), and according to the above formula (2) and formula (3), the following formula (13) can be obtained:
[0093] |Y(m,k)| 2 =λ ze (m,k)+λ zl (m,k)+λ v (m,k); (13)
[0094] Among them, |Y(m,k)| 2 represents the square of the amplitude spectrum of the first audio signal, and λ v (m,k) represents the energy of the above environmental noise audio signal; thus, the energy of the environmental noise audio signal can be calculated, and the above prior signal-to-noise ratio ε(m,k) can be calculated through the following formula (14):
[0095]
[0096] Step 402c2: The electronic device calculates the posterior signal-to-noise ratio corresponding to the target time frame according to the late reverberation energy of the reverberant audio signal, the energy of the environmental noise audio signal in the target time frame of the first audio signal, and the amplitude spectrum of the first audio signal in the target time frame.
[0097] Optionally, in the embodiments of the present application, after the electronic device obtains the above λ v (m,k), the above posterior signal-to-noise ratio ζ(m,k) can be calculated through the following formula (15):
[0098]
[0099] Step 402c3: The electronic device determines the reverberation suppression function corresponding to the target time frame in the N energy decay curves according to the prior signal-to-noise ratio corresponding to the target time frame and the posterior signal-to-noise ratio corresponding to the target time frame.
[0100] Optionally, in the embodiments of the present application, after the electronic device obtains the above prior signal-to-noise ratio λ v (m,k) and the posterior signal-to-noise ratio ζ(m,k), the above reverberation suppression function can be determined, and the reverberation suppression function can be expressed as the following formula (16):
[0101]
[0102] In the embodiments of the present application, since the electronic device can determine the above reverberation suppression function based on the calculated prior signal-to-noise ratio and posterior signal-to-noise ratio corresponding to the target time frame, the accuracy of the electronic device in determining the reverberation suppression function can be improved, so that the reverberant audio signal of the target time frame can be accurately suppressed by the reverberation suppression function.
[0103] In the embodiments of the present application, since the electronic device can calculate the reverberation weights corresponding to each of the above N linear fitting curves based on the slopes of each of the N linear fitting curves, and calculate the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal based on the obtained N reverberation weights to determine the above reverberation suppression function, the accuracy of the electronic device in determining the reverberation suppression function can be further improved.
[0104] Step 403: The electronic device suppresses the reverberation part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function corresponding to the target time frame in the N energy decay curves, so as to obtain a second audio signal.
[0105] In the embodiments of the present application, the second audio signal is: the estimated direct audio signal of the target time frame after suppressing the above reverberation part.
[0106] Optionally, in the embodiments of the present application, the above step 403 may be specifically implemented by the following step 403a and step 403b.
[0107] Step 403a: The electronic device performs a dot product operation on the reverberation suppression function corresponding to the target time frame in the N energy decay curves and the amplitude spectrum of the first audio signal at the target time frame to obtain a target amplitude spectrum.
[0108] Optionally, in the embodiments of the present application, the target amplitude spectrum is: the amplitude spectrum of the first audio signal after suppressing the reverberant audio signal, and the target amplitude spectrum can be calculated by the following formula (17):
[0109]
[0110] Step 403b: The electronic device performs an inverse Fourier transform on the target amplitude spectrum and the phase of the first audio signal at the target time frame to obtain a second audio signal.
[0111] Optionally, in the embodiments of the present application, the inverse Fourier transform can restore the audio signal from the frequency domain back to the time domain.
[0112] In the embodiments of the present application, since the electronic device can perform a dot product operation on the target amplitude spectrum obtained by multiplying the above reverberation suppression function by the amplitude spectrum of the first audio signal in the target time frame, and perform an inverse Fourier transform on the phase of the first audio signal in the target time frame to obtain the second audio signal, the reverberant audio signal can be accurately suppressed by this reverberation suppression function, thereby improving the robustness and flexibility of suppressing the reverberant audio signal.
[0113] It should be noted that the electronic device can suppress the reverberant audio signal in each time frame of the first audio signal through the above steps, and further can suppress the reverberant audio signal in each frequency band of the collected audio signal, so as to achieve reverberation suppression of the entire collected audio signal.
[0114] In the audio signal processing method provided in the embodiments of the present application, since the electronic device can divide one frequency band energy attenuation curve in the RIR energy attenuation curve into N segment energy attenuation curves and perform linear fitting respectively, and can determine the reverberation suppression function corresponding to the target time frame in the N segment energy attenuation curves based on the obtained N linear fitting curves, so as to suppress the reverberation part in the reverberant audio signal of the target time frame in the audio signal in the frequency band corresponding to the frequency band energy attenuation curve, the reverberant audio signal in each time frame can be accurately suppressed through the piecewise linear fitting with a small fitting error and the reverberation suppression function corresponding to each time frame, thereby improving the effect of suppressing the reverberant audio signal.
[0115] Next, the audio signal processing method provided in the embodiments of the present application will be exemplarily described with reference to the accompanying drawings.
[0116] Exemplarily, assuming that the sampling rate is 16 kHz, the frame length of the short-time Fourier transform is 512, and the frame offset is 160, then the time represented by one frame is 10 ms. Taking m1 = 2 and m2 = 5, if the 5th frame is the boundary between the early reverberant audio signal and the late reverberant audio signal, then the 1st frame to the 5th frame are the early reverberation part, and after the 5th frame is the late reverberation part. Without considering background noise, there is the following derivation:
[0117] The 1st frame:
[0118] λ z (1,k) = α(k)λ s (1,k)
[0119] λ ze (1,k) = λ z (1,k)
[0120] λ zl (1,k) = 0;
[0121] The 2nd frame:
[0122] λ z (2,k) = α(k)λ s (2,k) + α 2 (k)λ s (1,k)
[0123] λ ze (2,k) = λ z (2,k)
[0124] λ zl (2,k) = 0;
[0125] Frame 3:
[0126] λ z (3,k) = α(k)λ s (3,k) + α 2 (k)λ s (2,k) + α 2 (k)β(k)λ s (1,k)
[0127] λ ze (3,k) = λ z (3,k)
[0128] λ zl (3,k) = 0;
[0129] Frame 4:
[0130] λ z (4,k) = α(k)λ s (4,k) + α 2 (k)λ s (3,k) + α 2 (k)β(k)λ s (2,k) + α 2 (k)β 2 (k)λ s (1,k)
[0131] λ ze (4,k) = λ z (4,k)
[0132] λ zl (4,k) = 0;
[0133] Frame 5:
[0134] λ z (5,k) = α(k)λ s (5,k) + α 2 (k)λ s (4,k) + α 2 (k)β(k)λs (3,k) + α 2 (k)β 2 (k)λ s (2,k) + α 2 (k)β 3 (k)λ s (1,k)
[0135] λ ze (5,k) = λ z (5,k)
[0136] λ zl (5,k) = 0;
[0137] Frame 6:
[0138] λ z (6,k) = α(k)λ s (6,k) + α 2 (k)λ s (5,k) + α 2 (k)β(k)λ s (4,k) + α 2 (k)β 2 (k)λ s (3,k) + α 2 (k)β 3 (k)λ s (2,k) + α 2 (k)β 3 (k)γ(k)λ s (1,k)
[0139] λ ze (6,k) = α(k)λ s (6,k) + α 2 (k)λ s (5,k) + α 2 (k)β(k)λ s (4,k) + α 2 (k)β 2 (k)λ s (3,k) + α 2 (k)β 3 (k)λ s (2,k)
[0140] λ zl (6,k) = α 2 (k)β 3 (k)γ(k)λ s (1,k);
[0141] Frame 7:
[0142] λz (7, k) = α(k)λ s (7, k) + α 2 (k)λ s (6, k) + α 2 (k)β(k)λ s (5, k) + α 2 (k)β 2 (k)λ s (4, k) + α 2 (k)β 3 (k)λ s (3, k) + α 2 (k)β 3 (k)γ(k)λ s (2, k) + α 2 (k)β 3 (k)γ 2 (k)λ s (1, k)
[0143] λ ze (7, k) = α(k)λ s (7, k) + α 2 (k)λ s (6, k) + α 2 (k)β(k)λ s (5, k) + α 2 (k)β 2 (k)λ s (4, k) + α 2 (k)β 3 (k)λ s (3, k)
[0144] λ zl (7, k) = α 2 (k)β 3 (k)γ(k)λ s (2, k) + α 2 (k)β 3 (k)γ 2 (k)λ s (1, k);
[0145] ……
[0146] By analogy, from the above derivation, it can be seen that when m1 = 2 and m2 = 5, starting from the 6th frame, the number of terms of λ ze is constantly 5, and the number of terms of λ zl increases with the increase of the number of frames, but when the energy of the newly added term in each frame is less than the set threshold (usually -60 dB), it is not considered, that is, λ zlThe number of items is also constant at this time, corresponding to the above-mentioned m3 frame here and corresponding to the above formula (11).
[0147] Figure 6 The figure shows a schematic diagram of the effect of suppressing the reverberation part in the reverberant audio signal by using the audio signal processing method according to the embodiment of the present application, as Figure 6 shown. In region 61 is the spectrogram of clean speech (i.e., the direct audio signal), in region 62 is the reverberant speech (i.e., the reverberant audio signal) obtained by convolving the clean speech with the RIR, and in region 62 is the speech after dereverberation (i.e., the second audio signal); it can be seen that the speech after dereverberation basically restores the harmonic structure of the clean speech, and the reverberant speech is effectively suppressed, thereby improving the speech quality and intelligibility of the speech.
[0148] For the audio signal processing method provided by the embodiment of the present application, the execution subject can be an audio signal processing device. In the embodiment of the present application, taking the audio signal processing device executing the audio signal processing method as an example, the audio signal processing device provided by the embodiment of the present application is described.
[0149] Combined with Figure 7 , the embodiment of the present application provides an audio signal processing device 70. The audio signal processing device 70 may include a processing module 71, a determination module 72, and a suppression module 73. The processing module 71 can be used to divide the band energy attenuation curve into N segment energy attenuation curves, and perform linear fitting on each segment energy attenuation curve to obtain N linear fitting curves, where N is an integer greater than or equal to 2. The determination module 72 can be used to determine the reverberation suppression function corresponding to the target time frame in the N segment energy attenuation curves based on the N linear fitting curves processed by the processing module 71. The suppression module 73 can be used to suppress the reverberation part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function determined by the determination module 72 to obtain a second audio signal. Wherein, the first audio signal is: the audio signal in the audio signal within the frequency band corresponding to the band energy attenuation curve; the band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve.
[0150] In a possible implementation manner, the determination module 72 can specifically be used to calculate the reverberation weight corresponding to each linear fitting curve respectively based on the slope of each linear fitting curve among the above N linear fitting curves to obtain N reverberation weights; and calculate the early reverberation energy of the above reverberant audio signal and the late reverberation energy of the above reverberant audio signal based on the N reverberation weights; and determine the above reverberation suppression function based on the early reverberation energy and the late reverberation energy.
[0151] In a possible implementation, the determination module 72 can specifically be used to, for each time frame before the target time frame, calculate the remaining energy of the direct audio signal of the time frame in the target time frame in the first audio signal according to a time frame and the reverberation weight corresponding to the time frame, so as to obtain the remaining energy corresponding to each time frame; and calculate the above-mentioned early reverberation energy and late reverberation energy according to the remaining energy corresponding to each time frame.
[0152] In a possible implementation, the determination module 72 can specifically be used to calculate the prior signal-to-noise ratio corresponding to the target time frame according to the above-mentioned early reverberation energy, late reverberation energy, and the energy of the ambient noise audio signal of the target time frame in the first audio signal; and calculate the posterior signal-to-noise ratio corresponding to the target time frame according to the late reverberation energy, the energy of the ambient noise audio signal, and the amplitude spectrum of the first audio signal at the target time frame; and determine the above-mentioned reverberation suppression function according to the prior signal-to-noise ratio and the posterior signal-to-noise ratio.
[0153] In a possible implementation, the suppression module 73 can specifically be used to perform a dot product operation on the above-mentioned reverberation suppression function and the amplitude spectrum of the first audio signal at the target time frame to obtain a target amplitude spectrum; and perform an inverse Fourier transform on the target amplitude spectrum and the phase of the first audio signal at the target time frame to obtain a second audio signal.
[0154] In the audio signal processing device provided in the embodiments of the present application, since the audio signal processing device can divide an energy attenuation curve of a frequency band in the RIR energy attenuation curve into N energy attenuation curves and perform linear fitting respectively, and can determine the reverberation suppression function corresponding to the target time frame in the N energy attenuation curves based on the obtained N linear fitting curves, so as to suppress the reverberation part in the reverberant audio signal of the target time frame in the audio signal in the frequency band corresponding to the energy attenuation curve of the frequency band, therefore, accurate suppression of the reverberant audio signal of each time frame can be achieved through piecewise linear fitting with a small fitting error and the reverberation suppression function corresponding to each time frame, thereby improving the effect of suppressing the reverberant audio signal.
[0155] The audio signal processing device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0156] The audio signal processing device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0157] The audio signal processing device provided by the embodiments of the present application can implement Figures 4 to 6 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.
[0158] As Figure 8 shown, the embodiments of the present application further provide an electronic device 800, including a processor 801 and a memory 802. A program or instruction that can run on the processor 801 is stored on the memory 802. When the program or instruction is executed by the processor 801, it implements each step of the above-mentioned audio signal processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0159] It should be noted that the electronic device in the embodiments of the present application includes the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0160] Figure 9 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.
[0161] The electronic device 1000 includes, but is not limited to, components such as a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.
[0162] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for powering each component. The power source can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0163] Among them, the processor 1010 can be used to divide the band energy attenuation curve into N energy attenuation curves, perform linear fitting on each energy attenuation curve to obtain N linear fitting curves, where N is an integer greater than or equal to 2; and based on the N linear fitting curves obtained by processing, determine the reverberation suppression function corresponding to the target time frame in the N energy attenuation curves; and based on the determined reverberation suppression function, suppress the reverberation part in the reverberant audio signal of the target time frame in the first audio signal to obtain a second audio signal. Wherein, the first audio signal is: the audio signal in the audio signal within the frequency band corresponding to the band energy attenuation curve; the band energy attenuation curve is: one of the band energy attenuation curves in the RIR energy attenuation curve.
[0164] In a possible implementation, the processor 1010 can specifically be used to calculate the reverberation weight corresponding to each linear fitting curve based on the slope of each linear fitting curve among the above N linear fitting curves to obtain N reverberation weights; and based on the N reverberation weights, calculate the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal; and based on the early reverberation energy and the late reverberation energy, determine the above reverberation suppression function.
[0165] In a possible implementation, the processor 1010 can specifically be used for each time frame before the target time frame, calculate the remaining energy of the direct audio signal energy of the one time frame in the first audio signal at the target time frame according to one time frame and the reverberation weight corresponding to the one time frame, to obtain the remaining energy corresponding to each time frame; and calculate the above early reverberation energy and late reverberation energy according to the remaining energy corresponding to each time frame.
[0166] In a possible implementation, the processor 1010 may specifically be configured to calculate a priori signal-to-noise ratio corresponding to a target time frame according to the above-mentioned early reverberation energy, late reverberation energy, and the energy of the environmental noise audio signal of the target time frame in the first audio signal; and calculate a posteriori signal-to-noise ratio corresponding to the target time frame according to the late reverberation energy, the energy of the environmental noise audio signal, and the amplitude spectrum of the first audio signal at the target time frame; and determine the above-mentioned reverberation suppression function according to the a priori signal-to-noise ratio and the a posteriori signal-to-noise ratio.
[0167] In a possible implementation, the processor 1010 may specifically be configured to perform a dot product operation on the above-mentioned reverberation suppression function and the amplitude spectrum of the first audio signal at the target time frame to obtain a target amplitude spectrum; and perform an inverse Fourier transform on the target amplitude spectrum and the phase of the first audio signal at the target time frame to obtain a second audio signal.
[0168] In the electronic device provided in the embodiments of the present application, since the electronic device can divide an energy attenuation curve of a frequency band in the RIR energy attenuation curve into N energy attenuation curves and perform linear fitting on each of them respectively, and can determine a reverberation suppression function corresponding to a target time frame in the N energy attenuation curves based on the obtained N linear fitting curves, so as to suppress the reverberation part in the reverberant audio signal of the target time frame in the audio signal within the frequency band corresponding to the energy attenuation curve of the frequency band, it is possible to accurately suppress the reverberant audio signal of each time frame through piecewise linear fitting with a small fitting error and the reverberation suppression function corresponding to each time frame, thereby improving the effect of suppressing the reverberant audio signal.
[0169] For the beneficial effects of various implementation manners in this embodiment, reference may specifically be made to the beneficial effects of the corresponding implementation manners in the above method embodiments. To avoid repetition, they will not be elaborated here.
[0170] It should be understood that in the embodiments of the present application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of still pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0171] The memory 1009 can be used to store software programs and various data. The memory 1009 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0172] The processor 1010 can include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010.
[0173] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the audio signal processing method embodiment as described above and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0174] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.
[0175] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the audio signal processing method embodiment as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0176] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system or system-on-chip, etc.
[0177] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the audio signal processing method embodiment as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0178] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0180] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. An audio signal processing method, characterized in that, The method includes: Dividing the band energy decay curve into N energy decay curves, and performing linear fitting on each of the energy decay curves to obtain N linear fitting curves, where N is an integer greater than or equal to 2; Based on the N linear fitting curves, determining a reverberation suppression function corresponding to a target time frame in the N energy decay curves; Based on the reverberation suppression function, suppressing the reverberation part in the reverberant audio signal of the target time frame in the first audio signal to obtain a second audio signal; Wherein, the first audio signal is: the audio signal in the audio signal within the frequency band corresponding to the band energy decay curve; The band energy decay curve is: one of the band energy decay curves in the room impulse response RIR energy decay curve.
2. The method according to claim 1, wherein The determining the reverberation suppression function corresponding to the target time frame in the energy decay curve based on the N linear fitting curves includes: Based on the slope of each of the N linear fitting curves, calculating the reverberation weight corresponding to each of the linear fitting curves to obtain N reverberation weights; Based on the N reverberation weights, calculating the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal; Based on the early reverberation energy and the late reverberation energy, determining the reverberation suppression function.
3. The method according to claim 2, wherein The calculating the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal based on the N reverberation weights includes: For each time frame before the target time frame, according to a time frame and the reverberation weight corresponding to the time frame, calculating the remaining energy of the direct audio signal of the time frame in the first audio signal at the target time frame to obtain the remaining energy corresponding to each time frame; According to the remaining energy corresponding to each time frame, calculating the early reverberation energy and the late reverberation energy.
4. The method according to claim 2, wherein The determining the reverberation suppression function based on the early reverberation energy and the late reverberation energy includes: According to the early reverberation energy, the late reverberation energy, and the energy of the ambient noise audio signal of the target time frame in the first audio signal, calculating the a priori signal-to-noise ratio corresponding to the target time frame; According to the late reverberation energy, the energy of the ambient noise audio signal, and the amplitude spectrum of the first audio signal at the target time frame, calculating the a posteriori signal-to-noise ratio corresponding to the target time frame; According to the a priori signal-to-noise ratio and the a posteriori signal-to-noise ratio, determining the reverberation suppression function.
5. The method according to claim 1, wherein The suppressing the reverberation part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function to obtain a second audio signal includes: Performing a dot product operation on the reverberation suppression function and the amplitude spectrum of the first audio signal at the target time frame to obtain a target amplitude spectrum; Performing an inverse Fourier transform on the target amplitude spectrum and the phase of the first audio signal at the target time frame to obtain the second audio signal.
6. An audio signal processing device, characterized in that, The device includes a processing module, a determining module, and a suppressing module; The processing module is configured to divide the band energy attenuation curve into N energy attenuation curves, and perform linear fitting on each of the energy attenuation curves to obtain N linear fitting curves, where N is an integer greater than or equal to 2; The determining module is configured to determine a reverberation suppression function corresponding to a target time frame in the N energy attenuation curves based on the N linear fitting curves processed by the processing module; The suppressing module is configured to suppress the reverberation part in the reverberant audio signal of the target time frame in the first audio signal based on the reverberation suppression function determined by the determining module to obtain a second audio signal; Wherein, the first audio signal is: an audio signal in the audio signal within the frequency band corresponding to the band energy attenuation curve; The band energy attenuation curve is: a band energy attenuation curve in the RIR energy attenuation curve.
7. The apparatus according to claim 6, wherein The determining module is specifically configured to calculate a reverberation weight corresponding to each linear fitting curve based on the slope of each linear fitting curve among the N linear fitting curves to obtain N reverberation weights; and calculate the early reverberation energy of the reverberant audio signal and the late reverberation energy of the reverberant audio signal based on the N reverberation weights; and determine the reverberation suppression function based on the early reverberation energy and the late reverberation energy.
8. The apparatus according to claim 7, wherein The determining module is specifically configured to, for each time frame before the target time frame, calculate the remaining energy of the direct audio signal energy of the one time frame in the first audio signal at the target time frame according to the one time frame and the reverberation weight corresponding to the one time frame to obtain the remaining energy corresponding to each time frame; and calculate the early reverberation energy and the late reverberation energy according to the remaining energy corresponding to each time frame.
9. The apparatus according to claim 7, wherein The determining module is specifically configured to calculate a priori signal-to-noise ratio corresponding to the target time frame according to the early reverberation energy, the late reverberation energy, and the energy of the ambient noise audio signal of the target time frame in the first audio signal; and calculate a posteriori signal-to-noise ratio corresponding to the target time frame according to the late reverberation energy, the energy of the ambient noise audio signal, and the amplitude spectrum of the first audio signal at the target time frame; and determine the reverberation suppression function according to the a priori signal-to-noise ratio and the a posteriori signal-to-noise ratio.
10. The apparatus according to claim 6, wherein The suppressing module is specifically configured to perform a dot product operation on the reverberation suppression function and the amplitude spectrum of the first audio signal at the target time frame to obtain a target amplitude spectrum; and perform an inverse Fourier transform on the target amplitude spectrum and the phase of the first audio signal at the target time frame to obtain the second audio signal.
11. An electronic device, characterized in that, It includes a processor and a memory. The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the audio signal processing method described in any one of claims 1-5 are implemented.
12. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the steps of the audio signal processing method described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Audio de-reverberation method and device, equipment and storage medium
CN114283827A
Audio processing for speech
US20190035415A1