A noise removal method and related apparatus

CN122575392APending Publication Date: 2026-08-14ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,对于复杂场景下的瞬时噪声,噪声抑制效果有限

Benefits of technology

[0008]本申请的有益效果是:区别于现有技术的情况,本申请提出的噪声去除方法通过根据上一历史时刻的历史噪声幅度谱预测下一目标时刻对应的初始噪声幅度谱,并结合上一历史时刻的初始增益,确定目标时刻的初始增益。利用初始语音幅度谱的目标语音活动状态,重点关注未包括语音内容的噪声部分,得到调整后的目标噪声幅度谱。最终,结合目标噪声幅度谱、初始语音幅度谱和目标时刻的初始增益,得到用于进行滤波的目标增益。通过利用目标增益对原始幅度谱进行滤波,使得得到的目标语音幅度谱准确性更高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575392A_ABST
    Figure CN122575392A_ABST
Patent Text Reader

Abstract

This application discloses a noise removal method and related apparatus. The method includes: acquiring the original amplitude spectrum corresponding to a target time in the original audio signal; predicting the initial noise amplitude spectrum corresponding to the target time based on the historical noise amplitude spectrum of the previous historical time; determining the initial gain of the target time based on the original amplitude spectrum, the initial noise amplitude spectrum, and the initial gain of the historical time; determining the matching initial speech amplitude spectrum using the initial gain of the target time; adjusting the initial noise amplitude spectrum using the target speech activity state of the initial speech amplitude spectrum and the adjusted historical noise amplitude spectrum of the previous historical time to obtain the adjusted target noise amplitude spectrum; obtaining the target gain of the target time based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain of the target time; and filtering the original amplitude spectrum using the target gain to obtain the target speech amplitude spectrum. Through the above methods, this application can improve the noise removal effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a noise removal method and related apparatus. Background Technology

[0002] In the field of noise removal, the core objective is to accurately separate and suppress noise from noisy speech while preserving or even enhancing speech quality to the maximum extent. Traditional methods typically operate in the short time-frequency domain, dynamically estimating the noise spectrum through methods such as recursive smoothing to suppress noise. However, for transient noise in complex scenarios, the noise suppression effect is limited.

[0003] Therefore, how to improve the noise removal effect has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a noise removal method and related apparatus that can improve the noise removal effect.

[0005] To address the aforementioned technical problems, this application provides a noise removal method comprising: acquiring the original amplitude spectrum corresponding to a target time in the original audio signal; predicting the initial noise amplitude spectrum corresponding to the target time based on the historical noise amplitude spectrum of the previous historical time; determining the initial gain of the target time based on the original amplitude spectrum, the initial noise amplitude spectrum, and the initial gain of the historical time; determining the initial speech amplitude spectrum matching the target time using the initial gain of the target time; adjusting the initial noise amplitude spectrum using the target speech activity state of the initial speech amplitude spectrum and the adjusted historical noise amplitude spectrum of the previous historical time to obtain the adjusted target noise amplitude spectrum; obtaining the target gain of the target time based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain of the target time; and filtering the original amplitude spectrum using the target gain to obtain the target speech amplitude spectrum.

[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method mentioned in the above technical solution.

[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium having program instructions stored thereon, wherein the program instructions, when executed by a processor, implement the method mentioned in the above technical solution.

[0008] The beneficial effects of this application are as follows: Unlike existing technologies, the noise removal method proposed in this application predicts the initial noise amplitude spectrum corresponding to the next target time based on the historical noise amplitude spectrum of the previous historical time, and determines the initial gain of the target time by combining it with the initial gain of the previous historical time. Using the target speech activity state of the initial speech amplitude spectrum, the method focuses on the noise portion that does not include speech content, resulting in an adjusted target noise amplitude spectrum. Finally, by combining the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain of the target time, the target gain used for filtering is obtained. By using the target gain to filter the original amplitude spectrum, the accuracy of the obtained target speech amplitude spectrum is improved. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating one embodiment of the noise removal method of this application; Figure 2 yes Figure 1 The flowchart of step S101 corresponds to another embodiment; Figure 3 yes Figure 2 The flowchart of step S201 corresponds to another embodiment; Figure 4 yes Figure 1 The flowchart of step S102 corresponds to another embodiment; Figure 5 yes Figure 4 The flowchart of step S404 corresponds to another embodiment; Figure 6 yes Figure 1 The flowchart of step S103 corresponds to another embodiment; Figure 7 This is a flowchart illustrating another embodiment of the noise removal method of this application; Figure 8 This is a schematic diagram of the structure of one embodiment of the electronic device of this application; Figure 9 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments, and different embodiments can be adaptively combined. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0011] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the noise removal method of this application, which includes: S101: Obtain the original amplitude spectrum corresponding to the target time in the original audio signal, and predict the initial noise amplitude spectrum corresponding to the target time based on the historical noise amplitude spectrum of the previous historical time.

[0012] In one embodiment, the original audio signal to be processed is acquired, and the original audio signal is preprocessed to determine the original amplitude spectrum corresponding to the original audio signal at the target time. Based on the historical noise amplitude spectrum corresponding to the previous historical time adjacent to the target time, the initial noise amplitude spectrum corresponding to the target time is predicted.

[0013] In some implementation scenarios, the above preprocessing involves segmenting the original audio signal into frames to determine the original audio signal segments corresponding to different times. Short-time Fourier transforms are then performed on the original audio segments at different times in chronological order to obtain the corresponding original amplitude spectra. Since the noise characteristics change relatively little in a short time, the initial noise amplitude spectrum corresponding to the target time is smoothly predicted using the historical noise amplitude spectrum from the previous historical time.

[0014] In one embodiment, the preprocessing process may further include at least one of normalization and silence removal. For example, the original audio signal is first silenced, the original audio signal after removing the silence portion is then framed, and the initial noise amplitude spectrum corresponding to the target time is predicted based on the historical noise amplitude spectrum of the previous historical time.

[0015] Additionally, it should be noted that when the original amplitude spectrum at the target time corresponds to the first frame of the original audio signal segment in the original audio signal, the historical noise amplitude spectrum at the previous historical time is the predetermined initial amplitude spectrum. The initial amplitude spectrum can be 0.

[0016] S102: Based on the original amplitude spectrum, the initial noise amplitude spectrum, and the initial gain of the historical time, determine the initial gain of the target time, and use the initial noise amplitude spectrum and the initial gain of the target time to determine the initial speech amplitude spectrum that matches the target time.

[0017] In one embodiment, the initial gain corresponding to the previous historical time is obtained, and the initial signal-to-noise ratio (SNR) for the target time is calculated based on the obtained original amplitude spectrum, initial noise amplitude spectrum, and initial gain of the historical time. The initial gain for the target time is then determined based on the initial SNR.

[0018] Furthermore, the original amplitude spectrum at the target time is filtered using the initial gain at the target time to obtain the initial speech amplitude spectrum matching the target time. The initial speech amplitude spectrum is used to characterize the amplitude spectrum corresponding to the speech content after preliminary filtering.

[0019] S103: Using the target speech activity state of the initial speech amplitude spectrum and the historical noise amplitude spectrum adjusted from the previous historical moment, the initial noise amplitude spectrum is adjusted to obtain the adjusted target noise amplitude spectrum.

[0020] In one embodiment, the target speech activity state of the initial speech amplitude spectrum is obtained. Based on the target speech activity state of the initial speech amplitude spectrum and the adjusted historical noise amplitude spectrum from the previous historical moment, the initial noise amplitude spectrum is adjusted to obtain the adjusted target noise amplitude spectrum. The target speech activity state is determined based on the energy distribution corresponding to the initial speech amplitude spectrum, and the target speech activity state is used to characterize whether speech content exists at different frequency points in the initial noise amplitude spectrum.

[0021] In some implementation scenarios, when it is determined that there is speech content at the corresponding frequency point of the target time based on the speech activity state, a larger target smoothing parameter is determined to estimate the noise at the target time based on the historical noise amplitude spectrum. When there is no speech content at the corresponding frequency point of the target time, a smaller target smoothing parameter is determined to estimate the noise based on the observation results at the target time.

[0022] S104: Based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain at the target time, the target gain at the target time is obtained. The original amplitude spectrum is then filtered using the target gain to obtain the target speech amplitude spectrum.

[0023] In one embodiment, the signal-to-noise ratio (SNR) is calculated based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain at the target time. The target gain at the target time is then obtained by filtering the original amplitude spectrum using the target gain.

[0024] The noise removal method proposed in this application predicts the initial noise amplitude spectrum corresponding to the next target time based on the historical noise amplitude spectrum of the previous historical time, and determines the initial gain of the target time by combining it with the initial gain of the previous historical time. Using the target speech activity state of the initial noise amplitude spectrum, the method focuses on the noise portion that does not include speech content, resulting in an adjusted target noise amplitude spectrum. Finally, the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain of the target time are combined to obtain the target gain used for filtering. By using the target gain to filter the original amplitude spectrum, the accuracy of the obtained target speech amplitude spectrum is improved.

[0025] In one implementation, please refer to Figure 2 , Figure 2 yes Figure 1 The flowchart of step S101 corresponds to another embodiment. Specifically, the implementation process of step S101 includes: S201: Obtain the initial speech activity state corresponding to the original amplitude spectrum, and determine the initial smoothing parameters matching the original amplitude spectrum based on the initial speech activity state. The initial speech activity state is used to characterize whether the original amplitude spectrum includes speech content.

[0026] In one embodiment, the short-time energy magnitude corresponding to each frequency point in the original amplitude spectrum is determined, and an initial speech activity state is established. Based on the initial speech activity state, an initial smoothing parameter matching the original amplitude spectrum is determined. Specifically, when the original amplitude spectrum at the target time includes speech content, the initial smoothing parameter takes a larger value; or, when the original amplitude spectrum at the target time does not include speech content, the initial smoothing parameter takes a smaller value.

[0027] In some implementation scenarios, the sum of short-time energy values ​​corresponding to all frequency points in the original amplitude spectrum is obtained, and these sums are compared with a preset energy threshold. If the sum of energy values ​​is greater than or equal to the energy threshold, it is determined that the original amplitude spectrum includes speech content, and the initial smoothing parameter is set to a first value. Alternatively, if the sum of energy values ​​is less than the energy threshold, it is determined that the original amplitude spectrum does not include speech content, and the initial smoothing parameter is set to a second value. The first value is greater than the second value.

[0028] S202: Obtain the historical noise amplitude spectrum of the previous historical moment, and predict the initial noise amplitude spectrum based on the initial smoothing parameters, the historical noise amplitude spectrum and the original amplitude spectrum.

[0029] In one embodiment, the historical noise amplitude spectrum of the previous historical time is obtained, and the corresponding original power spectrum is determined based on the original amplitude spectrum of the target time. Based on the aforementioned historical noise amplitude spectrum, initial smoothing parameters, and original power spectrum, the initial noise amplitude spectrum is predicted.

[0030] In some implementation scenarios, the specific formula for calculating the initial noise amplitude spectrum is as follows:

[0031] in, Indicates the target time The initial noise amplitude spectrum, Indicates the initial smoothing parameters. This represents the historical noise amplitude spectrum at the previous historical moment. Represents the original amplitude spectrum. This represents the original power spectrum corresponding to the original amplitude spectrum.

[0032] The above scheme uses the initial speech activity state to obtain the initial smoothing parameter, and uses the initial smoothing parameter to predict the initial noise amplitude spectrum at the target time. This allows the initial noise amplitude spectrum to be predicted by combining the historical noise amplitude spectrum of the previous historical time when there is no speech, so as to avoid treating the speech content as noise.

[0033] Please see Figure 3 , Figure 3 yes Figure 2 The flowchart of step S201 corresponds to another embodiment. Specifically, the implementation process of step S201 includes: S301: Obtain the local frequency range corresponding to the current frequency in the original amplitude spectrum, and determine the reference power spectrum corresponding to the local frequency range.

[0034] In one implementation, all frequencies corresponding to the original amplitude spectrum at the target time are traversed, and the current frequency is frequency-extended to obtain a local frequency range with a fixed bandwidth. A reference power spectrum corresponding to this local frequency range is then determined from the original amplitude spectrum.

[0035] In some implementation scenarios, a preset reference frequency range is obtained, which corresponds to a minimum reference frequency and a maximum reference frequency. The difference between the current frequency and each frequency in the aforementioned reference frequency range is calculated, and the local frequency range is determined based on the difference calculation result. For example, the reference frequency range is determined by the minimum reference frequency. and maximum reference frequency Determined, the current frequency The difference between the original frequency spectrum and the reference frequency range is calculated for each frequency within the range. Based on the difference calculation results, the corresponding local frequency range is obtained, and the reference power spectrum corresponding to the local frequency range is determined from the original amplitude spectrum. Determining the local frequency range allows for processing of the original amplitude spectrum within that range, thereby improving the efficiency of determining the initial smoothing parameters.

[0036] S302: Obtain the historical local frequency band energy corresponding to the previous historical moment, and determine the reference local frequency band energy corresponding to the current frequency at the target moment based on the reference power spectrum and the historical local frequency band energy.

[0037] In one embodiment, the historical local frequency band energy and a first preset weight corresponding to the previous historical moment are obtained. Based on the preset weight, the reference power spectrum, and the historical local frequency band energy, the reference local frequency band energy corresponding to the current frequency at the target moment is determined. The specific calculation formula for the reference local frequency band energy is as follows:

[0038] in, Indicates the reference local frequency band energy. Indicates the first preset weight. Indicates the use of length is Window statistical energy comprehensive value, This represents the local frequency band energy corresponding to the previous historical moment.

[0039] S303: Obtain the reference energy threshold based on the reference local frequency band energy and the historical local frequency band energy corresponding to multiple adjacent historical times before the target time.

[0040] In one embodiment, after obtaining the reference local frequency band energy corresponding to the current target time, the reference energy threshold is obtained by filtering from the historical local frequency band energies corresponding to the target time and multiple adjacent historical times before the target time.

[0041] In some implementation scenarios, multiple historical moments preceding and adjacent to the target moment are used as reference historical moments. The minimum value of the local frequency band energy corresponding to the target moment and each reference historical moment is used as the reference energy threshold. The specific selection formula for the reference energy threshold is as follows:

[0042] Among them, the above The specific value can be set according to the actual scenario.

[0043] In some implementation scenarios, after determining the minimum value of the historical local frequency band energy corresponding to the target time and each reference historical time, a second preset weight is obtained. The minimum value is then multiplied by the second preset weight to obtain the reference energy threshold. The specific value of the second preset weight can be set according to the actual scenario.

[0044] In one embodiment, after obtaining the target local energy corresponding to the current target time, multiple historical times before the target time and adjacent to the target time are used as reference historical times. The average value of the historical local frequency band energy corresponding to the target time and each reference historical time is obtained. The average value is multiplied by a second preset weight to obtain a reference energy threshold.

[0045] S304: Based on the comparison between the reference local frequency band energy corresponding to the current frequency and the reference energy threshold, determine the initial speech activity state corresponding to the original amplitude spectrum.

[0046] In one embodiment, the reference local frequency band energy corresponding to the current frequency at the target time is compared with a reference energy threshold to determine the initial speech activity state corresponding to the original amplitude spectrum.

[0047] In some implementation scenarios, when the reference local frequency band energy corresponding to the current frequency at the target time is greater than or equal to the reference energy threshold, it indicates that speech content exists at the current frequency in the original amplitude spectrum. Alternatively, when the reference local frequency band energy corresponding to the current frequency at the target time is less than the reference energy threshold, it indicates that no speech content exists at the current frequency in the original amplitude spectrum.

[0048] S305: In response to the presence of speech content at the current frequency in the original amplitude spectrum, the initial smoothing parameter is determined to be the first value.

[0049] In one implementation, when the current frequency in the original amplitude spectrum corresponds to speech content, the initial smoothing parameter is determined to be a first value.

[0050] S306: In response to the fact that the current frequency in the original amplitude spectrum does not include speech content, the initial smoothing parameter is determined to be a second value; wherein the first value is greater than the second value.

[0051] In one embodiment, when there is no speech content at the current frequency in the original amplitude spectrum, the initial smoothing parameter is determined to be a second value, and the first value is greater than the second value.

[0052] Please see Figure 4 , Figure 4 yes Figure 1 The flowchart of step S102 corresponds to another embodiment. Specifically, the implementation process of step S102 includes: S401: Obtain the predicted speech power spectrum based on the original amplitude spectrum and the initial noise amplitude spectrum.

[0053] In one embodiment, the original power spectrum corresponding to the original amplitude spectrum is obtained, and the predicted speech power spectrum is obtained based on the original power spectrum and the initial noise amplitude spectrum. The specific calculation formula for the predicted speech power spectrum is as follows:

[0054] in, Indicates the predicted speech power spectrum. Represents the original amplitude spectrum The corresponding original power spectrum, This represents the initial noise amplitude spectrum.

[0055] S402: Based on the initial gain, historical noise amplitude spectrum, and the ratio between the predicted speech power spectrum and the initial noise amplitude spectrum of the previous historical moment, obtain the initial signal-to-noise ratio corresponding to the target moment.

[0056] In one embodiment, a first ratio between the predicted speech power spectrum and the initial noise amplitude spectrum is obtained. A first product between the initial gain at the previous historical time and the historical noise power spectrum at the previous historical time is obtained, and a second ratio between the first product and the historical noise amplitude spectrum at the previous historical time is determined. Based on preset parameters, the first ratio, and the second ratio, the initial signal-to-noise ratio (SNR) corresponding to the target time is obtained. The specific formula for calculating the initial SNR is as follows:

[0057]

[0058] in, This represents the initial signal-to-noise ratio. This represents the initial gain at the previous historical moment. This represents the historical noise power spectrum at the previous historical moment. This represents the historical noise amplitude spectrum at the previous historical moment. Indicates the predicted speech power spectrum. This represents the initial noise amplitude spectrum. And, and This is a preset constant, and its specific value can be obtained through estimation or through multiple experiments. Additionally, it should be noted that when the original amplitude spectrum at the target time corresponds to the first frame of the original audio signal segment, the initial gain at the previous historical time is the predetermined initial estimated gain, and the specific value of the initial estimated gain can be set according to the actual situation.

[0059] S403: Based on the initial signal-to-noise ratio, obtain the initial gain corresponding to the target time.

[0060] In one embodiment, after determining the initial signal-to-noise ratio (SNR), the initial gain corresponding to the target time is calculated using the initial SNR. Wherein, the initial gain corresponding to the target time... The specific calculation formula is as follows:

[0061] S404: Determine the initial speech amplitude spectrum matched at the target time using the initial gain corresponding to the target time.

[0062] In one embodiment, the original amplitude spectrum is filtered using the initial gain corresponding to the target time to obtain the initial speech amplitude spectrum matching the target time. The initial speech amplitude spectrum... The specific calculation formula is as follows:

[0063] The above scheme combines the initial gain of the previous historical moment and the historical noise amplitude spectrum to determine the initial gain corresponding to the target moment. Then, the original amplitude spectrum is filtered using the initial gain corresponding to the target moment, so as to efficiently predict the initial speech amplitude spectrum that matches the target moment.

[0064] Please see Figure 5 , Figure 5 yes Figure 4 The flowchart of step S404 corresponds to another embodiment. Specifically, the specific implementation process of step S404 includes: S501: Based on the initial gain and original amplitude spectrum corresponding to the target time, obtain the initial speech amplitude spectrum that matches the target time.

[0065] In one embodiment, the original amplitude spectrum is filtered using the initial gain corresponding to the target time to obtain an initial amplitude spectrum that matches the target time. The specific calculation formula for the initial amplitude spectrum can be found in the corresponding embodiments described above, and will not be elaborated upon here.

[0066] S502: Obtain multiple candidate fundamental frequencies, and based on each candidate fundamental frequency and the sampling rate of the original audio signal, obtain the harmonic frequencies corresponding to the initial speech amplitude spectrum.

[0067] In one embodiment, multiple preset candidate fundamental frequencies are acquired, and peak detection is performed on the initial speech amplitude spectrum based on each candidate fundamental frequency and the sampling rate of the original audio signal to determine the harmonic frequencies corresponding to the initial speech amplitude spectrum.

[0068] In some implementation scenarios, peak detection is performed on the initial speech amplitude spectrum using the candidate fundamental frequency and the sampling rate of the original audio signal to determine the feature factors corresponding to the initial speech amplitude spectrum. The specific calculation formulas for these feature factors are as follows:

[0069] in, Indicates characteristic factor; This represents the candidate fundamental frequency, and its specific value can be 30, 30.1, ... 500, etc. This represents the sampling rate of the original audio signal, which is a preset parameter greater than or equal to 0 and less than or equal to 1. Indicates to and The ratio is rounded down. This represents the initial speech amplitude spectrum.

[0070] Furthermore, the harmonic frequencies of the initial speech amplitude spectrum are determined based on the frequencies at which the feature factors reach their maximum values. Wherein, the harmonic frequencies... The specific calculation formula is as follows:

[0071]

[0072] S503: Obtain the preset attenuation weight, and use the attenuation weight to enhance the harmonic frequencies in the initial speech amplitude spectrum to obtain the adjusted initial speech amplitude spectrum.

[0073] In one embodiment, a preset attenuation weight is obtained, and the harmonic frequencies in the initial speech amplitude spectrum are enhanced using the attenuation weight to obtain the harmonic-enhanced initial speech amplitude spectrum.

[0074] In some implementation scenarios, for all detected harmonics, the initial speech amplitude spectrum at the corresponding frequency is enhanced using attenuation weights until the initial speech amplitude spectrum after harmonic enhancement is obtained.

[0075] In some implementation scenarios, the specific calculation formula for harmonic enhancement using attenuation weights is as follows:

[0076] in, This represents the initial speech amplitude spectrum after harmonic enhancement. This represents the attenuation weight. Preferably, the attenuation weight is a fixed value less than 1.

[0077] The above scheme performs harmonic detection on the initial speech amplitude spectrum and uses preset attenuation weights to enhance the harmonics of the initial speech amplitude spectrum, thereby repairing and strengthening the naturalness of the speech that has been damaged or blurred due to noise reduction processing and improving the speech quality.

[0078] Please see Figure 6 , Figure 6 yes Figure 1 The flowchart of step S103 corresponds to another embodiment. Specifically, the implementation process of step S103 includes: S601: Obtain the target speech activity state of the initial speech amplitude spectrum, and determine the target smoothing parameter that matches the initial noise amplitude spectrum based on the target speech activity state.

[0079] In one embodiment, after acquiring the initial speech amplitude spectrum, the target speech activity state of the initial speech amplitude spectrum is determined. Based on the target speech activity state, a target smoothing parameter matching the initial noise amplitude spectrum is determined.

[0080] In some implementation scenarios, the local frequency range corresponding to the current frequency in the initial speech amplitude spectrum after harmonic enhancement is obtained, and the target power spectrum corresponding to this local frequency range is determined. The historical target local frequency band energy corresponding to the previous historical moment is obtained. Based on the target power spectrum and the historical target local frequency band energy, the target local frequency band energy corresponding to the current frequency at the target moment is determined. The specific calculation formula for the target local frequency band energy is as follows:

[0081] in, This represents the target's local frequency band energy. This represents the local frequency band energy of the historical target corresponding to the previous historical moment. This represents the initial speech power spectrum corresponding to the initial speech amplitude spectrum.

[0082] Furthermore, multiple historical moments preceding and adjacent to the target moment are used as reference historical moments. The minimum target local frequency band energy corresponding to the target moment and each reference historical moment is multiplied by a preset weight to obtain the target energy threshold. When the target local frequency band energy corresponding to the current frequency at the target moment is greater than or equal to the target energy threshold, it indicates that there is speech content at the current frequency in the initial speech amplitude spectrum, and the target smoothing parameter is determined to be the third value. Alternatively, when the target local frequency band energy corresponding to the current frequency at the target moment is less than the target energy threshold, it indicates that there is no speech content at the current frequency in the initial speech amplitude spectrum, and the target smoothing parameter is determined to be the fourth value. The third value is greater than the fourth value.

[0083] Optionally, the third value can be the same as the first value in the corresponding embodiments described above, and the fourth value can be the same as the second value in the corresponding embodiments described above. Furthermore, the specific process for determining the target smoothing parameter in this embodiment can refer to the process for determining the initial smoothing parameter in the corresponding embodiments described above.

[0084] S602: Based on the target smoothing parameters, the initial noise amplitude spectrum, and the historical noise amplitude spectrum adjusted at the previous historical moment, obtain the adjusted target noise amplitude spectrum.

[0085] In one embodiment, the target noise amplitude spectrum adjusted for the target time is determined based on the target smoothing parameter, the initial noise amplitude spectrum at the target time, and the historical noise amplitude spectrum adjusted at the previous historical time.

[0086] In some implementation scenarios, the specific calculation formula for the above target noise amplitude spectrum is as follows:

[0087] in, This represents the target noise amplitude spectrum after adjustment at the target time. Indicates the target smoothing parameter. This represents the historical noise amplitude spectrum after adjustment at the previous historical moment.

[0088] The above scheme, by determining the target speech activity state of the initial speech amplitude spectrum and updating the initial noise amplitude spectrum, makes the adjusted target noise amplitude spectrum more accurate.

[0089] Please see Figure 7 , Figure 7 This is a flowchart illustrating another embodiment of the noise removal method of this application. Specifically, the specific process for obtaining the target speech amplitude spectrum mentioned in any of the above embodiments may further include: S701: Obtain the initial speech power spectrum corresponding to the initial speech amplitude spectrum and the historical target gain of the previous historical moment.

[0090] In one implementation, the initial speech power spectrum is obtained based on the initial speech amplitude spectrum. Additionally, the historical target gain from the previous iteration is obtained.

[0091] In some implementation scenarios, the initial speech amplitude spectrum in this embodiment can be obtained by directly filtering the original amplitude spectrum using the initial gain corresponding to the target time. Alternatively, the initial speech amplitude spectrum can also be obtained by harmonic enhancement using preset attenuation weights, wherein the specific process of harmonic enhancement can be operated in steps S501 to S503.

[0092] S702: Based on historical target gain, target noise amplitude spectrum and initial speech power spectrum, obtain the target signal-to-noise ratio at the target time, and use the target signal-to-noise ratio to determine the reference gain at the target time.

[0093] In one embodiment, the target signal-to-noise ratio (SNR) at the target time is calculated based on the historical target gain, the target noise amplitude spectrum, and the initial speech power spectrum. The reference gain at the target time is then calculated using the target SNR.

[0094] In some implementation scenarios, the target ratio between the initial speech power spectrum and the target noise amplitude spectrum at the target time is obtained. Based on the historical target gain and the target ratio, the target signal-to-noise ratio at the target time is determined. The specific formula for calculating the target signal-to-noise ratio is as follows:

[0095]

[0096] in, The target signal-to-noise ratio at a given moment. Characterizing historical target gain, The adjusted noise amplitude spectrum representing the previous historical moment. Characterizing the initial speech power spectrum, Characterizes the target noise amplitude spectrum. Among them, The specific acquisition process can be found in the process of acquiring the target noise amplitude spectrum at the target time.

[0097] Furthermore, a preset lower gain limit is obtained. Using the lower gain limit and the target signal-to-noise ratio (SNR) at the target time, a reference gain at the target time is determined. This ensures that the reference gain fluctuates with the target SNR, avoiding the impact of a fixed gain on the continuity of the filtering results. The specific formula for calculating the reference gain is as follows:

[0098] in, The reference gain characterizing the target time; The lower limit of the gain is preset, and its specific value can be set according to the actual scenario. In this embodiment, it is set to a value between 0.01 and 0.05.

[0099] S703: Based on the initial gain and reference gain at the target time, determine the target gain at the target time, and use the target gain to filter the original amplitude spectrum to obtain the target speech amplitude spectrum.

[0100] In one implementation, for each frequency corresponding to the target time, the larger of the initial gain and the reference gain is used as the corresponding target gain. The original amplitude spectrum is filtered using the target gain to obtain the target speech amplitude spectrum corresponding to the speech content.

[0101] In some implementation scenarios, the target gain corresponding to different frequencies at the target time is obtained by filtering using the initial gain and the reference gain. The specific filtering formula is as follows:

[0102] in, This represents the target gain corresponding to the frequency at the target time. Indicates the initial gain. Indicates the reference gain.

[0103] The above scheme determines the reference gain based on the obtained target signal-to-noise ratio, and selects the larger value from the initial gain and the reference gain as the target gain for the corresponding frequency, which helps to improve the effect of subsequent filtering.

[0104] In one embodiment, to improve noise removal efficiency, after obtaining the reference gain through the above-described corresponding embodiments, the original amplitude spectrum is directly filtered using the reference gain to obtain the target speech amplitude spectrum.

[0105] Please see Figure 8 , Figure 8 This is a schematic diagram of one embodiment of the electronic device of this application. The electronic device includes a memory 10 and a processor 20 coupled to each other. The memory 10 stores program instructions, and the processor 20 executes the program instructions to implement the methods mentioned in any of the above embodiments. Specifically, the electronic device includes, but is not limited to, desktop computers, laptops, tablets, servers, etc., and is not limited thereto. In addition, the processor 20 may also be called a CPU (Center Processing Unit). The processor 20 may be an integrated circuit chip with signal processing capabilities. The processor 20 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. In addition, the processor 20 may be implemented by integrated circuit chips.

[0106] Please see Figure 9 , Figure 9 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 30 stores program instructions 40 that can be executed by a processor. When the program instructions 40 are executed by the processor, they implement the methods mentioned in any of the above embodiments.

[0107] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0108] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0109] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0110] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0111] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A noise removal method, characterized in that, include: Obtain the original amplitude spectrum corresponding to the target time in the original audio signal, and predict the initial noise amplitude spectrum corresponding to the target time based on the historical noise amplitude spectrum of the previous historical time. Based on the original amplitude spectrum, the initial noise amplitude spectrum, and the initial gain of the historical time, the initial gain of the target time is determined, and the initial speech amplitude spectrum matching the target time is determined using the initial gain of the target time. The initial noise amplitude spectrum is adjusted using the target speech activity state of the initial speech amplitude spectrum and the historical noise amplitude spectrum adjusted at the previous historical moment to obtain the adjusted target noise amplitude spectrum. Based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain at the target time, the target gain at the target time is obtained. The original amplitude spectrum is then filtered using the target gain to obtain the target speech amplitude spectrum.

2. The noise removal method according to claim 1, characterized in that, The prediction of the initial noise amplitude spectrum corresponding to the target time based on the historical noise amplitude spectrum of the previous historical time includes: Obtain the initial speech activity state corresponding to the original amplitude spectrum, and determine the initial smoothing parameter matching the original amplitude spectrum based on the initial speech activity state; wherein, the initial speech activity state is used to characterize whether the original amplitude spectrum includes speech content; The historical noise amplitude spectrum of the previous historical moment is obtained, and the initial noise amplitude spectrum is predicted based on the initial smoothing parameter, the historical noise amplitude spectrum and the original amplitude spectrum.

3. The noise removal method according to claim 2, characterized in that, The step of obtaining the initial speech activity state corresponding to the original amplitude spectrum and determining the initial smoothing parameters matching the original amplitude spectrum based on the initial speech activity state includes: Obtain the local frequency range corresponding to the current frequency in the original amplitude spectrum, and determine the reference power spectrum corresponding to the local frequency range; Obtain the historical local frequency band energy corresponding to the previous historical moment, and determine the reference local frequency band energy corresponding to the current frequency at the target moment based on the reference power spectrum and the historical local frequency band energy; Based on the reference local frequency band energy and the historical local frequency band energy corresponding to multiple adjacent historical moments before the target time, a reference energy threshold is obtained. Based on the comparison between the reference local frequency band energy corresponding to the current frequency and the reference energy threshold, the initial speech activity state corresponding to the original amplitude spectrum is determined; In response to the fact that the current frequency in the original amplitude spectrum corresponds to speech content, the initial smoothing parameter is determined to be a first value; In response to the fact that the current frequency in the original amplitude spectrum does not include speech content, the initial smoothing parameter is determined to be a second value; wherein the first value is greater than the second value.

4. The noise removal method according to claim 1, characterized in that, The step of determining the initial gain of the target time based on the original amplitude spectrum, the initial noise amplitude spectrum, and the initial gain of the historical time, and determining the initial speech amplitude spectrum matching the target time using the initial gain of the target time, includes: Based on the original amplitude spectrum and the initial noise amplitude spectrum, the predicted speech power spectrum is obtained; Based on the initial gain at the historical moment, the historical noise amplitude spectrum, and the ratio between the predicted speech power spectrum and the initial noise amplitude spectrum, the initial signal-to-noise ratio corresponding to the target moment is obtained; Based on the initial signal-to-noise ratio, obtain the initial gain corresponding to the target time. Using the initial gain corresponding to the target time, the initial speech amplitude spectrum matching the target time is determined.

5. The noise removal method according to claim 4, characterized in that, The step of determining the initial speech amplitude spectrum matching the target time using the initial gain corresponding to the target time includes: Based on the initial gain corresponding to the target time and the original amplitude spectrum, obtain the initial speech amplitude spectrum that matches the target time; Multiple candidate fundamental frequencies are obtained, and based on each candidate fundamental frequency and the sampling rate of the original audio signal, the harmonic frequencies corresponding to the initial speech amplitude spectrum are obtained; A preset attenuation weight is obtained, and the harmonic frequencies in the initial speech amplitude spectrum are enhanced using the attenuation weight to obtain the adjusted initial speech amplitude spectrum.

6. The noise removal method according to claim 1, characterized in that, The step of adjusting the initial noise amplitude spectrum using the target speech activity state of the initial speech amplitude spectrum and the adjusted historical noise amplitude spectrum from the previous historical moment to obtain the adjusted target noise amplitude spectrum includes: Obtain the target speech activity state of the initial speech amplitude spectrum, and determine the target smoothing parameter that matches the initial noise amplitude spectrum based on the target speech activity state; Based on the target smoothing parameters, the initial noise amplitude spectrum, and the adjusted historical noise amplitude spectrum from the previous historical moment, the adjusted target noise amplitude spectrum is obtained.

7. The noise removal method according to claim 1 or 4, characterized in that, The process of obtaining the target gain at the target time based on the target noise amplitude spectrum, the initial speech amplitude spectrum, and the initial gain at the target time, and then filtering the original amplitude spectrum using the target gain to obtain the target speech amplitude spectrum, includes: Obtain the initial speech power spectrum corresponding to the initial speech amplitude spectrum and the historical target gain of the previous historical moment; Based on the historical target gain, the target noise amplitude spectrum, and the initial speech power spectrum, the target signal-to-noise ratio at the target time is obtained, and the reference gain at the target time is determined using the target signal-to-noise ratio; Based on the initial gain and reference gain at the target time, the target gain at the target time is determined, and the original amplitude spectrum is filtered using the target gain to obtain the target speech amplitude spectrum.

8. The noise removal method according to claim 7, characterized in that, Determining the target gain at the target time based on the initial gain and reference gain at the target time includes: For each frequency corresponding to the target time, the larger value between the initial gain and the reference gain is taken as the corresponding target gain.

9. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, the memory storing program instructions, and the processor executing the program instructions to implement the method as described in any one of claims 1-8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method as described in any one of claims 1-8.