Noise reduction methods, devices, computer equipment, and storage media with dynamically controllable noise reduction range
By calculating the short-time average energy and noise energy of the frequency domain signal, and dynamically adjusting the noise reduction range using an adjustable parameter gain function, the problem of narrow gain range in existing technologies is solved, achieving a balance between flexible noise reduction control and speech preservation.
Patent Information
- Application Number
- CN202210306321.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Existing gain algorithms have a narrow gain range in audio noise reduction processing, making it difficult to flexibly adjust the noise reduction range and thus making it difficult to balance noise reduction intensity and speech preservation.
By calculating the short-time average energy of the frequency domain signal and the short-time average energy of the input speech signal, the noise reduction range is dynamically controlled using an adjustable gain function, including α, β, and γ parameters, to adjust the noise reduction intensity and the degree of speech preservation.
It enables dynamic control of the noise reduction range, improves the flexibility of noise reduction effect and speech preservation quality, and enhances the adaptability of the noise reduction method.
Smart Images

Figure CN114974196B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio noise reduction technology, and in particular to a noise reduction method, apparatus, computer device, and storage medium that can dynamically control the noise reduction range. Background Technology
[0002] An audio clip typically includes noise and speech components. For example, in a narration, the noise component is the background sound, including other people's voices, wind noise, engine noise, and electronic interference. Therefore, audio processing is needed to remove noise, commonly known as noise reduction. When performing noise reduction, a balance needs to be struck between noise reduction intensity and speech preservation using a gain function. This prevents excessive noise reduction from resulting in low speech preservation, or insufficient noise reduction from producing an ineffective result. Common gain algorithms include Wiener filtering and spectral reduction functions; however, these algorithms have narrow gain ranges and cannot flexibly adjust the noise reduction range. Summary of the Invention
[0003] This invention provides a noise reduction method, apparatus, computer device, and storage medium that can dynamically control the noise reduction range, which can not only remove noise from speech but also dynamically control the noise reduction range.
[0004] In a first aspect, embodiments of the present invention provide a noise reduction method with dynamically controllable noise reduction range, the method comprising:
[0005] Obtain the frequency domain signal X(k,m) of the input speech signal, and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index;
[0006] According to the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy
[0007] The short-time speech energy and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range;
[0008] The frequency domain signal X(k,m) is amplified using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal.
[0009] Secondly, embodiments of the present invention also provide a noise reduction device capable of dynamically controlling the noise reduction range, the device comprising:
[0010] The first calculation unit is used to acquire the frequency domain signal X(k,m) of the input speech signal and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index;
[0011] Short-time speech energy calculation unit, used to calculate based on the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy
[0012] The speech gain calculation unit is used to calculate the short-time speech energy. and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range;
[0013] The output speech signal calculation unit is used to amplify the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal.
[0014] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.
[0016] This invention provides a noise reduction method, apparatus, computer device, and storage medium with dynamically controllable noise reduction range. The method includes: acquiring a frequency domain signal X(k,m) of an input speech signal, and calculating the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index, and m is the short-time Fourier transform time index; according to the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy The short-time speech energy and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range; the frequency domain signal X(k,m) is amplified using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal. This embodiment of the invention can calculate the short-time average energy of the frequency domain signal X(k,m) of the input speech signal. and short-time noise average energy This allows for the acquisition of short-term speech energy. Then average the energy of the short-time noise. and short-term speech energy Substituting the gain function W provided in the embodiment of the present invention, the speech gain W(ω) is obtained. The obtained speech gain W(ω) includes three adjustable parameters: α, β and γ. The noise reduction intensity, the degree of speech preservation and the ratio between the noise reduction intensity and the degree of speech preservation can be adjusted by adjusting the values of α, β and γ, thereby realizing dynamic adjustment of the noise reduction range. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the noise reduction method with dynamically controllable noise reduction range provided in an embodiment of the present invention.
[0019] Figure 2 and Figure 3 This is a graph of the input signal-to-noise ratio versus the output gain provided in an embodiment of the present invention;
[0020] Figure 4 This is a flowchart illustrating a noise reduction method with dynamically controllable noise reduction range provided in another embodiment of the present invention.
[0021] Figure 5 This is a schematic block diagram of a noise reduction device with dynamically controllable noise reduction range provided in an embodiment of the present invention;
[0022] Figure 6 This is a schematic block diagram of a noise reduction device with dynamically controllable noise reduction range provided in another embodiment of the present invention;
[0023] Figure 7This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] It should be understood that, when used in this specification and the appended claims, the terms “comprising” and “including” indicate the presence of the described features, integrals, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0027] Please see Figure 1 , Figure 1 This is a flowchart illustrating a noise reduction method with dynamically controllable noise reduction range provided in an embodiment of the present invention. Figure 1 As shown, the method includes steps S110 to S140.
[0028] S110, acquire the frequency domain signal X(k,m) of the input speech signal, and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index.
[0029] In this embodiment of the invention, the frequency domain signal describes the relationship between frequency and amplitude, typically represented on a coordinate system with frequency on the horizontal axis and amplitude on the vertical axis. Frequency domain signals and time domain signals are mutually convertible, and conversion methods include, but are not limited to, Fourier transform, discrete transform, and improved discrete cosine transform. The frequency domain signal X(k,m) of the input speech signal can be obtained from the time domain signal through Fourier transform. After obtaining the frequency domain signal X(k,m), the short-time average energy in the frequency domain signal X(k,m) can be calculated. and short-time noise average energy Speech input signals are generally non-stationary random processes that vary over time. While speech is time-varying, it exhibits short-time correlation. This correlation stems from the inertia of the human vocal organs, meaning the state of speech does not change abruptly. Therefore, when calculating energy, we typically calculate the short-time energy (short-time average energy) and the short-time noise average energy. The frequency domain signal X(k,m) generally includes two indices: k, which is the discrete spectrum index, and m, which is the short-time Fourier transform time index.
[0030] In some embodiments, such as this embodiment, the step of acquiring the frequency domain signal X(k,m) of the input speech signal may include the following steps: acquiring the time domain signal x(n) of the input speech signal, where n is a discrete-time index; and converting the time domain signal into the frequency domain signal X(k,m) through a short-time Fourier transform.
[0031] In this embodiment of the invention, the speech signal that can be directly obtained is generally a time-domain signal. Therefore, the time-domain signal x(n) of the input speech signal can be obtained first, and then the time-domain signal x(n) can be converted into the frequency-domain signal X(k,m) through short-time Fourier transform.
[0032] In some embodiments, such as this embodiment, the noise reduction method with dynamically controllable noise reduction range may further include the following steps: calculating the zero-crossing rate r of each short-time Fourier transform time index m in the frequency domain signal X(k,m), and determining whether there is a short-time Fourier transform time index m with a zero-crossing rate r greater than a preset value a; if there is a short-time Fourier transform time index m with a zero-crossing rate r greater than the preset value a, then counting the number N of all short-time Fourier transform time indices m with a zero-crossing rate r greater than the preset value a.
[0033] In this embodiment of the invention, the zero-crossing rate is a characteristic parameter in the time-domain analysis of speech signals. It refers to the number of times the signal crosses zero within each frame. In the case of discrete-time speech signals, if adjacent samples have different algebraic signs, it is said that a zero-crossing has occurred, and therefore the number of zero-crossings can be calculated. The number of zero-crossings per unit time is called the zero-crossing rate. Therefore, the zero-crossing rate can reflect the frequency information of the signal to a certain extent. The preset value 'a' can be set according to specific circumstances. When the zero-crossing rate 'r' is greater than the preset value 'a', it indicates that the corresponding short-time Fourier transform time index 'm' has a high number of zero-crossings, indicating that the signal segment corresponding to the short-time Fourier transform time index 'm' contains only noise, and can be used to statistically analyze the average energy of short-time noise. Assuming that there are 5 short-time Fourier transform time indices 'm' in the frequency domain signal X(k,m) with zero-crossing rates 'r' greater than 'a', and they are m1, m2, m3, m4, and m5 respectively, then N = 5, I = {m1, m2, m3, m4, and m5}.
[0034] In some embodiments, such as this one, the calculation of the short-time average energy of the frequency domain signal X(k,m) is... and short-time noise average energy The steps include the following: According to the formula Calculate the average energy of the short-time noise. Where I is the set of all short-time Fourier transform time indices m for which the zero-crossing rate r is greater than the preset value a; according to the formula Calculate the short-time average energy
[0035] In this embodiment of the invention, given that the number of short-time Fourier transform time indices m with a zero-crossing rate r greater than a preset value is N, and the set is I, the average energy of short-time noise is... Short-time average energy
[0036] S120, based on the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy
[0037] In this embodiment of the invention, after calculating the short-time average energy... and short-time noise average energy Then, short-time average energy can be used. Subtract the average energy of short-time noise Obtain short-time speech energy
[0038] S130, the short-time speech energy and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range.
[0039] In this embodiment of the invention, the gain function Then the gain function
[0040]
[0041] Average energy of short-time noise and short-term speech energy Substituting into formula (1) yields the speech gain W(ω). Here, α, β, and γ are all adjustable parameters, α≥0, β≥0, γ≥0. The noise reduction intensity and speech preservation level can be adjusted by changing the values of these three parameters. A larger α, a larger β, and a larger γ result in greater noise reduction. For example... Figure 2 As shown, Figure 2 This is a graph of the input signal-to-noise ratio (SNR) versus the output gain. Different formulas correspond to different curves, and each curve represents a different degree of noise reduction under different noise intensities. A lower SNR indicates greater noise, and a smaller gain indicates greater noise reduction, but there may be significant speech loss. It can be seen that different algorithms tend to increase gain as the SNR decreases. The formula for Wiener filtering is... The formula for Wiener filtering with adjustable speech loss (SDW-SWF) is: The formula corresponding to the spectral subtraction is: The formula corresponding to the power subtraction spectrum is: like Figure 3 As shown, Figure 3 Even when using the same input signal-to-noise ratio (SNR) versus output gain curve, different formulas will produce different curves.
[0042] The formula for Wiener filtering is: The formula corresponding to the spectral subtraction is: parametric2 is the gain function W provided in the embodiment of the present invention. It can be seen that the gain function W provided in the embodiment of the present invention also provides slope control of the curve, which can control the slope of the curve to adjust the degree of inclination of the curve, and can increase or decrease the dynamic range of noise reduction.
[0043] S140, the frequency domain signal X(k,m) is amplified using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal.
[0044] In this embodiment of the invention, after obtaining the speech gain W(ω), the speech gain W(ω) is multiplied by the frequency domain signal X(k,m) to complete the gain of the frequency domain signal X(k,m), thereby obtaining the frequency domain signal Y(k,m) of the output speech signal.
[0045] In some embodiments, such as this embodiment, step S140 may include the following step: calculating the short-time average gain of the speech gain W(ω). in, Using the short-time average gain The frequency domain signal X(k,m) is amplified to obtain the frequency domain signal Y(k,m).
[0046] In this embodiment of the invention, before applying the voice gain W(ω) to the frequency domain signal X(k,m), the short-time average gain of the voice gain W(ω) can be calculated first. Then use short-time average gain Gaining the frequency domain signal X(k,m) can also be done directly using the speech gain W(ω).
[0047] In some embodiments, such as this one, after step S140, the following step may be included: converting the frequency domain signal Y(k,m) into a time domain signal y(n) by inverse short-time Fourier transform to obtain the time domain signal of the output speech signal.
[0048] In this embodiment of the invention, after obtaining the frequency domain signal Y(k,m), the frequency domain signal Y(k,m) can be converted into the time domain signal y(n) by the inverse short-time Fourier transform, and then the time domain signal y(n) is output to complete the processing of the speech signal.
[0049] Figure 4 This is another embodiment of the noise reduction method that allows for dynamic control of the noise reduction range, such as... Figure 4 As shown, the noise reduction method with dynamically controllable noise reduction range in this embodiment includes steps S210-S250. Steps S210-S240 are similar to steps S110-S140 in the above embodiment and will not be described again here. The following details the additional step S250 in this embodiment.
[0050] S250, the frequency domain signal Y(k,m) is converted into a time domain signal y(n) by improving the discrete cosine transform to obtain the time domain signal of the output speech signal.
[0051] In this embodiment of the invention, the frequency domain signal Y(k,m) can be converted into the time domain signal y(n) by improving the discrete cosine transform, and then the time domain signal y(n) can be output to complete the processing of the speech signal.
[0052] Figure 5 This is a schematic block diagram of a noise reduction device 100 with dynamically controllable noise reduction range provided in an embodiment of the present invention. Figure 5 As shown, corresponding to the above-described noise reduction method with dynamically controllable noise reduction range, the present invention also provides a noise reduction device 100 with dynamically controllable noise reduction range. This noise reduction device 100 includes a unit for performing the above-described noise reduction method with dynamically controllable noise reduction range. Specifically, please refer to... Figure 5The noise reduction device 100, which can dynamically control the noise reduction range, includes a first calculation unit 110, a short-time speech energy calculation unit 120, a speech gain calculation unit 130, and an output speech signal calculation unit 140.
[0053] The first calculation unit 110 is used to acquire the frequency domain signal X(k,m) of the input speech signal and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index, and m is the short-time Fourier transform time index; the short-time speech energy calculation unit 120 is used to calculate the short-time average energy. and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy The speech gain calculation unit 130 is used to calculate the short-time speech energy. and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β and γ are all adjustable parameters used to control the noise reduction range; the output speech signal calculation unit 140 is used to amplify the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal.
[0054] An embodiment of the present invention also provides a noise reduction device with dynamically controllable noise reduction range. It adds a first judgment unit and a first statistical unit to the above embodiment.
[0055] The first judgment unit is used to calculate the zero-crossing rate r of each short-time Fourier transform time index m in the frequency domain signal X(k,m), and determine whether there is a short-time Fourier transform time index m whose zero-crossing rate r is greater than a preset value a; the first statistics unit is used to count the number N of all short-time Fourier transform time indices m whose zero-crossing rate r is greater than the preset value a if there is a short-time Fourier transform time index m whose zero-crossing rate r is greater than the preset value a.
[0056] In some embodiments, such as this one, the first computing unit 110 includes a second computing unit and a third computing unit.
[0057] The second calculation unit is used to calculate according to the formula. Calculate the average energy of the short-time noise. Where I is the set of all short-time Fourier transform time indices m for which the zero-crossing rate r is greater than the preset value a; the third calculation unit is used to calculate according to the formula Calculate the short-time average energy
[0058] In another embodiment, the output voice signal calculation unit 140 includes a fourth calculation unit and a first processing unit.
[0059] The fourth calculation unit is used to calculate the short-time average gain of the speech gain W(ω). in, The first processing unit is used to utilize the short-time average gain. The frequency domain signal X(k,m) is amplified to obtain the frequency domain signal Y(k,m).
[0060] An embodiment of the present invention also provides a noise reduction device with dynamically controllable noise reduction range. This device adds a first conversion unit to the above embodiment.
[0061] The first conversion unit is used to convert the frequency domain signal Y(k,m) into the time domain signal y(n) through the inverse short-time Fourier transform to obtain the time domain signal of the output speech signal.
[0062] An embodiment of the present invention also provides a noise reduction device with dynamically controllable noise reduction range. This device adds a first acquisition unit and a second conversion unit to the above embodiment.
[0063] The first acquisition unit is used to acquire the time-domain signal x(n) of the input speech signal, where n is a discrete-time index; the second conversion unit is used to convert the time-domain signal into the frequency-domain signal X(k,m) through short-time Fourier transform.
[0064] Figure 6 This is a schematic block diagram of a noise reduction device 200 with dynamically controllable noise reduction range provided in another embodiment of the present invention. Figure 6 As shown, corresponding to the above-described noise reduction method with dynamically controllable noise reduction range, the present invention also provides a noise reduction device 200 with dynamically controllable noise reduction range. This noise reduction device 200 includes a unit for performing the above-described noise reduction method with dynamically controllable noise reduction range. Specifically, please refer to... Figure 6 The noise reduction device 200, which can dynamically control the noise reduction range, includes a first calculation unit 110, a short-time speech energy calculation unit 120, a speech gain calculation unit 130, an output speech signal calculation unit 140, and a third conversion unit 150.
[0065] The first calculation unit 110 is used to acquire the frequency domain signal X(k,m) of the input speech signal and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index, and m is the short-time Fourier transform time index; the short-time speech energy calculation unit 120 is used to calculate the short-time average energy. and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy The speech gain calculation unit 130 is used to calculate the short-time speech energy. and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β and γ are all adjustable parameters used to control the noise reduction range; the output speech signal calculation unit 140 is used to amplify the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal; the third conversion unit 150 is used to convert the frequency domain signal Y(k,m) into the time domain signal y(n) by improving the discrete cosine transform to obtain the time domain signal of the output speech signal.
[0066] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the noise reduction device and each unit with dynamically controllable noise reduction range can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0067] The aforementioned noise reduction device with dynamically controllable noise reduction range can be implemented as a computer program, which can, for example... Figure 7 It runs on the computer device shown.
[0068] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of this application. It can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0069] See Figure 7 The computer device 500 includes a processor 502, a memory, and an interface 507 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0070] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it causes the processor 502 to execute a noise reduction method with dynamically controllable noise reduction range.
[0071] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0072] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a noise reduction method that can dynamically control the noise reduction range.
[0073] This interface 505 is used for communication with other devices. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0074] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:
[0075] Obtain the frequency domain signal X(k,m) of the input speech signal, and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index;
[0076] According to the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy
[0077] The short-time speech energy and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range;
[0078] The frequency domain signal X(k,m) is amplified using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal.
[0079] In one embodiment, the processor 502 further performs the following steps:
[0080] Calculate the zero-crossing rate r of each short-time Fourier transform time index m in the frequency domain signal X(k,m), and determine whether there is a short-time Fourier transform time index m whose zero-crossing rate r is greater than a preset value a.
[0081] If there exists a short-time Fourier transform time index m whose zero-crossing rate r is greater than the preset value a, then count the number N of all short-time Fourier transform time indices m whose zero-crossing rate r is greater than the preset value a.
[0082] In one embodiment, the processor 502 calculates the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy The specific steps are as follows:
[0083] According to the formula Calculate the average energy of the short-time noise. Where I is the set of all short-time Fourier transform time indices m for which the zero-crossing rate r is greater than the preset value a;
[0084] According to the formula Calculate the short-time average energy
[0085] In one embodiment, when the processor 502 implements the step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal, the specific steps are as follows:
[0086] Calculate the short-time average gain of the speech gain W(ω). in,
[0087] Using the short-time average gain The frequency domain signal X(k,m) is amplified to obtain the frequency domain signal Y(k,m).
[0088] In one embodiment, after implementing the step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal, the processor 502 further includes the following step:
[0089] The frequency domain signal Y(k,m) is converted into the time domain signal y(n) by inverse short-time Fourier transform to obtain the time domain signal of the output speech signal.
[0090] In one embodiment, after implementing the step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal, the processor 502 further includes the following step:
[0091] The frequency domain signal Y(k,m) is converted into the time domain signal y(n) by improving the discrete cosine transform to obtain the time domain signal of the output speech signal.
[0092] In one embodiment, when implementing the step of acquiring the frequency domain signal X(k,m) of the input speech signal, the processor 502 specifically implements the following steps:
[0093] Obtain the time-domain signal x(n) of the input speech signal, where n is a discrete-time index;
[0094] The time-domain signal is converted into the frequency-domain signal X(k,m) by short-time Fourier transform.
[0095] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (FSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0096] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0097] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program implements any embodiment of the noise reduction method described above for dynamically controllable noise reduction range.
[0098] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0099] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0100] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0101] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0103] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0104] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.
[0105] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A noise reduction method with dynamically controllable noise reduction range, characterized in that, The method includes: Obtain the frequency domain signal X(k,m) of the input speech signal, and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index; According to the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, short-time speech energy =Short-time average energy -Short-time noise average energy The short-time speech energy and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range; The frequency domain signal X(k,m) is amplified using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal; Calculate the zero-crossing rate r of each short-time Fourier transform time index m in the frequency domain signal X(k,m), and determine whether there is a short-time Fourier transform time index m whose zero-crossing rate r is greater than a preset value a. If there exists a short-time Fourier transform time index m whose zero-crossing rate r is greater than the preset value a, then count the number N of all short-time Fourier transform time indices m whose zero-crossing rate r is greater than the preset value a.
2. The noise reduction method with dynamically controllable noise reduction range as described in claim 1, characterized in that, The calculation of the short-time average energy of the frequency domain signal X(k,m) is described. and short-time noise average energy The steps include: According to the formula Calculate the average energy of the short-time noise. Where I is the set of all short-time Fourier transform time indices m for which the zero-crossing rate r is greater than the preset value a; According to the formula Calculate the short-time average energy 3. The noise reduction method with dynamically controllable noise reduction range as described in claim 1, characterized in that, The step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal includes: Calculate the short-time average gain of the speech gain W(ω). in, Using the short-time average gain The frequency domain signal X(k,m) is amplified to obtain the frequency domain signal Y(k,m).
4. The noise reduction method with dynamically controllable noise reduction range as described in claim 1, characterized in that, After the step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal, the method further includes: The frequency domain signal Y(k,m) is converted into the time domain signal y(n) by inverse short-time Fourier transform to obtain the time domain signal of the output speech signal.
5. The noise reduction method with dynamically controllable noise reduction range as described in claim 1, characterized in that, After the step of amplifying the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal, the method further includes: The frequency domain signal Y(k,m) is converted into the time domain signal y(n) by improving the discrete cosine transform to obtain the time domain signal of the output speech signal.
6. The noise reduction method with dynamically controllable noise reduction range as described in claim 1, characterized in that, The step of obtaining the frequency domain signal X(k,m) of the input speech signal further includes: Obtain the time-domain signal x(n) of the input speech signal, where n is a discrete-time index; The time-domain signal is converted into the frequency-domain signal X(k,m) by short-time Fourier transform.
7. A noise reduction device with dynamically controllable noise reduction range, characterized in that, The device includes: The first calculation unit is used to acquire the frequency domain signal X(k,m) of the input speech signal and calculate the short-time average energy of the frequency domain signal X(k,m). and short-time noise average energy Where k is the discrete spectrum index and m is the short-time Fourier transform time index; Short-time speech energy calculation unit, used to calculate based on the short-time average energy and the short-time noise average energy Calculate short-time speech energy Among them, the short-time speech energy =Short-time average energy -Short-time noise average energy The speech gain calculation unit is used to calculate the short-time speech energy. and the short-time noise average energy Substituting the gain function W, we obtain the speech gain W(ω), where the gain function... α, β, and γ are all adjustable parameters used to control the noise reduction range; The output speech signal calculation unit is used to amplify the frequency domain signal X(k,m) using the speech gain W(ω) to obtain the frequency domain signal Y(k,m) of the output speech signal; The first judgment unit is used to calculate the zero-crossing rate r of each short-time Fourier transform time index m in the frequency domain signal X(k,m), and to determine whether there is a short-time Fourier transform time index m whose zero-crossing rate r is greater than a preset value a. The first statistical unit is used to count the number N of all short-time Fourier transform time indices m whose zero-crossing rate r is greater than the preset value a if there exists a short-time Fourier transform time index m whose zero-crossing rate r is greater than the preset value a.
8. A computer device, characterized in that, The computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to perform the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Static noise reduction method and device thereof, computer equipment and storage medium
CN114005456A