Noise reduction gain processing method and apparatus, medium, device, and speech noise reduction method

By normalizing and convolutionally smoothing the noise reduction gain of the low-frequency part of noisy speech, the problems of high-frequency speech damage and large computational load are solved, achieving a highly efficient noise reduction effect.

CN115910088BActive Publication Date: 2026-01-02武汉斗鱼鱼乐网络科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211574861.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2026-01-02
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Existing technologies suffer from severe speech damage in the high-frequency range when performing noise reduction on noisy speech, and require a large amount of computation. In particular, when estimating noise across the entire frequency band, directly applying the noise reduction gain from the low-frequency range to the high-frequency range results in poor performance.

Method used

By normalizing and convolutionally smoothing the noise reduction gain of the low-frequency part of the noisy speech, the noise reduction gain of the high-frequency part is obtained. The normalized window function of convolutional smoothing is then used to process the noise reduction gain of the low-frequency part, thereby reducing speech impairment in the high-frequency part and reducing the computational load.

Benefits of technology

It effectively reduces speech damage in the high-frequency part of noisy speech, while reducing the consumption of computing resources and improving the noise reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910088B_ABST
    Figure CN115910088B_ABST
Patent Text Reader

Abstract

The application provides a noise reduction gain processing method and device, a medium, an electronic device and a voice noise reduction method. The method comprises the following steps: obtaining a noise reduction gain of a low-frequency part of noisy voice; selecting a window function, performing normalization processing on the selected window function to obtain a normalized window function; and performing convolution smoothing processing on the noise reduction gain of the low-frequency part of the noisy voice and the normalized window function to obtain a noise reduction gain of a high-frequency part of the noisy voice. The noise reduction gain of the high-frequency part of the noisy voice obtained by the application can effectively reduce the voice damage of the high-frequency part of the noisy voice when the noise reduction gain is applied to the noise reduction of the high-frequency part of the noisy voice, and noise estimation does not need to be performed on the full frequency band, and the calculation amount is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of voice, in particular, relates to a noise reduction gain processing method and device, medium, equipment and voice noise reduction method. BACKGROUND

[0002] The noisy voice is formed by the pure voice and the mixed noise, and the noise can be additive or non-additive. In order to perform noise reduction processing on the noisy voice, the noise estimation method is usually used to obtain the noise reduction gain, and then the noise reduction gain is used to obtain the noise reduction voice signal.

[0003] At present, in the actual processing process of the noisy voice, if the noise estimation is performed on the full frequency band of the noisy voice, the calculation amount is large. In order to reduce the calculation amount of noise estimation, the noise estimation is usually performed only on the low frequency part of the noisy voice to obtain the noise reduction gain of the low frequency part of the noisy voice, and then the noise reduction gain of the low frequency part of the noisy voice is applied to the high frequency part of the noisy voice.

[0004] For example, for the 48kHz noisy voice, the noise estimation is performed only on the low frequency part of 50Hz-8kHz to obtain the noise reduction gain of 50Hz-8kHz, and then the noise reduction gain of 50Hz-8kHz is directly applied to the high frequency part of 8kHz-24kHz. For example, the minimum value of the noise reduction gain in a frame of the low frequency part of the noisy voice is directly applied to the high frequency part, but this direct application method has certain defects and can cause serious voice damage to the high frequency part of the noisy voice.

[0005] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0006] The purpose of the present application is to provide a noise reduction gain processing method, device, medium, equipment and voice noise reduction method. The present application can perform normalization and convolution smoothing processing on the noise reduction gain of the low frequency part of the noisy voice to obtain the noise reduction gain of the high frequency part of the noisy voice, so as to perform noise reduction processing on the high frequency part of the noisy voice by using the noise reduction gain, thereby reducing the voice damage of the high frequency part of the noisy voice.

[0007] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned partly through practice of the present application.

[0008] According to one aspect of the embodiment of the present application, a noise reduction gain processing method is provided, comprising the following steps:

[0009] obtaining the noise reduction gain of the low frequency part of the noisy voice;

[0010] Select a window function and normalize it to obtain a normalized window function;

[0011] The noise reduction gain of the low-frequency part of the noisy speech is obtained by convolution smoothing the noise reduction gain and normalized window function.

[0012] In one embodiment of this application, based on the foregoing scheme, the window function is one of Hamming window, Hanning window, triangular window, Gaussian window, and rectangular window.

[0013] In one embodiment of this application, based on the foregoing scheme, the formula for the window function is as follows:

[0014]

[0015] , where N is the window length of the window function and n is the number of points of the window function.

[0016] In one embodiment of this application, based on the aforementioned scheme, 1 / 24 of the window length of the low-frequency portion of the noisy speech is less than or equal to 1 / 12 of the window length of the low-frequency portion of the noisy speech.

[0017] In one embodiment of this application, based on the foregoing scheme, the normalization process of the selected window function to obtain the normalized window function includes: calculating the normalized window function using the following formula:

[0018]

[0019] n = 0, 1, ..., (N-1), where w normalized w(n) is the normalized window function, w(i) is the window function, n is the point of the window function and the normalized window function, and i is the point of the window function.

[0020] In one embodiment of this application, based on the foregoing scheme, the step of performing convolutional smoothing on the noise reduction gain of the low-frequency portion of the noisy speech and the normalized window function to obtain the noise reduction gain of the high-frequency portion of the noisy speech includes: calculating the noise reduction gain of the high-frequency portion of the noisy speech using the following formula:

[0021] in, Let g(k) be the noise reduction gain for the high-frequency portion of the noisy speech, g(k) be the noise reduction gain for the low-frequency portion of the noisy speech, and k represent the frequency point corresponding to the noisy speech. (FFT) SIZE Let be the window length for the low-frequency portion of the noisy speech, and i be a point of the normalized window function.

[0022] According to an aspect of the embodiments of the present application, a device for processing noise reduction gain is provided, the device comprising: an obtaining unit configured to obtain noise reduction gain of a low frequency part of noisy speech; a normalization unit configured to select a window function, and normalize the selected window function to obtain a normalized window function; and a data processing unit configured to perform convolution smoothing processing on the noise reduction gain of the low frequency part of noisy speech and the normalized window function to obtain noise reduction gain of a high frequency part of noisy speech.

[0023] According to an aspect of the embodiments of the present application, a computer readable storage medium having stored thereon a computer program comprising executable instructions which, when executed by a processor, implement the method for processing noise reduction gain as described in the above embodiments.

[0024] According to an aspect of the embodiments of the present application, an electronic device is provided, comprising: one or more processors; and a memory configured to store executable instructions of the processor, which, when executed by the one or more processors, cause the one or more processors to implement the method for processing noise reduction gain as described in the above embodiments.

[0025] According to an aspect of the embodiments of the present application, a method for processing noise reduction gain is provided, comprising the steps of:

[0026] estimating noise reduction gain of a low frequency part of noisy speech;

[0027] using the noise reduction gain processing method described above, obtaining noise reduction gain of a high frequency part of noisy speech according to the noise reduction gain of the low frequency part of noisy speech;

[0028] performing noise reduction processing on the low frequency part of noisy speech using the noise reduction gain of the low frequency part of noisy speech to obtain a low frequency part of noise-reduced speech, and performing noise reduction processing on the high frequency part of noisy speech using the noise reduction gain of the high frequency part of noisy speech to obtain a high frequency part of noise-reduced speech;

[0029] combining the low frequency part of noise-reduced speech and the high frequency part of noise-reduced speech to obtain a noise-reduced speech signal.

[0030] In an embodiment of the present application, based on the foregoing scheme, before performing noise reduction processing on the high frequency part of noisy speech using the noise reduction gain of the high frequency part of noisy speech, the low frequency part of noisy speech and the high frequency part of noisy speech are time-delayed and aligned according to the time delay generated by estimating the noise reduction gain of the low frequency part of noisy speech.

[0031] In an embodiment of the present application, based on the foregoing scheme, the noisy speech is divided into at least two sub-bands in the time domain, the low frequency part of noisy speech is a sub-band with lower frequency, and the high frequency part of noisy speech is a sub-band with higher frequency.

[0032] According to an aspect of the embodiments of the present application, a voice noise reduction device is provided, the device comprising:

[0033] a low frequency noise reduction gain calculation unit configured to estimate a noise reduction gain of a low frequency part of a noisy voice;

[0034] a high frequency noise reduction gain calculation unit configured to calculate a noise reduction gain of a high frequency part of the noisy voice according to the noise reduction gain of the low frequency part of the noisy voice by using the voice noise reduction method;

[0035] a noise reduction processing unit configured to perform noise reduction processing on the low frequency part of the noisy voice by using the noise reduction gain of the low frequency part of the noisy voice to obtain a noise-reduced low frequency part of the voice, and perform noise reduction processing on the high frequency part of the noisy voice by using the noise reduction gain of the high frequency part of the noisy voice to obtain a noise-reduced high frequency part of the voice;

[0036] a synthesis unit configured to synthesize the noise-reduced low frequency part of the voice and the noise-reduced high frequency part of the voice to obtain a noise-reduced voice signal.

[0037] In the technical solution of the embodiments of the present application, the noise reduction gain of the low frequency part of the noisy voice is convoluted and smoothed with the normalized window function to obtain the noise reduction gain of the high frequency part of the noisy voice, so that the noise reduction gain of the high frequency part of the noisy voice is between 0 and 1. When the noise reduction gain of the high frequency part of the noisy voice is applied to the noise reduction of the high frequency part of the noisy voice, the voice damage of the high frequency part of the noisy voice is effectively reduced, and the noise reduction gain of the high frequency part of the noisy voice does not need to be directly calculated, so that the calculation amount is reduced and the calculation resource is saved.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0039] The drawings incorporated into the specification and forming a part thereof, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application. It is clear that the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. In the drawings:

[0040] Figure 1 a flow chart of a noise reduction gain processing method according to an embodiment of the present application;

[0041] Figure 2 a block diagram of a noise reduction gain processing device according to an embodiment of the present application;

[0042] Figure 3 a schematic diagram of a computer readable storage medium according to an embodiment of the present application;

[0043] Figure 4 A schematic diagram of a system structure of an electronic device according to an embodiment of the present application.

[0044] Figure 5 A flow chart of a voice noise reduction method according to an embodiment of the present application;

[0045] Figure 6 A block diagram of a voice noise reduction device according to an embodiment of the present application;

[0046] Figure 7 A noise reduction framework based on STFT analysis according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided as non-limiting examples so that this disclosure will be thorough and complete, and will fully convey the scope of the example implementations to those skilled in the art. Like reference numerals may refer to like elements throughout the description of the figures.

[0048] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the

[0049] The block diagrams in the drawings show only the functionality of the features and can not imply a necessity of any specific arrangements. That is, the functionality of the features can be implemented in software, hardware, or a combination thereof, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0050] The flow diagrams depicted in the drawings show the functionality of the example implementations and are not necessarily the steps of a computer program. That is, the functionality of the flow diagrams can be implemented in software, hardware, or a combination thereof, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0051] It should be noted that the "multiple" mentioned in the present document refers to two or more than two. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents that the front and rear associated objects are in an "or" relationship.

[0052] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the objects thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described.

[0053] The implementation details of the technical solutions of the embodiments of the present application are described in detail as follows:

[0054] First of all, it should be noted that the noise reduction gain processing scheme proposed in the present application can be applied to the related technical field of noisy speech, for example, for noisy speech, a noise estimation method is used to estimate the noise reduction gain of noisy speech, and the estimation result of the noise reduction gain directly affects the noise reduction effect of noisy speech. Therefore, the accuracy of noise reduction gain estimation is particularly important for the noise reduction effect of noisy speech.

[0055] According to one aspect of the present application, a noise reduction gain processing method is provided, Figure 1 For the flowchart of the noise reduction gain processing method shown in the embodiments of the present application, the noise reduction gain processing method can be executed by a device with computing processing function, and the noise reduction gain processing method includes the following steps, which are introduced in detail as follows:

[0056] Obtain the noise reduction gain of the low frequency part of the noisy speech.

[0057] In the present application, the noise reduction gain of the low frequency part of the noisy speech can be determined by a noise estimation method, which can use quantile noise estimation method, histogram noise estimation method, minimum value tracking noise estimation method, or recursive smoothing algorithm noise estimation method.

[0058] It should be emphasized that the minimum value tracking noise estimation method includes a frame minimum value tracking noise estimation method and a continuous spectrum minimum value tracking noise estimation method. The present application estimates the noise reduction gain of the low frequency part of the noisy speech directly, that is, directly calculates, and then obtains the noise reduction gain of the high frequency part of the noisy speech by derivation, rather than deriving the former by estimating the latter. The reason is that the effective components of the noisy speech are mainly concentrated in the low frequency part of 50Hz-8kHz, and the high frequency part exceeding 8kHz is relatively less important.

[0059] With reference to the foregoing Figure 1 , a window function is selected, and the selected window function is normalized to obtain a normalized window function.

[0060] In the present application, the selection of the window function is not limited, and the window function is normalized by dividing each point of the window function by the sum of all points of the window function. The purpose of normalizing the window function is to ensure that the noise reduction gain of the high frequency part of the noisy speech obtained by convolution smoothing of the noise reduction gain of the low frequency part of the noisy speech and the normalized window function is between 0 and 1, so as to avoid the noise reduction gain of the high frequency part of the noisy speech being greater than 1, thereby bringing in other noise and making it difficult to complete the noise reduction work.

[0061] With reference to the foregoing Figure 1 , the noise reduction gain of the high frequency part of the noisy speech is obtained by convolution smoothing of the noise reduction gain of the low frequency part of the noisy speech and the normalized window function.

[0062] In the present application, the noise reduction gain of the high frequency part of the noisy speech is obtained by convolution smoothing of the noise reduction gain of the low frequency part of the noisy speech and the normalized window function, that is, the noise reduction gain of the low frequency part of the noisy speech is convolved with the normalized window function, so as to smooth the noise reduction gain of the low frequency part of the noisy speech and avoid directly applying the noise reduction gain of the low frequency part of the noisy speech to the high frequency part of the noisy speech, which may cause damage to the speech of the high frequency part of the noisy speech. At the same time, the noise estimation of the high frequency part of the noisy speech is not directly performed, which effectively reduces the calculation amount of the noise estimation and saves the calculation resources.

[0063] In an embodiment of the present application, the window function can be a Hamming window, a Hanning window, a triangular window, a Gaussian window, or a rectangular window.

[0064] In the present application, the window function and the noise reduction gain of the low frequency part of the noisy speech are matched, and the noise reduction gain of the low frequency part of the noisy speech is processed by convolution smoothing to obtain the noise reduction gain of the high frequency part of the noisy speech.

[0065] In an embodiment of the present application, the formula of the window function is as follows:

[0066]

[0067] , wherein N is the window length of the window function, and n is the point of the window function.

[0068] Preferably, 1 / 24 of the window length of the low frequency part of the noisy speech≤N≤1 / 12 of the window length of the low frequency part of the noisy speech.

[0069] Specifically, when the noisy speech is framed, each frame is 10-30 ms, if the noisy speech is 48 kHz, each frame is 10 ms, when the noisy speech is windowed, the frame length of the noisy speech is 48000*0.01=480, the low frequency part of the 48 kHz noisy speech is set as 0-8 kHz, according to the Nyquist sampling theorem, the effective upper limit of the spectrum is 24 kHz, so the frame length of the 0-8 kHz band is 480 / 3=160, the window length is the sum of the frame length and the frame shift, and usually the frame shift is half of the frame length, so the window length of the 0-8 kHz band is 160+160 / 2=240, and the value of N is 10-20.

[0070] If the noisy speech is 32 kHz, each frame is 10 ms, when the noisy speech is windowed, the frame length of the noisy speech is 32000*0.01=320, the low frequency part of the 32 kHz noisy speech is set as 0-8 kHz, according to the Nyquist sampling theorem, the effective upper limit of the spectrum is 16 kHz, so the frame length of the 0-8 kHz band is 320 / 2=160, the window length is the sum of the frame length and the frame shift, and usually the frame shift is half of the frame length, so the window length of the 0-8 kHz band is 160+160 / 2=240, and the value of N is 10-20.

[0071] In an embodiment of the present application, the normalization processing on the selected window function is performed to obtain a normalized window function, including: calculating the normalized window function by using the following formula:

[0072]

[0073] n=0, 1, …, (N-1), wherein, w normalized (n) is the normalized window function, w(i) is the window function, n is the point of the window function and the normalized window function, and i is the point of the window function.

[0074] Specifically, the normalized window function is the quotient value of each point of the window function and the sum of all points of the window function, that is, the sum of all points of the window function is obtained Then, the value of each point of the window function is divided by to obtain the value w normalizet (n) of each point of the normalized window function, when n=0, the value w

[0075] In an embodiment of the present application, the noise reduction gain of the high frequency part of the noisy speech is obtained by performing convolution smoothing processing on the noise reduction gain of the low frequency part of the noisy speech and the normalized window function, including: calculating the noise reduction gain of the high frequency part of the noisy speech by using the following formula:

[0076] wherein, is the noise reduction gain of the high frequency part of the noisy speech, g(k) is the noise reduction gain of the low frequency part of the noisy speech, k represents the frequency point corresponding to the noisy speech, FFT SIZE is the window length of the low frequency part of the noisy speech, and i is the point of the normalized window function.

[0077] Specifically, FFT SIZE is the window length of the low frequency part of the noisy speech, and N is not the same as M, according to the above embodiment, then FFT SIZE = 240, here, k takes the value of 0-120, the main reason is that when performing FFT transformation, the figure is in a symmetrical state, only the noise reduction gain of the high frequency part of the noisy speech is required, the noise reduction gain of the low frequency part of the noisy speech is the convolution of the noise reduction gain and the normalized window function, when k = 0, the noise reduction gain of the high frequency part of the noisy speech is obtained

[0078] In the above embodiment, the Hanning window is used to convolve and smooth the noise reduction gain of the low frequency part of the noisy speech, if the Hamming window is used to convolve and smooth, then the noise reduction gain of the high frequency part of the noisy speech is convolved and smoothed by using similar calculation formula, and the window function formula is:

[0079]

[0080] M represents the window length of the window function, m is the point of the window function,

[0081] α is usually 0.46, if α is 0.46, then the normalized window function is:

[0082]

[0083] m = 0, 1,..., (M-1), w(i) is the window function, m is the point of the window function and the normalized window function, and i is the point of the window function,

[0084] The calculation formula of the noise reduction gain of the high frequency part of the noisy speech is as follows:

[0085]

[0086] is the noise reduction gain of the high frequency part of the noisy speech, g(j) is the noise reduction gain of the low frequency part of the noisy speech, j represents the frequency point corresponding to the noisy speech, FFT SIZE is the window length of the low frequency part of the noisy speech, and i is the point of the normalized window function.

[0087] It should be pointed out that if the triangular window, rectangular window, Gaussian window or other window functions are used, similar formula can also be used to convolve and smooth the noise reduction gain of the high frequency part of the noisy speech, which is not limited here.

[0088] And, the points of the window function and the points of the normalized window function are one-to-one corresponding.

[0089] Figure 2 A block diagram of a noise reduction gain processing device according to an embodiment of the present application is shown.

[0090] Referring to Figure 2 As shown in the figure, the noise reduction gain processing device 100 according to an embodiment of the present application comprises an acquisition unit 101, a normalization unit 102 and a data processing unit 103.

[0091] The acquisition unit 101 is configured to acquire a noise reduction gain of a low frequency part of noisy speech; the normalization unit 102 is configured to select a window function, normalize the selected window function to obtain a normalized window function; and the data processing unit 103 is configured to perform convolution smoothing processing on the noise reduction gain of the low frequency part of noisy speech and the normalized window function to obtain a noise reduction gain of a high frequency part of noisy speech.

[0092] As another aspect, the present application also provides a computer readable storage medium having stored thereon a program product capable of implementing the noise reduction gain processing method described above in the specification. In some possible implementation manners, various aspects of the present application can also be implemented in the form of a program product, which includes program codes for causing terminal equipment to perform the steps described in the above "Exemplary Method" section according to various exemplary embodiments of the present application when the program product is run on the terminal equipment.

[0093] Referring to Figure 3 As shown in the figure, a program product 200 for implementing the above method according to an embodiment of the present application is described, which can adopt a portable compact disc read-only memory (CD-ROM) and include program codes, and can be run on terminal equipment such as a personal computer. However, the program product of the present application is not limited to this, and in this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, apparatus or device.

[0094] The program product can employ any combination of one or more computer readable media or storage media. The computer readable media or storage media can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0095] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein. The propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport program code.

[0096] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0097] Program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, application specific circuitry, or field programmable gate array (FPGA) circuitry, includes the circuitry employed in the microprocessors, optical chips, central processing units (CPU), graphics processing units (GPU), digital signal processors (DSP), digital signal processing devices (DSPD), or digital signal processing device (DSPD), programmable logic devices (PLD), programmable logic (PL), controllers, state machines, gated logic, discrete hardware components, or any other processing circuitry, functioning individually or in combination, and are

[0098] As another aspect, the present application provides an electronic device capable of implementing the above method.

[0099] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a program product. Therefore, each aspect of the present application can be specifically implemented as follows: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.

[0100] The electronic device 300 according to this embodiment of the present application will be described below with reference to Figure 4 Figure 4 The electronic device 300 shown is merely an example and should not impose any limitation on the function and scope of use of the embodiments of the present application.

[0101] As Figure 4 shown, the electronic device 300 is in the form of a general computing device. The components of the electronic device 300 can include, but are not limited to, the at least one processing unit 310 described above, the at least one storage unit 320 described above, and a bus 330 connecting different system components, including the storage unit 320 and the processing unit 310.

[0102] The storage unit stores program codes which can be executed by the processing unit 310, so that the processing unit 310 performs the steps according to various exemplary embodiments of the present application described in the "Embodiment Method" part of the present specification.

[0103] The storage unit 320 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 321 and / or a cache memory 322, and can further include a read-only memory (ROM) 323.

[0104] The storage unit 320 can further include program / utilities 324 having a set of (at least one) program modules 325, such as an operating system, one or more application programs, other program modules, and program data, each of which or some combination of which can include implementation of a network environment.

[0105] The bus 330 can represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0106] ​The electronic device 300 can also communicate with one or more external devices 400 such as a keyboard, a pointing device, a Bluetooth device, etc.; and / or one or more devices that enable a user to interact with the electronic device 300 and / or one or more devices (e.g. routers, modems, etc.) that enable the electronic device 300 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 350. Still yet, the electronic device 300 can communicate with one or more networks such as a local area network (LAN), a wide area network (WAN), and / or the Internet through network adapter 360. As depicted, network adapter 360 communicates with the other components of the electronic device 300 via bus 330. It should be appreciated that although not shown, other hardware and / or software modules could be used in connection with the electronic device 300. Examples, include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0107] Figure 5 Flow chart of the voice noise reduction method according to the embodiment of the present application

[0108] As another aspect, referring to Figure 5 , a voice noise reduction method is provided, which can be executed by a device with computing processing function, and the voice noise reduction method comprises the following steps:

[0109] The noise reduction gain of the low frequency part of the noisy voice is estimated, and specifically, the noise reduction gain of the low frequency part of the noisy voice is estimated by using a noise estimation method. In this embodiment, one of the quantile noise estimation method, the histogram noise estimation method, the minimum value tracking noise estimation method and the recursive smoothing algorithm noise estimation method can be used to estimate the noise reduction gain of the low frequency part of the noisy voice.

[0110] The noise reduction gain of the high frequency part of the noisy voice is obtained according to the noise reduction gain of the low frequency part of the noisy voice by using the noise reduction gain processing method. Specifically, a window function is selected, the selected window function is normalized to obtain a normalized window function, and the noise reduction gain of the low frequency part of the noisy voice and the normalized window function are convolved and smoothed to obtain the noise reduction gain of the high frequency part of the noisy voice.

[0111] The low frequency part of the noisy voice is processed by using the noise reduction gain of the low frequency part of the noisy voice to obtain the low frequency part of the voice after noise reduction, and the high frequency part of the noisy voice is processed by using the noise reduction gain of the high frequency part of the noisy voice to obtain the high frequency part of the voice after noise reduction.

[0112] The low frequency part of the voice after noise reduction and the high frequency part of the voice after noise reduction are synthesized to obtain the voice signal after noise reduction.

[0113] In one embodiment of this application, before denoising the high-frequency part of the noisy speech using the denoising gain of the high-frequency part of the noisy speech, the delay generated by estimating the denoising gain of the low-frequency part of the noisy speech, which is generated during the estimation of the denoising gain by the noise estimation method, is used to align the delay between the low-frequency part of the noisy speech and the high-frequency part of the noisy speech. That is, after aligning the delay between the high-frequency part of the noisy speech and the low-frequency part of the noisy speech, the denoising gain of the high-frequency part of the noisy speech is applied to the high-frequency part of the noisy speech.

[0114] Specifically, according to the above embodiment, when k = 0, the following is obtained: Will Before applying the initial point to the high-frequency part of the noisy speech, delay alignment is performed based on the delay generated during noise estimation of the low-frequency part of the noisy speech, and then... The initial point applied to the high-frequency part of noisy speech.

[0115] In one embodiment of this application, the noisy speech is divided into at least two sub-bands in the time domain, with the low-frequency part of the noisy speech being the lower-frequency sub-band and the high-frequency part of the noisy speech being the higher-frequency sub-band.

[0116] To better understand this embodiment, a specific example is provided. Specifically, the noisy speech is a 48kHz speech signal, which is divided into three sub-bands: 0–8kHz, 8–16kHz, and 16–24kHz. The 0–8kHz sub-band can be considered the low-frequency portion of the noisy speech, while the 8–16kHz and 16–24kHz sub-bands represent the high-frequency portions. The noise reduction gain estimated based on the 0–8kHz sub-band is the noise reduction gain for the low-frequency portion of the noisy speech. After calculation using the aforementioned noise reduction gain processing method, the noise reduction gain for the high-frequency portion of the noisy speech is obtained. This noise reduction gain for the high-frequency portion is then applied to the 8–16kHz and 16–24kHz sub-bands. Two sub-bands of 4kHz are used for noise reduction, and two sub-bands of 8-16kHz and 16-24kHz are used for noise reduction. Alternatively, the two sub-bands of 0-8kHz and 8-16kHz can be used for the low-frequency part of the noisy speech, and the two sub-bands of 16-24kHz can be used for the high-frequency part of the noisy speech. In this case, the noise reduction gain is estimated based on one of the 0-8kHz and 8-16kHz sub-bands as the noise reduction gain for the low-frequency part of the noisy speech. After calculation using the above noise reduction gain processing method, the noise reduction gain for the high-frequency part of the noisy speech is obtained. The noise reduction gain for the high-frequency part of the noisy speech is then applied to the 16-24kHz sub-band for noise reduction.

[0117] See Figure 7 , Figure 7A noise reduction framework based on STFT analysis is provided, specifically, a noisy speech k(n) is divided into frames to obtain x i (m) after which FFT processing is performed to separate the amplitude X i (k) and the phase to obtain the amplitude spectrum |X i (k) and the original phase According to the amplitude X i (k), a noise estimation method is used to estimate the noise reduction gain gain(k), which is multiplied by the amplitude spectrum |X i (k) of the noisy speech to obtain the amplitude spectrum of the noise-reduced speech The amplitude spectrum and the original phase are subjected to IFFT (inverse fast Fourier transform) transformation to obtain the noise-reduced speech signal The noise reduction gain estimated according to the amplitude is the content involved in the noise estimation method in the above embodiment.

[0118] In order to better understand the embodiment, a noise reduction method of speech based on a noise reduction framework of STFT analysis is proposed, specifically, a 48 kHz noisy speech is decomposed in the time domain, since the effective upper limit of the spectrum of the 48 kHz noisy speech is 24 kHz, three subbands of 0-8 kHz, 8-16 kHz and 16-24 kHz are obtained, the 0-8 kHz, 8-16 kHz and 16-24 kHz subbands are divided into frames, each frame is 10 ms, a window function is used to perform windowing processing on the 48 kHz noisy speech, it should be noted that the window function here is not the same as the window function used for the noise reduction gain of the low frequency part of the noisy speech, the window function here is to facilitate the FFT transformation of the 48 kHz noisy speech, the frame length of each subband is 48000*0.01 / 3=160, the window length of each subband is the sum of the frame length 160 and the frame shift, the frame shift is usually half of the frame length, so the window length of each subband is 240, then FFT processing is performed to separate the amplitude and the phase to obtain the amplitude spectrum of the three subbands, a noise estimation method is used to estimate the noise reduction gain of the 0-8 kHz subband to obtain the noise reduction gain g(k) of the 0-8 kHz subband, a window function is selected, the formula of the window function is:

[0119]

[0120] wherein N=240*(1 / 24)=10, then w(n) is normalized by using the following formula:

[0121]

[0122] wherein N=240*(1 / 24)=10, then w normalized(n) and g(k) are convoluted to smooth:

[0123]

[0124] obtaining the noise reduction gain of the 8-16 kHz and 16-24 kHz subbands FFT SIZE for 240, multiplying g(k) and the amplitude spectrum of the 0-8 kHz subband to obtain the amplitude spectrum of the 0-8 kHz subband after noise reduction, using the amplitude spectrum of the 0-8 kHz subband after noise reduction and the phase of the 0-8 kHz subband, performing IFFT (inverse fast Fourier transform) to obtain the speech signal of the 0-8 kHz subband after noise reduction, aligning the 8-16 kHz subband and the 16-24 kHz subband with the 0-8 kHz subband by time delay, the time delay being from the estimation process of the noise reduction gain of the 0-8 kHz subband by the noise estimation method, and then multiplying the amplitude spectrum of the 8-16 kHz subband and the amplitude spectrum of the 16-24 kHz subband with the amplitude spectrum of the 0-8 kHz subband after noise reduction to obtain the amplitude spectrum of the 8-16 kHz subband after noise reduction and the amplitude spectrum of the 16-24 kHz subband after noise reduction, using the amplitude spectrum of the 8-16 kHz subband after noise reduction and the phase of the 8-16 kHz subband to perform IFFT (inverse fast Fourier transform) to obtain the speech signal of the 8-16 kHz subband after noise reduction, using the amplitude spectrum of the 16-24 kHz subband after noise reduction and the phase of the 16-24 kHz subband to perform IFFT (inverse fast Fourier transform) to obtain the speech signal of the 16-24 kHz subband after noise reduction, and then using a subband synthesis function to synthesize the speech signal of the 0-8 kHz subband after noise reduction, the speech signal of the 8-16 kHz subband after noise reduction and the speech signal of the 16-24 kHz subband after noise reduction into a speech signal after noise reduction. respectively with the amplitude spectrum of the 8-16 kHz subband and the amplitude spectrum of the 16-24 kHz subband to obtain the amplitude spectrum of the 8-16 kHz subband after noise reduction and the amplitude spectrum of the 16-24 kHz subband after noise reduction, using the amplitude spectrum of the 8-16 kHz subband after noise reduction and the phase of the 8-16 kHz subband to perform IFFT (inverse fast Fourier transform) to obtain the speech signal of the 8-16 kHz subband after noise reduction, using the amplitude spectrum of the 16-24 kHz subband after noise reduction and the phase of the 16-24 kHz subband to perform IFFT (inverse fast Fourier transform) to obtain the speech signal of the 16-24 kHz subband after noise reduction, and then using a subband synthesis function to synthesize the speech signal of the 0-8 kHz subband after noise reduction, the speech signal of the 8-16 kHz subband after noise reduction and the speech signal of the 16-24 kHz subband after noise reduction into a speech signal after noise reduction.

[0125] Specifically, the subband synthesis function can use any existing subband synthesis function.

[0126] As another aspect, referring to Figure 6 The application provides a speech noise reduction device 500, which comprises:

[0127] a low-frequency noise reduction gain calculation unit 501 configured to estimate the noise reduction gain of the low-frequency part of the noisy speech;

[0128] a high-frequency noise reduction gain calculation unit 502 configured to use the above speech noise reduction method to obtain the noise reduction gain of the high-frequency part of the noisy speech according to the noise reduction gain of the low-frequency part of the noisy speech;

[0129] The noise reduction processing unit 503 performs noise reduction processing on the low frequency part of the noisy speech using the noise reduction gain of the low frequency part of the noisy speech to obtain a low frequency part of the noise-reduced speech, and performs noise reduction processing on the high frequency part of the noisy speech using the noise reduction gain of the high frequency part of the noisy speech to obtain a high frequency part of the noise-reduced speech.

[0130] The synthesis unit 504 synthesizes the low frequency part of the noise-reduced speech and the high frequency part of the noise-reduced speech to obtain a noise-reduced speech signal.

[0131] Specifically, the high frequency noise reduction gain calculation unit 502 can be the noise reduction gain processing device 100.

[0132] It should be noted that the main purpose of the present application is to reduce the calculation amount in the noise reduction gain processing of the high frequency part of the noisy speech and not to cause great speech damage to the high frequency part of the noisy speech. The main method is to perform convolution smoothing on the noise reduction gain of the low frequency part of the noisy speech, rather than directly applying the noise reduction gain of the low frequency part of the noisy speech to the high frequency part of the noisy speech. The main tool of the convolution is the normalized normalization window function, that is, the parameters of the convolution are the noise reduction gain of the low frequency part of the noisy speech and the normalized normalization window function.

[0133] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or on a network, and includes a plurality of instructions to make a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) execute the method according to the embodiments of the present application.

[0134] In addition, the above-described figures are only schematic illustrations of the processes included in the method according to the example embodiments of the present application, and are not for limiting purposes. It is easy to understand that the processes shown in the above-described figures do not indicate or limit the time sequence of the processes. In addition, it is also easy to understand that the processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0135] It should be understood that the present application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the application is limited only by the claims that follow.

Claims

1. A method of noise reduction gain processing, the method comprising: The method comprises the following steps: obtaining a noise reduction gain of a low frequency part of the noisy speech; selecting a window function, and performing normalization processing on the selected window function to obtain a normalized window function; performing convolution smoothing processing on the noise reduction gain of the low frequency part of the noisy speech and the normalized window function to obtain a noise reduction gain of a high frequency part of the noisy speech.

2. A noise reduction gain processing method according to claim 1, characterized in that, The normalization processing on the selected window function to obtain the normalized window function comprises: calculating the normalized window function by using the following formula: n = 0, 1,..., (N-1), where w normalized (n) is a normalized window function, w(n) and w(i) are window functions, n is a point of the window function and the normalized window function, i is a point of the window function, and N is a window length of the window function.

3. A noise reduction gain processing method according to claim 2, wherein, The convolution smoothing processing on the noise reduction gain of the low frequency part of the noisy speech and the normalized window function to obtain the noise reduction gain of the high frequency part of the noisy speech comprises: calculating the noise reduction gain of the high frequency part of the noisy speech by using the following formula: wherein, is the noise reduction gain for the high frequency portion of the noisy speech, g(k) is the noise reduction gain for the low frequency portion of the noisy speech, k represents the frequency bin corresponding to the noisy speech, FFT SIZE is the window length for the low frequency portion of the noisy speech, i is the point of the normalized window function.

4. The noise reduction gain processing method of claim 1, wherein, The noisy speech is divided into at least two subbands in the time domain, the low frequency part of the noisy speech is a subband with a lower frequency, and the high frequency part of the noisy speech is a subband with a higher frequency.

5. A noise reduction gain processing apparatus characterized by comprising: The device comprises: an obtaining unit configured to obtain a noise reduction gain of a low frequency part of the noisy speech; a normalization unit configured to select a window function, and perform normalization processing on the selected window function to obtain a normalized window function; a data processing unit configured to perform convolution smoothing processing on the noise reduction gain of the low frequency part of the noisy speech and the normalized window function to obtain a noise reduction gain of a high frequency part of the noisy speech.

6. A computer readable storage medium characterized by, A computer program is stored thereon, and the computer program comprises executable instructions. When the executable instructions are executed by a processor, a noise reduction gain processing method according to any one of claims 1-4 is implemented.

7. An electronic device, comprising: comprise: one or more processors; a memory configured to store executable instructions of the processor, and when the executable instructions are executed by the one or more processors, the one or more processors implement a noise reduction gain processing method according to any one of claims 1-4.

8. A voice noise reduction method, characterized by, The method comprises the following steps: estimating a noise reduction gain of a low frequency part of the noisy speech; obtaining a noise reduction gain of a high frequency part of the noisy speech by using the noise reduction gain processing method according to any one of claims 1-4; performing noise reduction processing on the low frequency part of the noisy speech by using the noise reduction gain of the low frequency part of the noisy speech to obtain a low frequency part of the noise-reduced speech, and performing noise reduction processing on the high frequency part of the noisy speech by using the noise reduction gain of the high frequency part of the noisy speech to obtain a high frequency part of the noise-reduced speech; synthesizing the low frequency part of the noise-reduced speech and the high frequency part of the noise-reduced speech to obtain a noise-reduced speech signal.

9. The voice noise reduction method of claim 8, wherein, Before the noise reduction processing on the high frequency part of the noisy speech by using the noise reduction gain of the high frequency part of the noisy speech, the low frequency part of the noisy speech and the high frequency part of the noisy speech are time-aligned according to a time delay generated by the estimation of the noise reduction gain of the low frequency part of the noisy speech.

10. A voice noise reduction device, characterized by, The device comprises: a low frequency noise reduction gain calculation unit configured to estimate a noise reduction gain of a low frequency part of the noisy speech; a high frequency noise reduction gain calculation unit configured to obtain a noise reduction gain of a high frequency part of the noisy speech by using the noise reduction gain of the low frequency part of the noisy speech according to the speech noise reduction method in claim 8 or 9; The noise reduction processing unit adopts a noise reduction gain of a low frequency part of the noisy speech to perform noise reduction processing on the low frequency part of the noisy speech, and obtains a low frequency part of the speech after noise reduction; adopts a noise reduction gain of a high frequency part of the noisy speech to perform noise reduction processing on the high frequency part of the noisy speech, and obtains a high frequency part of the speech after noise reduction. The synthesis unit synthesizes the low frequency part of the speech after noise reduction and the high frequency part of the speech after noise reduction, and obtains a speech signal after noise reduction.

Citation Information

Patent Citations

  • Speech enhancement method and device, electronic equipment and storage medium

    CN113299308A

  • Training method of frequency band gain model and voice noise reduction method for vehicle-mounted scene

    CN113782011A