Voice noise reduction method, device and storage medium
By dividing the speech signal into frequency domain components and judging the distance to the sound source, and adjusting the noise reduction intensity of high and low frequency components, the speech distortion problem when the sound source is far away is solved, and the quality and robustness of the speech signal are improved.
Patent Information
- Application Number
- CN202111410909.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-11-25
AI Technical Summary
In existing technologies, the speech signal is weak when the sound source is far away, which means that the problem of speech distortion after noise reduction has not been effectively solved.
By acquiring the frequency domain components of the speech signal and dividing them into high-frequency and low-frequency components, the noise reduction intensity of different frequency components is adjusted according to the sound source distance and noise level. The noise reduction intensity of the high-frequency component is no higher than that of the low-frequency component when the sound source distance is not lower than the threshold. Automatic gain control and noise estimation algorithms are used to determine the noise level in order to optimize the noise reduction strategy.
It improves the quality of speech signals, solves the problem of speech distortion when the sound source is far away, and enhances the noise reduction effect and robustness of speech signals.
Smart Images

Figure CN114333885B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, and in particular to a voice noise reduction method and device and a storage medium. BACKGROUND
[0002] Noise generally exists in a voice signal, and in order to obtain a voice signal with high quality, noise reduction processing is usually required.
[0003] In related technologies, the level of noise reduction is adjusted by estimating the size of noise, and the stronger the noise, the higher the level of noise reduction. However, when the distance of a sound source is far, the collected voice signal is weak, and there is a problem of voice distortion after noise reduction.
[0004] At present, there is no effective solution to the problem of voice distortion when the distance of a sound source is far in related technologies. SUMMARY
[0005] A voice noise reduction method, device and storage medium are provided in the present embodiment to solve the problems in related technologies.
[0006] In a first aspect, a voice noise reduction method is provided in the present embodiment, comprising:
[0007] obtaining a voice signal, extracting a frequency domain component of the voice signal, and dividing the frequency domain component into a high frequency component and a low frequency component according to the frequency size, wherein the frequency of the high frequency component is greater than the frequency of the low frequency component;
[0008] obtaining the distance of a sound source of the voice signal;
[0009] performing noise reduction on the voice signal, and in the case that the distance of the sound source is not lower than a preset threshold, the noise reduction intensity of the high frequency component is not higher than the noise reduction intensity of the low frequency component.
[0010] In some embodiments thereof, obtaining the distance of a sound source of the voice signal comprises:
[0011] performing frame processing on the voice signal to obtain a plurality of voice signal frames;
[0012] obtaining a gain value of at least two voice signal frames in the plurality of voice signal frames, and calculating an average gain value of the at least two voice signal frames;
[0013] determining the distance of a sound source of the voice signal according to the average gain value of the at least two voice signal frames.
[0014] In some embodiments thereof, obtaining a gain value of at least two voice signal frames in the plurality of voice signal frames, and calculating an average gain value of the at least two voice signal frames comprises:
[0015] obtaining gain values of a current speech signal frame and N previous speech signal frames in the plurality of speech signal frames, and calculating average gain values of the current speech signal frame and the N previous speech signal frames, wherein N is a positive integer.
[0016] In some embodiments, the speech signal is subjected to noise reduction, and in the case where the sound source distance is not lower than a preset threshold, the noise reduction intensity of the high-frequency component is not higher than the noise reduction intensity of the low-frequency component.
[0017] performing noise estimation on the speech signal to obtain estimated noise of the speech signal;
[0018] determining a noise level of a current environment according to the estimated noise, wherein the noise level comprises a first level and a second level, and noise of the first level is greater than noise of the second level;
[0019] adjusting noise reduction intensity of a high-frequency component and a low-frequency component in the speech signal according to the noise level.
[0020] In some embodiments, adjusting noise reduction intensity of a high-frequency component and a low-frequency component in the speech signal according to the noise level comprises:
[0021] in the case where the noise level of the current environment is the first level, performing low-intensity noise reduction on the high-frequency component of the speech signal, and performing high-intensity noise reduction on the low-frequency component of the speech signal.
[0022] In some embodiments, adjusting noise reduction intensity of a high-frequency component and a low-frequency component in the speech signal according to the noise level comprises:
[0023] in the case where the noise level of the current environment is the second level, performing low-intensity noise reduction on the high-frequency component and the low-frequency component of the speech signal respectively.
[0024] In some embodiments, the method further comprises:
[0025] performing noise estimation on the speech signal to obtain estimated noise of the speech signal;
[0026] determining a noise level of a current environment according to the estimated noise, wherein the noise level comprises a first level and a second level, and noise of the first level is greater than noise of the second level;
[0027] in the case where the sound source distance is lower than a preset threshold, and the noise level of the current environment is the first level, performing high-intensity noise reduction on the high-frequency component and the low-frequency component of the speech signal respectively.
[0028] In some embodiments, the method further comprises:
[0029] In the case that the noise level of the current environment is the second level, the high frequency components and the low frequency components of the speech signal are subjected to low-intensity noise reduction respectively.
[0030] In some embodiments, the speech signal is subjected to noise estimation to obtain estimated noise of the speech signal, and determining the noise level of the current environment according to the estimated noise comprises:
[0031] The speech signal is subjected to frame processing to obtain a plurality of speech signal frames;
[0032] Each of the speech signal frames is subjected to noise estimation to obtain estimated noise of each of the speech signal frames;
[0033] It is determined whether noise energy of continuous M speech signal frames in the speech signal frames exceeds an energy threshold, wherein M is a positive integer;
[0034] In the case that it is determined that noise energy of continuous M speech signal frames in the speech signal frames exceeds the energy threshold, it is determined that the noise level of the current environment is the first level; and,
[0035] In the case that it is determined that noise energy of continuous M speech signal frames in the speech signal frames does not exceed the energy threshold, it is determined that the noise level of the current environment is the second level.
[0036] In some embodiments, the noise reduction comprises low-intensity noise reduction and high-intensity noise reduction, wherein the gain adjustment range of the low-intensity noise reduction is smaller than the gain adjustment range of the high-intensity noise reduction.
[0037] In some embodiments, dividing the frequency domain components into high frequency components and low frequency components according to frequency size comprises:
[0038] The speech signal is subjected to frame processing to obtain a plurality of speech signal frames, and frequency domain information of each of the speech signal frames is extracted;
[0039] According to the frequency domain information, speech signal with frequency not less than a frequency threshold is divided into high frequency components, and speech signal with frequency less than the frequency threshold is divided into low frequency components.
[0040] In a second aspect, an electronic device is provided in the present embodiments, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the speech noise reduction method of the first aspect when executing the computer program.
[0041] In a third aspect, the present embodiment provides a storage medium having stored thereon a computer program which, when executed by a processor, implements the speech noise reduction method of the first aspect.
[0042] Compared with the related art, the speech noise reduction method provided in the present embodiment solves the problem of speech distortion when the sound source distance is far, and improves the quality of the speech signal, by obtaining a speech signal, extracting a frequency domain component of the speech signal, and dividing the frequency domain component into a high frequency component and a low frequency component according to the frequency size, wherein the frequency of the high frequency component is greater than the frequency of the low frequency component; obtaining a sound source distance of the speech signal; and performing noise reduction on the speech signal, and in the case where the sound source distance is not lower than a preset threshold, the noise reduction intensity of the high frequency component is not higher than the noise reduction intensity of the low frequency component.
[0043] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0044] The accompanying drawings, which are included to provide a further understanding of the present application, illustrate embodiments of the present application and together with the description serve to explain the present application. The drawings are not intended to be limiting in any way, and it is contemplated that various embodiments of the present application can be practiced without some or all of these specific details. In the drawings:
[0045] Figure 1 FIG. 1 is a hardware structure block diagram of a terminal of the speech noise reduction method of an embodiment of the present application;
[0046] Figure 2 FIG. 2 is a flowchart of the speech noise reduction method of an embodiment of the present application;
[0047] Figure 3 FIG. 3 is a flowchart of a method of determining the distance of a sound source of a preferred embodiment of the present application;
[0048] Figure 4 FIG. 4 is a flowchart of the speech noise reduction method of a preferred embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to more clearly understand the objects, technical solutions and advantages of the present application, the present application is described and explained below in conjunction with the drawings and embodiments.
[0050] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the same meaning as those commonly understood by a person of ordinary skill in the art to which the present application belongs. The terms "one", "a", "an", "the", "these", and similar terms in the present application do not indicate quantity of limitation, and they can be singular or plural. The terms "include", "contain", "have", and any variant thereof in the present application are intended to cover non-exclusive inclusion; for example, a process, method, and system, product or device containing a series of steps or modules (units) are not limited to the listed steps or modules (units), but can include steps or modules (units) not listed, or can include other steps or modules (units) inherent to the process, method, product or device. The terms "connect", "connected", "couple" and similar terms in the present application are not limited to physical or mechanical connection, but can include electrical connection, whether direct or indirect. The term "multiple" in the present application refers to two or more. The term "and / or" describes the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. Generally, the character " / " represents an "or" relationship between the associated objects. The terms "first", "second", "third" and the like in the present application are only used to distinguish similar objects, and do not represent a specific order of the objects.
[0051] The method embodiments provided in the present embodiment can be executed in a terminal, a computer or a similar computing device. For example, the method embodiments are executed on a terminal, Figure 1 is a hardware structure diagram of the terminal of the voice noise reduction method of the present embodiment. As shown in Figure 1 , the terminal can include one or more (only one is shown in Figure 1 ) processor 102 and memory 104 for storing data, wherein the processor 102 can include but not limited to processing devices such as microprocessor MCU or programmable logic device FPGA. The above terminal can also include a transmission device 106 for communication function and an input / output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above terminal. For example, the terminal can include more or less components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0052] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as the computer program corresponding to the voice noise reduction method in the embodiment. The processor 102 can execute various functional applications and data processing, i.e., implement the method described above, by running the computer programs stored in the memory 104. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories remotely arranged with respect to the processor 102, which can be connected to the terminal through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0053] The transmission device 106 is used to receive or send data via a network. The network described above includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.
[0054] In the embodiment, a voice noise reduction method is provided, Figure 2 A flowchart of the voice noise reduction method of the embodiment is shown in FIG. 2, which includes the following steps: Figure 2
[0055] In step S201, a voice signal is acquired, a frequency domain component of the voice signal is extracted, and the frequency domain component is divided into a high frequency component and a low frequency component according to the frequency size, wherein the frequency of the high frequency component is greater than the frequency of the low frequency component.
[0056] Optionally, the division of the high frequency component and the low frequency component is implemented by the following method:
[0057] The voice signal is subjected to frame processing to obtain a plurality of voice signal frames, and frequency domain information of each voice signal frame is extracted; according to the frequency domain information, the voice signal with a frequency not less than a frequency threshold is divided into a high frequency component, and the voice signal with a frequency less than the frequency threshold is divided into a low frequency component.
[0058] In step S202, a sound source distance of the voice signal is acquired.
[0059] Optionally, the sound source distance of the voice signal is directly obtained by measuring the distance between the sound source position and the voice signal collection position.
[0060] Preferably, the sound source distance of the voice signal is obtained by the following method:
[0061] The voice signal is frame processed to obtain a plurality of voice signal frames, gain values of at least two voice signal frames are obtained from the plurality of voice signal frames, and an average gain value of the at least two voice signal frames is calculated; and the sound source distance of the voice signal is determined according to the average gain value of the at least two voice signal frames.
[0062] The gain value is a gain value of the current voice signal frame obtained by an automatic gain control method (AGC) when performing automatic gain control on the current voice signal frame. The automatic gain control method is that an automatic control circuit automatically adjusts the gain of a power amplifier circuit by identifying the intensity of the collected voice signal, that is, when a person speaks close to the microphone and the voice signal intensity is high, the gain of the power amplifier circuit is reduced, and when a person speaks far from the microphone and the voice intensity is low, the power amplifier gain is increased.
[0063] Preferably, the gain values of the at least two voice signal frames are obtained from the plurality of voice signal frames, and the average gain value of the at least two voice signal frames is calculated by the following method:
[0064] The gain values of the current voice signal frame and the previous N voice signal frames are obtained from the plurality of voice signal frames, and the average gain value of the current voice signal frame and the previous N voice signal frames is calculated, where N is a positive integer.
[0065] In step S203, the voice signal is noise reduced, and in the case that the sound source distance is not lower than a preset threshold, the noise reduction intensity of the high frequency component is not higher than that of the low frequency component.
[0066] Preferably, in the case that the average gain value is not lower than a preset threshold, the noise reduction intensity of the high frequency component is not higher than that of the low frequency component.
[0067] In the related voice noise reduction technology, since the high frequency of the voice signal attenuates quickly, the signal-to-noise ratio of the high frequency voice signal collected from a distant voice signal is very low, part of the high frequency voice signal is submerged in the noise, and in the noise reduction process, the high frequency signal is also attenuated with the noise, resulting in distortion of the high frequency part of the final signal. Through the above steps S201 to S203, the voice signal with a distant sound source distance is divided into a high frequency component and a low frequency component according to the frequency size, a lower intensity noise reduction is performed on the high frequency component with a higher frequency, and a higher intensity noise reduction is performed on the low frequency component with a lower frequency. The voice distortion problem caused by the same intensity noise reduction on the high frequency component and the low frequency component when the voice signal has a distant sound source distance is solved, and the quality of the voice signal is improved.
[0068] In some embodiments, the voice signal is denoised, and in the case that the sound source distance is not less than a preset threshold, the denoising intensity of the high-frequency component is not higher than that of the low-frequency component, which includes:
[0069] The voice signal is noise-estimated to obtain an estimated noise of the voice signal; a noise level of the current environment is determined according to the estimated noise, wherein the noise level includes a first level and a second level, and the noise of the first level is greater than that of the second level; and the denoising intensity of the high-frequency component and the low-frequency component in the voice signal is adjusted according to the noise level.
[0070] The noise estimation can be directly implemented by using a noise estimation algorithm, which can be any existing mature noise estimation algorithm, such as a minimum control recursive average algorithm (MCRA).
[0071] The denoising level is determined by the sound source distance and the noise level, and different denoising strategies are set for frequency domain signals of different frequencies according to different sound source distances and noise levels, thereby improving the robustness of the denoising method.
[0072] Optionally, the noise level of the current environment is determined by the following method:
[0073] The voice signal is frame-processed to obtain a plurality of voice signal frames; the voice signal frames are noise-estimated to obtain an estimated noise of each voice signal frame; it is judged whether the noise energy of continuous M voice signal frames in the voice signal frame exceeds an energy threshold, wherein M is a positive integer; in the case that it is judged that the noise energy of continuous M voice signal frames in the voice signal frame exceeds the energy threshold, the noise level of the current environment is determined as the first level; and in the case that it is judged that the noise energy of continuous M voice signal frames in the voice signal frame does not exceed the energy threshold, the noise level of the current environment is determined as the second level.
[0074] Specifically, the denoising intensity of the high-frequency component and the low-frequency component in the voice signal is adjusted according to the noise level, which includes:
[0075] In the case that the noise level of the current environment is the first level, the high-frequency component of the voice signal is low-intensity denoised, and the low-frequency component of the voice signal is high-intensity denoised.
[0076] In the case that the noise level of the current environment is the second level, the high-frequency component and the low-frequency component of the voice signal are respectively low-intensity denoised.
[0077] In some embodiments, a voice signal denoising method is also provided when the sound source distance is not far, which includes:
[0078] The noise of the voice signal is estimated to obtain an estimated noise of the voice signal; a noise level of a current environment is determined according to the estimated noise, wherein the noise level comprises a first level and a second level, and the noise of the first level is greater than the noise of the second level; in a case that the distance of the sound source is lower than a preset threshold and the noise level of the current environment is the first level, high-intensity noise reduction is performed on the high-frequency component and the low-frequency component of the voice signal respectively.
[0079] Optionally, in a case that the noise level of the current environment is the second level, low-intensity noise reduction is performed on the high-frequency component and the low-frequency component of the voice signal respectively.
[0080] Optionally, the determination of the noise level of the current environment is realized by the following method:
[0081] The voice signal is subjected to frame processing to obtain a plurality of voice signal frames; the noise of each voice signal frame is estimated to obtain an estimated noise of each voice signal frame; it is judged whether the noise energy of continuous M voice signal frames in the voice signal frame exceeds an energy threshold, wherein M is a positive integer; in a case that it is judged that the noise energy of continuous M voice signal frames in the voice signal frame exceeds the energy threshold, it is determined that the noise level of the current environment is the first level; and in a case that it is judged that the noise energy of continuous M voice signal frames in the voice signal frame does not exceed the energy threshold, it is determined that the noise level of the current environment is the second level.
[0082] In some embodiments, the noise reduction comprises low-intensity noise reduction and high-intensity noise reduction, wherein the gain adjustment range of the low-intensity noise reduction is smaller than the gain adjustment range of the high-intensity noise reduction.
[0083] The voice noise reduction can be realized by using all noise reduction algorithms based on Wiener filtering, such as the open source algorithm SPEEX, and such algorithms all need to calculate a noise reduction gain for each frequency point, and the gain range is 0-1, and the noise reduction is completed by multiplying the frequency domain signal by the noise reduction gain.
[0084] Optionally, when the noise reduction is low-intensity noise reduction, the gain of the noise reduction is not processed; and when the noise reduction is high-intensity noise reduction, the noise reduction gain is subjected to exponential power adjustment, such as adjustment to the square value or the cubic value of the gain.
[0085] As shown in FIG. 1, the voice signal is subjected to automatic gain control. Figure 3 As shown in FIG. 2, it is a flow chart of a method for determining the distance of the sound source according to a preferred embodiment of the present application, comprising the following steps:
[0086] S301, the voice signal is subjected to automatic gain control.
[0087] Specifically, the speech signal is segmented into frames to obtain multiple speech signal frames; automatic gain control is applied to the multiple speech signal frames to obtain the gain values of the multiple speech signal frames; the gain values of the current speech signal frame and the previous N speech signal frames are obtained, and the average gain value of the current speech signal frame and the previous N speech signal frames is calculated, where N is a positive integer.
[0088] S302, determine whether the average gain value is greater than the preset threshold. If yes, proceed to step S303; otherwise, proceed to step S304.
[0089] Specifically, when the average gain value of the current speech signal frame and the previous N speech signal frames is greater than a preset threshold, the sound source is considered to be far away from the microphone, i.e., a distant sound source; when the average gain value of the current speech signal frame and the previous N speech signal frames is not greater than the preset threshold, the sound source is considered to be close to the microphone, i.e., a near sound source.
[0090] S303, the current sound source is determined to be a distant sound source.
[0091] S304, the current sound source is determined to be a distant sound source.
[0092] like Figure 4 The diagram shown is a flowchart of a preferred embodiment of the speech denoising method of this application. The following is a brief description of the speech denoising method of the preferred embodiment of this application:
[0093] S401 performs noise estimation on the speech signal to obtain the estimated noise value of the speech signal.
[0094] Specifically, the speech signal is segmented into frames to obtain multiple speech signal frames; noise is estimated for each speech signal frame to obtain the noise energy value of each speech signal frame as the estimated noise value.
[0095] S402, determine whether the noise value is greater than the threshold. If yes, proceed to step S403; otherwise, proceed to step S405.
[0096] Specifically, it is determined whether there are M consecutive speech signal frames in the speech signal frame whose noise energy exceeds the energy threshold, where M is a positive integer; if it is determined that there are M consecutive speech signal frames in the speech signal frame whose noise energy exceeds the energy threshold, the noise level of the current environment is determined to be the first level; and if it is determined that there are no M consecutive speech signal frames in the speech signal frame whose noise energy exceeds the energy threshold, the noise level of the current environment is determined to be the second level, where the noise of the first level is greater than the noise of the second level.
[0097] S403, determine whether it is a distant sound source. If yes, proceed to step S404; otherwise, proceed to step S406.
[0098] Specifically, gain values of a current speech signal frame and N previous speech signal frames are obtained, and average gain values of the current speech signal frame and the N previous speech signal frames are calculated, where N is a positive integer. When the average gain values are greater than a preset threshold, the sound source is a far distance sound source, otherwise, the sound source is a near distance sound source.
[0099] S404, low-level noise reduction is performed on high-frequency components in the speech signal, and high-level noise reduction is performed on low-frequency components in the speech signal.
[0100] S405, low-level noise reduction is performed on the speech signal.
[0101] S406, high-level noise reduction is performed on the speech signal.
[0102] In the preferred embodiment, all noise reduction algorithms based on Wiener filtering are used to implement speech noise reduction, for example, the open source algorithm SPEEX. Such algorithms all need to calculate a noise reduction gain for each frequency point, and the gain range is 0-1. The noise reduction is completed by multiplying the frequency domain signal by the noise reduction gain. When low-intensity noise reduction is performed, the noise reduction gain is not processed. When high-intensity noise reduction is performed, the noise reduction gain is exponentially adjusted, for example, the gain can be adjusted to the square value or the cube value of the gain.
[0103] It is found through research that the related art changes the noise reduction level by changing the related parameters in the calculation process. The main disadvantage of this noise reduction method is that when the noise changes greatly or the human voice is high and low, the noise reduction level will change repeatedly, the background noise after processing is not uniform, and the listening experience is jarring. By changing the parameters in the calculation process to change the noise reduction level, the noise reduction gain needs to be calculated repeatedly.
[0104] In some embodiments of the present application, the gain value obtained by using the automatic gain control method and the estimated noise value are used together to determine the noise reduction level, and different noise reduction strategies are performed on different frequency domain signals for different situations, thereby improving the robustness of the adaptive noise reduction level control system.
[0105] In the embodiment, an electronic device is also provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor implements the speech noise reduction method of the first aspect when executing the computer program.
[0106] Optionally, the electronic device can further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0107] It is noted that the steps shown in the above-described flow or in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0108] In addition, in combination with the voice noise reduction method provided in the above-mentioned embodiments, a storage medium can also be provided in this embodiment to implement. The storage medium has a computer program stored thereon; the computer program is executed by a processor to implement any one of the voice noise reduction methods in the above-mentioned embodiments.
[0109] It should be understood that the specific embodiments described herein are merely illustrative of this application and should not be used to limit its scope. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0110] Obviously, the accompanying drawings are only some examples or embodiments of the present application, and those of ordinary skill in the art can also apply the present application to other similar situations without creative labor. In addition, it can be understood that although the work done in the development process may be complex and long, some design, manufacture or production changes made by those of ordinary skill in the art according to the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application.
[0111] The term "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiments can be included in at least one embodiment of the present application. The presence of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean independence or alternatives to other embodiments. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0112] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A voice noise reduction method, characterized by, The method comprises: obtaining a voice signal, extracting frequency domain components of the voice signal, and dividing the frequency domain components into high frequency components and low frequency components according to frequency size, wherein the frequency of the high frequency components is greater than the frequency of the low frequency components; obtaining the sound source distance of the voice signal; obtaining the sound source distance of the voice signal comprises: performing frame processing on the voice signal to obtain a plurality of voice signal frames; obtaining the gain value of automatic gain control of at least two voice signal frames in the plurality of voice signal frames, and calculating the average gain value of the at least two voice signal frames; determining the sound source distance of the voice signal according to the average gain value of the at least two voice signal frames; de-noising the voice signal, and in the case that the sound source distance is not lower than a preset threshold, the de-noising intensity of the high frequency components is not higher than the de-noising intensity of the low frequency components; de-noising the voice signal, and in the case that the sound source distance is not lower than a preset threshold, the de-noising intensity of the high frequency components is not higher than the de-noising intensity of the low frequency components, which comprises: performing noise estimation on the voice signal to obtain the estimated noise of the voice signal; using the gain value and the estimated noise value together to determine the de-noising level.
2. The voice noise reduction method according to claim 1, characterized in that, obtaining the gain value of at least two voice signal frames in the plurality of voice signal frames, and calculating the average gain value of the at least two voice signal frames comprises: obtaining the gain value of the current voice signal frame and the previous N voice signal frames in the plurality of voice signal frames, and calculating the average gain value of the current voice signal frame and the previous N voice signal frames, wherein N is a positive integer.
3. The voice noise reduction method according to claim 1, wherein, de-noising the voice signal, and in the case that the sound source distance is not lower than a preset threshold, the de-noising intensity of the high frequency components is not higher than the de-noising intensity of the low frequency components, which comprises: performing noise estimation on the voice signal to obtain the estimated noise of the voice signal; determining the noise level of the current environment according to the estimated noise, wherein the noise level comprises a first level and a second level, and the noise of the first level is greater than the noise of the second level; adjusting the de-noising intensity of the high frequency components and the low frequency components in the voice signal according to the noise level.
4. The voice noise reduction method according to claim 3, characterized in that, adjusting the de-noising intensity of the high frequency components and the low frequency components in the voice signal according to the noise level, which comprises: in the case that the noise level of the current environment is the first level, performing low-intensity de-noising on the high frequency components of the voice signal, and performing high-intensity de-noising on the low frequency components of the voice signal.
5. The voice noise reduction method according to claim 3, characterized in that, adjusting the de-noising intensity of the high frequency components and the low frequency components in the voice signal according to the noise level, which comprises: in the case that the noise level of the current environment is the second level, performing low-intensity de-noising on the high frequency components and the low frequency components of the voice signal, respectively.
6. The voice noise reduction method according to claim 1, wherein, The method further comprises: performing noise estimation on the voice signal to obtain the estimated noise of the voice signal; determining the noise level of the current environment according to the estimated noise, wherein the noise level comprises a first level and a second level, and the noise of the first level is greater than the noise of the second level; In a case where the sound source distance is below a preset threshold and the noise level of the current environment is the first level, high-intensity noise reduction is performed on the high-frequency component and the low-frequency component of the speech signal respectively.
7. The voice noise reduction method according to claim 6, characterized in that, The method further comprises: In a case where the noise level of the current environment is the second level, low-intensity noise reduction is performed on the high-frequency component and the low-frequency component of the speech signal respectively.
8. The voice noise reduction method according to claim 3 or 6, characterized by, The noise estimation of the speech signal obtains an estimated noise of the speech signal, and the determination of the noise level of the current environment according to the estimated noise comprises: frame processing of the speech signal to obtain a plurality of speech signal frames; noise estimation of each speech signal frame to obtain an estimated noise of each speech signal frame; determination of whether the noise energy of continuous M speech signal frames in the speech signal frame exceeds an energy threshold, wherein M is a positive integer; In a case where it is determined that the noise energy of continuous M speech signal frames in the speech signal frame exceeds the energy threshold, it is determined that the noise level of the current environment is the first level; and, In a case where it is determined that the noise energy of continuous M speech signal frames in the speech signal frame does not exceed the energy threshold, it is determined that the noise level of the current environment is the second level.
9. The voice noise reduction method of claim 1, wherein, The noise reduction comprises low-intensity noise reduction and high-intensity noise reduction, wherein the gain adjustment range of the low-intensity noise reduction is smaller than the gain adjustment range of the high-intensity noise reduction.
10. The voice noise reduction method of claim 1, wherein, The division of the frequency domain component into a high-frequency component and a low-frequency component according to the frequency size comprises: frame processing of the speech signal to obtain a plurality of speech signal frames, and extraction of frequency domain information of each speech signal frame; According to the frequency domain information, the speech signal with a frequency not less than a frequency threshold is divided into a high-frequency component, and the speech signal with a frequency less than the frequency threshold is divided into a low-frequency component. 11.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to run the computer program to execute the speech noise reduction method of any one of claims 1 to 10.
12. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the speech noise reduction method of any one of claims 1 to 10.
Citation Information
Patent Citations
Audio signal processing method and device and electronic equipment
CN112969130A
User-adaptable hearing aid comprising an initialization module
US20100183177A1