Methods, devices, terminal equipment and storage media for suppressing high-frequency noise
By segmenting and suppressing high-frequency burst noise in voice calls, and using multi-segment suppression curves and frequency band noise reduction coefficients, the interference of burst noise on voice quality is resolved, and high-frequency noise is effectively suppressed while voice quality is guaranteed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING GOERTEK ACOUSTICS TECH CO LTD
- Filing Date
- 2023-03-13
- Publication Date
- 2026-05-26
AI Technical Summary
During voice calls, sudden high-frequency noise can interfere with voice quality, and existing technologies struggle to effectively suppress it without damaging the human voice signal.
When the target audio signal is classified as burst noise, it is segmented using a multi-segment suppression curve. The noise is then suppressed by combining a preset frequency band and a noise reduction coefficient to generate a noise-reduced audio signal.
It effectively suppresses high-frequency burst noise, ensures voice call quality, and reduces interference with human voice signals.
Smart Images

Figure CN116312606B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio processing technology, and particularly relates to a method, apparatus, terminal device, and computer-readable storage medium for suppressing high-frequency noise. Background Technology
[0002] With the continuous development of technology, people's lives are almost inseparable from electronic devices.
[0003] Currently, people have a need and habit of making voice calls at home, on the street, on the subway, or in the office. However, sudden noises are inevitable in these situations, such as car horns on the street, announcements on the subway, and door closing warnings. These sudden noises can affect the voice quality when users make voice calls using electronic devices. Generally speaking, sudden noises or their harmonics are mainly concentrated in the mid-to-high frequency range, while the main characteristic information of human voices is mainly concentrated in the mid-to-low frequency range. Therefore, how to effectively suppress high-frequency sudden noises while ensuring voice quality has become a technical problem that urgently needs to be solved in the field of audio processing technology. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, terminal device, and computer-readable storage medium for suppressing high-frequency noise. It aims to provide a high-frequency noise suppression scheme that effectively suppresses high-frequency burst noise while ensuring voice call quality.
[0005] To achieve the above objectives, the present invention provides a method for suppressing high-frequency noise, the method comprising:
[0006] Acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal;
[0007] When the noise type is burst noise, the target audio signal is subjected to multi-segment suppression processing based on the noise multi-segment suppression curve to obtain a noise-reduced frequency signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve.
[0008] Optionally, the method further includes:
[0009] The preset noise suppression default curve is segmented according to each preset frequency band to obtain the curve for each frequency band;
[0010] Based on the frequency band curves and the noise reduction coefficients corresponding to each preset frequency band, the target frequency band curves are determined.
[0011] The target frequency band curves are combined according to their respective frequency bands to obtain the noise multi-segment suppression target curve.
[0012] Optionally, the step of determining the target frequency band curve based on the frequency band curve and the noise reduction coefficient corresponding to each preset frequency band includes:
[0013] The target frequency band curve is obtained by multiplying the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band.
[0014] Optionally, the higher the preset frequency band, the smaller the noise reduction coefficient corresponding to the preset frequency band.
[0015] Optionally, the step of performing multi-segment suppression processing on the target audio signal based on the noise multi-segment suppression curve to obtain the noise-reduced audio signal includes:
[0016] The noise multi-segment suppression curve and the target audio signal are multiplied to obtain the noise-reduced frequency signal.
[0017] Optionally, the method further includes:
[0018] An audio signal training set is established based on multiple audio signals of known noise types. The initial noise model is trained based on the audio signal training set to obtain the target noise model.
[0019] Optionally, the step of training a pre-constructed initial noise model based on the audio signal training set to obtain a target noise model includes:
[0020] The target audio signal in the audio signal training set is input into a pre-constructed initial noise model;
[0021] The speech features of the target audio signal are extracted by the feature extraction module of the initial noise model, and the noise signal in the target audio signal is separated.
[0022] The initial noise model's spectrum analysis module performs spectrum analysis on the noise signal to determine the frequency range of the noise signal, and determines the noise type of the noise signal based on the frequency range.
[0023] The model parameters of the initial noise model are adjusted according to the noise type to obtain the target noise model.
[0024] Furthermore, to achieve the above objectives, the present invention also provides a high-frequency noise suppression device, the high-frequency noise suppression device comprising:
[0025] The noise type module is used to acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal;
[0026] A multi-segment suppression module is used to perform multi-segment suppression processing on the target audio signal based on a noise multi-segment suppression curve when the noise type is burst noise, to obtain a noise-reduced audio signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve.
[0027] In addition, to achieve the above objectives, the present invention also provides a terminal device, the terminal device comprising: a memory, a processor, and a high-frequency noise suppression program stored in the memory and executable on the processor, wherein the high-frequency noise suppression program of the terminal device, when executed by the processor, implements the steps of the high-frequency noise suppression method as described above.
[0028] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a high-frequency noise suppression program, which, when executed by a processor, implements the steps of the high-frequency noise suppression method described above.
[0029] This invention provides a method, apparatus, terminal device, and computer-readable storage medium for suppressing high-frequency noise. The method involves acquiring a target audio signal, inputting the target audio signal into a target noise model for classification to obtain the noise type of the target audio signal; when the noise type is burst noise, performing multi-segment suppression processing on the target audio signal based on a multi-segment noise suppression curve to obtain a reduced noise frequency signal. The multi-segment noise suppression curve is obtained by segmenting a preset default noise suppression curve.
[0030] This invention provides a high-frequency noise suppression scheme that detects the target audio signal recorded by the headset microphone, inputs the target audio signal into a target noise model for classification, and obtains the noise type of the target audio signal. Then, when the noise type is burst noise, multi-segment suppression processing is performed on the target audio signal based on a multi-segment noise suppression curve to obtain a reduced-noise frequency signal. The multi-segment noise suppression curve is obtained by segmenting a preset default noise suppression curve. Thus, this invention provides a high-frequency noise suppression scheme that segments-suppresses high-frequency burst noise, thereby effectively suppressing high-frequency burst noise while ensuring voice call quality. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the device structure of the terminal device hardware operating environment involved in the embodiments of the present invention;
[0032] Figure 2 This is a schematic flowchart of the first embodiment of the high-frequency noise suppression method of the present invention;
[0033] Figure 3 This is a schematic diagram of the frequency domain noise multi-segment suppression algorithm involved in an embodiment of the high-frequency noise suppression method of the present invention;
[0034] Figure 4 This is a schematic diagram of the noise environment spectrum involved in an embodiment of the high-frequency noise suppression method of the present invention;
[0035] Figure 5 This is a schematic diagram of the noise reduction effect spectrum of an embodiment of the high-frequency noise suppression method of the present invention;
[0036] Figure 6 This is a functional module diagram of an embodiment of the high-frequency noise suppression device of the present invention.
[0037] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0038] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0039] Reference Figure 1 , Figure 1 This is a schematic diagram of the hardware operating environment of the terminal device involved in the embodiment of the present invention.
[0040] It should be noted that the terminal device in this embodiment of the invention can be left or right earphones, or a terminal device integrating a system composed of left and right earphones, applied in the field of earphone audio technology. Specifically, the terminal device can be a smartphone, PC (Personal Computer), tablet computer, portable computer, etc. No specific limitations are imposed here.
[0041] like Figure 1As shown, the terminal device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0042] Those skilled in the art will understand that Figure 1 The terminal device structure shown does not constitute a limitation on the terminal device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0043] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a high-frequency noise suppression program.
[0044] exist Figure 1 In the terminal shown, network interface 1004 is mainly used to connect to the backend server and communicate with it; user interface 1003 is mainly used to connect to the client and communicate with it; and processor 1001 can be used to call the high-frequency noise suppression program stored in memory 1005 and perform the following operations:
[0045] Acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal;
[0046] When the noise type is burst noise, the target audio signal is subjected to multi-segment suppression processing based on the noise multi-segment suppression curve to obtain a noise-reduced frequency signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve.
[0047] Furthermore, the processor 1001 can also be used to call the high-frequency noise suppression program stored in the memory 1005 to perform the following operations:
[0048] The preset noise suppression default curve is segmented according to each preset frequency band to obtain the curve for each frequency band;
[0049] Based on the frequency band curves and the noise reduction coefficients corresponding to each preset frequency band, the target frequency band curves are determined.
[0050] The target frequency band curves are combined according to their respective frequency bands to obtain the noise multi-segment suppression target curve.
[0051] Furthermore, the operation of determining the target frequency band curve based on the frequency band curve and the noise reduction coefficient corresponding to each preset frequency band includes:
[0052] The target frequency band curve is obtained by multiplying the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band.
[0053] Furthermore, the higher the preset frequency band, the smaller the noise reduction coefficient corresponding to the preset frequency band.
[0054] Furthermore, the operation of performing multi-segment suppression processing on the target audio signal based on the noise multi-segment suppression curve to obtain the noise-reduced audio signal includes:
[0055] The noise multi-segment suppression curve and the target audio signal are multiplied to obtain the noise-reduced frequency signal.
[0056] Furthermore, the processor 1001 can also be used to call the high-frequency noise suppression program stored in the memory 1005 to perform the following operations:
[0057] An audio signal training set is established based on multiple audio signals of known noise types. The initial noise model is trained based on the audio signal training set to obtain the target noise model.
[0058] Furthermore, the operation of training the pre-constructed initial noise model based on the audio signal training set to obtain the target noise model includes:
[0059] The target audio signal in the audio signal training set is input into a pre-constructed initial noise model;
[0060] The speech features of the target audio signal are extracted by the feature extraction module of the initial noise model, and the noise signal in the target audio signal is separated.
[0061] The initial noise model's spectrum analysis module performs spectrum analysis on the noise signal to determine the frequency range of the noise signal, and determines the noise type of the noise signal based on the frequency range.
[0062] The model parameters of the initial noise model are adjusted according to the noise type to obtain the target noise model.
[0063] Based on the above structure, various embodiments of the high-frequency noise suppression method are proposed.
[0064] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating a first embodiment of the high-frequency noise suppression method of the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the high-frequency noise suppression method of the present invention may execute the steps shown or described in a different order. In this embodiment, the executing entity of the high-frequency noise suppression method can be a device such as headphones, a personal computer, or a smartphone; this is not limited in this embodiment. For ease of description, the executing entity is omitted from the description of each embodiment. In this embodiment, the high-frequency noise suppression method includes:
[0065] Step S10: Obtain the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal;
[0066] The system acquires the real-time audio signal received by the headphone microphone (hereinafter referred to as the target audio signal for distinction), and inputs the target audio signal into a pre-trained noise model (hereinafter referred to as the target noise model for distinction) for classification. Based on the target noise model, the noise type of the target audio signal is obtained.
[0067] In one feasible implementation, based on the target audio signal obtained by the detected headphone microphone, the target audio signal is input to the target noise model. The target noise model outputs the noise type of the target audio signal according to the input target audio signal. The noise type of the audio signal may be, but is not limited to, wind noise, road noise, and sudden noise. Sudden noise may be, but is not limited to, the door closing alarm sound on the subway.
[0068] Step S20: When the noise type is burst noise, the target audio signal is subjected to multi-segment suppression processing based on the noise multi-segment suppression curve to obtain a noise-reduced frequency signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve.
[0069] When the noise type of the target audio signal is detected as burst noise, multi-segment suppression processing is performed on the target audio signal to generate a processed audio signal (hereinafter referred to as the noise-reduced frequency signal for distinction). The noise multi-segment suppression curve is obtained by segmenting a preset suppression curve (hereinafter referred to as the noise suppression default curve for distinction).
[0070] In one feasible implementation, when the noise type output by the target noise model is detected to be burst noise, it is determined that the detected target audio signal contains burst noise. Then, a Fourier transform is performed on the target audio signal to obtain a frequency domain audio signal. Based on a pre-set noise suppression target curve, multi-segment suppression processing is applied to the corresponding frequency domain audio signal to generate a processed frequency domain noise-reduced signal. According to the inverse Fourier transform principle, the frequency domain noise-reduced signal is converted into a time domain noise-reduced signal. After obtaining the time domain noise-reduced signal, it is played through a headphone speaker. That is, the time domain noise-reduced signal is used in real time to replace the target audio signal with burst noise. It should be noted that the aforementioned time domain noise-reduced signal is the noise-reduced frequency signal.
[0071] It should be noted that the target audio signal mentioned above is a time-domain signal, while the frequency-domain audio signal is a frequency-domain signal. In other words, the target audio signal can be a time-domain audio curve, and the frequency-domain audio signal can be a frequency-domain audio curve. The Fourier transform is primarily used to convert a real time-domain signal into a virtual mathematical structure, namely, the frequency domain.
[0072] In one feasible implementation, such as Figure 4 The diagram shown is a spectrum of a noise environment without any noise reduction treatment. Figure 4 In the upper part, it can be seen that besides the obvious presence of human voices, there is a lot of noise, based on Figure 4 The lower half shows that noise exists from 0Hz to 15kHz, such as Figure 5 As shown, the noise reduction effect spectrum diagram is based on Figure 5 The upper part can be seen Figure 5 The CCP displayed three signal segments, which can be understood as three sentences spoken and / or heard by the user during a voice call. The intervals between the three segments are blank spaces that would exist without speech, which have been noise-reduced to be noise-free. Figure 5 The lower half shows that, in comparison Figure 4 The lower half of the audio signal was significantly weakened, achieving a noise reduction effect.
[0073] Furthermore, in a feasible embodiment, step S20, the step of "performing multi-segment suppression processing on the target audio signal based on the noise multi-segment suppression curve to obtain a noise-reduced audio signal" includes:
[0074] Step S201: Multiply the noise multi-segment suppression curve and the target audio signal to obtain the noise-reduced frequency signal.
[0075] The pre-set noise multi-segment suppression curve and the target audio signal are fitted together to obtain the fitted noise-reduced frequency signal.
[0076] It should be noted that the amplitude range of the default noise suppression curve in this scheme is between 0 and 1.
[0077] In one feasible implementation, the noise multi-segment suppression curve is multiplied with the target audio signal, thereby weakening the frequency domain audio signal and obtaining the noise-reduced audio signal.
[0078] Furthermore, in one feasible embodiment, the high-frequency noise suppression method of the present invention further includes:
[0079] Step A10: Divide the preset noise suppression default curve into segments according to each preset frequency band to obtain the curve for each frequency band;
[0080] The preset noise curve (hereinafter referred to as the noise suppression default curve) is divided into multiple noise curves (hereinafter referred to as frequency band curves) based on the preset frequency band (hereinafter referred to as the preset frequency band for distinction).
[0081] In one feasible implementation, the preset frequency band can be 0Hz to 2KHz, 2KHz to 4KHz, and 4KHz to 8KHz. The default noise suppression curve is determined according to the factory parameters of the headphones. Therefore, by dividing the preset default noise suppression curve according to the preset frequency band, the frequency band curves of the three frequency bands can be obtained.
[0082] It should be noted that this solution does not restrict how the default noise suppression curve is obtained.
[0083] Step A20: Determine the target frequency band curve based on the frequency band curves and the noise reduction coefficients corresponding to each preset frequency band.
[0084] The noise reduction coefficients corresponding to each frequency band curve are determined based on the curves of each frequency band, and the target frequency band curves are calculated based on the curves of each frequency band and the noise reduction coefficients.
[0085] In one feasible implementation, the noise reduction coefficient corresponding to the frequency band curve from 0Hz to 2kHz is set to 1, the noise reduction coefficient corresponding to the frequency band curve from 2kHz to 4kHz is set to 0.005, and the noise reduction coefficient corresponding to the frequency band curve from 4kHz to 8kHz is set to 0.002. Then, based on the pre-set noise reduction coefficients corresponding to each frequency band curve, the target frequency band curves are obtained. It should be understood that the frequency band from 0Hz to 2kHz is mainly distributed with human voice, so it is not advisable to suppress noise too much, which could lead to speech damage. Therefore, the noise reduction coefficient for this frequency band is set to 1. Also, the higher the frequency of the band, the smaller the noise reduction coefficient, so as to achieve greater noise reduction intensity for higher frequency noise.
[0086] It should be noted that the high-frequency noise suppression method of the present invention does not limit the preset frequency band and the number of frequency bands. It should be understood that, based on different design needs of actual applications, in different feasible implementation methods, the preset frequency band can be any range that meets the actual needs, and the number of frequency bands can be any number that meets the actual needs.
[0087] Step A30: Combine the target frequency band curves according to their respective frequency bands to obtain the noise multi-segment suppression target curve.
[0088] The target frequency band curves are combined according to their respective frequency bands. After the combination is completed, the complete multi-segment noise suppression target curve is obtained.
[0089] In one feasible implementation, the three target frequency band curves are ordered and connected according to their respective frequency bands. Specifically, the right endpoint of the first target frequency band curve is connected to the left endpoint of the second target frequency band curve, and the right endpoint of the second target frequency band curve is connected to the left endpoint of the third target frequency band curve, to obtain a complete multi-band noise suppression target curve that includes three frequency bands: 0Hz to 2KHz, 2KHz to 4KHz, and 4KHz to 8KHz.
[0090] In one feasible implementation, such as Figure 3 As shown in the diagram, the frequency domain noise multi-segment suppression algorithm first detects the target audio signal received by the headphone microphone. Then, it uses a noise model to estimate the type of the audio signal and determine the noise type of the target audio signal. When the noise type is burst noise, it performs an FFT (Fast Fourier transform) on the target audio signal. The Fourier transform yields the frequency domain audio signal used for calculation. The frequency domain audio signal is then fitted using a pre-set noise multi-segment suppression target curve, i.e., the two curves are multiplied to obtain the frequency domain audio signal. Finally, the frequency domain audio signal is subjected to an IFFT (Inverse Fast Fourier Transform) to convert the frequency domain audio signal into a time domain audio signal, which is then played through the headphone speaker, thus completing the noise reduction of burst noise.
[0091] Further, in one feasible embodiment, step A20 includes:
[0092] Step A201: Multiply the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band to obtain the target frequency band curve.
[0093] The target frequency band curve is obtained by multiplying the curve of each frequency band by the noise reduction coefficient corresponding to each frequency band curve.
[0094] In one feasible implementation, the frequency band curve from 0Hz to 2kHz is multiplied by a noise reduction coefficient of 1 to obtain the first target frequency band curve; the frequency band curve from 2kHz to 4kHz is multiplied by a noise reduction coefficient of 0.005 to obtain the second target frequency band curve; and the frequency band curve from 4kHz to 8kHz is multiplied by a noise reduction coefficient of 0.002 to obtain the third target frequency band curve.
[0095] Furthermore, the higher the frequency band, the smaller the noise reduction coefficient corresponding to that preset frequency band.
[0096] In one feasible implementation, the higher the frequency band of the audio signal, the smaller the preset noise reduction coefficient of that frequency band, and the greater the noise reduction intensity.
[0097] In this embodiment, the high-frequency noise suppression method of the present invention acquires the real-time target audio signal received by the headphone microphone and inputs the target audio signal into a pre-trained target noise model for classification, and obtains the noise type of the target audio signal based on the target noise model; when the noise type of the target audio signal is detected to be burst noise, multi-segment suppression processing is performed on the target audio signal to generate a processed noise-reduced frequency signal, wherein the noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve; the preset noise multi-segment suppression curve and the target audio signal are curve-fitted to obtain the fitted noise-reduced frequency signal; the preset noise suppression default curve is segmented according to a preset preset frequency band to obtain multiple segmented frequency band curves; the noise reduction coefficient corresponding to each frequency band curve is determined according to each frequency band curve, and each target frequency band curve is calculated based on each frequency band curve and each noise reduction coefficient; the target frequency band curves are combined according to their respective frequency bands, and after combination, the completed noise multi-segment suppression target curve is obtained; each frequency band curve is multiplied by its respective noise reduction coefficient to obtain each target frequency band curve.
[0098] Thus, in this embodiment of the invention, the target audio signal recorded by the headset microphone is detected, and the target audio signal is input into a target noise model for classification to obtain the noise type of the target audio signal. Then, when the noise type is burst noise, multi-segment suppression processing is performed on the target audio signal based on a noise multi-segment suppression curve to obtain a noise-reduced frequency signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve. Therefore, this invention provides a high-frequency noise suppression scheme that segments-suppresses high-frequency burst noise, thereby effectively suppressing high-frequency burst noise while ensuring voice call quality.
[0099] Furthermore, based on the first embodiment of the high-frequency noise suppression method of the present invention described above, a second embodiment of the high-frequency noise suppression method of the present invention is proposed.
[0100] In this embodiment, the high-frequency noise suppression method of the present invention further includes:
[0101] Step B10: Establish an audio signal training set based on multiple audio signals of known noise types, and train the pre-constructed initial noise model based on the audio signal training set to obtain the target noise model.
[0102] An audio signal training set is established, and the pre-constructed noise model (hereinafter referred to as the initial noise model for distinction) is trained based on the audio signal training set to obtain the trained target noise model.
[0103] In one feasible implementation, the audio signal training set consists of multiple audio signals, including but not limited to various types of audio signals such as wind noise, road noise, and burst noise. Based on the audio signal training set, a pre-constructed initial noise model is trained. It should be noted that the initial noise model is an untrained model structure. After the model training is completed, a target noise model is obtained, which is used to output the noise type of the input audio signal.
[0104] Furthermore, in a feasible embodiment, step B10 above, the step of "training the pre-constructed initial noise model based on the audio signal training set to obtain the target noise model" includes:
[0105] Step B101: Input the target audio signal from the audio signal training set into the pre-constructed initial noise model;
[0106] The audio signals from the audio signal training set (hereinafter referred to as the target audio signals for distinction) are input into the pre-built initial noise model.
[0107] It should be noted that the high-frequency noise suppression method of the present invention does not limit how the spectral energy of the audio signal is specifically determined.
[0108] Step B102: Extract speech features of the target audio signal through the feature extraction module of the initial noise model, and separate the noise signal in the target audio signal;
[0109] The speech features of the target audio signal are extracted by the feature extraction module of the initial noise model, and the noise signal of the target audio signal is separated.
[0110] In one feasible implementation, by extracting speech features from the target audio signal, the speech signal in the target audio signal is determined, thereby separating the noise signal in the target audio signal other than the speech signal.
[0111] Step B103: Perform spectrum analysis on the noise signal using the spectrum analysis module of the initial noise model to determine the spectrum range of the noise signal, and determine the noise type of the noise signal based on the spectrum range;
[0112] The spectrum analysis module of the initial noise model determines the spectrum energy range of the target audio signal, and detects whether the spectrum energy range corresponding to the spectrum energy of the target audio signal is within a preset spectrum energy range (hereinafter referred to as the preset spectrum energy range for distinction). If the spectrum energy range corresponding to the spectrum energy of the target audio signal is within the preset spectrum energy range, the noise type of the target audio signal is determined to be burst noise.
[0113] In one feasible implementation, the initial noise model has preset spectral energy ranges. The spectral energy range in which the spectral energy of the high-frequency audio signal is located is taken as the preset spectral energy range. The high-frequency audio signal is the burst audio signal. Then, when the spectral energy range corresponding to the spectral energy of the target audio signal is the preset spectral energy range, the noise type of the target audio signal is determined to be burst noise.
[0114] Step B104: Adjust the model parameters of the initial noise model according to the noise type to obtain the target noise model.
[0115] The model parameters of the initial noise model are adjusted based on the noise type obtained from model training to obtain the target noise model.
[0116] In this embodiment, the high-frequency noise suppression method of the present invention establishes an audio signal training set and trains a pre-constructed initial noise model based on the audio signal training set to obtain a trained target noise model; the target audio signal in the audio signal training set is input into the pre-constructed initial noise model; the speech features of the target audio signal are extracted by the feature extraction module of the initial noise model to separate the noise signal of the target audio signal; the spectral analysis module of the initial noise model determines the spectral energy range of the target audio signal, detects whether the spectral energy range corresponding to the spectral energy of the target audio signal is a preset spectral energy range, and determines that the noise type of the target audio signal is burst noise when the spectral energy range corresponding to the spectral energy of the target audio signal is a preset spectral energy range; the model parameters of the initial noise model are adjusted according to the noise type obtained by model training to obtain the target noise model.
[0117] Thus, by using a pre-built noise model, it is possible to detect whether the target audio signal collected by the headphone microphone needs to undergo multi-segment noise suppression processing, thereby establishing the key judgment conditions for multi-segment noise suppression processing.
[0118] Furthermore, embodiments of the present invention also provide a high-frequency noise suppression device.
[0119] Please refer to Figure 6 , Figure 6 This is a functional module diagram of an embodiment of the high-frequency noise suppression device of the present invention, as shown below. Figure 6 As shown, the high-frequency noise suppression device of the present invention includes:
[0120] The noise type module 10 is used to acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal;
[0121] The multi-segment suppression module 20 is used to perform multi-segment suppression processing on the target audio signal based on the noise multi-segment suppression curve when the noise type is burst noise, so as to obtain a noise-reduced frequency signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve.
[0122] Furthermore, the high-frequency noise suppression device of the present invention further includes:
[0123] The segmentation module is used to segment the preset noise suppression default curve according to each preset frequency band to obtain the curve for each frequency band.
[0124] The target frequency band curve module is used to determine each target frequency band curve based on each frequency band curve and the noise reduction coefficient corresponding to each preset frequency band.
[0125] The combination module is used to combine the target frequency band curves according to their respective frequency bands to obtain the noise multi-segment suppression target curve.
[0126] Furthermore, the target frequency band curve module includes:
[0127] The target frequency band curve unit is used to multiply the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band to obtain the target frequency band curve.
[0128] Furthermore, the higher the preset frequency band, the lower the noise reduction coefficient corresponding to the preset frequency band.
[0129] Furthermore, the multi-segment suppression module 20 also includes:
[0130] The noise reduction frequency signal module is used to multiply the noise multi-segment suppression curve and the target audio signal to obtain the noise reduction frequency signal.
[0131] Furthermore, the high-frequency noise suppression device of the present invention further includes:
[0132] The model training module is used to establish an audio signal training set based on multiple audio signals of known noise types, and to train a pre-constructed initial noise model based on the audio signal training set to obtain the target noise model.
[0133] Furthermore, the model training module is also used to input the target audio signal from the audio signal training set into a pre-constructed initial noise model; extract the speech features of the target audio signal through the feature extraction module of the initial noise model, and separate the noise signal in the target audio signal; perform spectral analysis on the noise signal through the spectral analysis module of the initial noise model to determine the spectral range of the noise signal, and determine the noise type of the noise signal based on the spectral range; adjust the model parameters of the initial noise model according to the noise type to obtain the target noise model.
[0134] Furthermore, the target noise model unit is also used to detect whether the spectral energy range is a preset spectral energy range; when the spectral energy range is the preset spectral energy range, the noise type of the target audio signal is determined to be burst noise.
[0135] The present invention also provides a computer storage medium storing a high-frequency noise suppression program, which, when executed by a processor, implements the steps of the high-frequency noise suppression program method as described in any of the above embodiments.
[0136] The specific embodiments of the computer storage medium of the present invention are basically the same as the embodiments of the high-frequency noise suppression program method of the present invention described above, and will not be repeated here.
[0137] The present invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the high-frequency noise suppression method of the present invention as described in any of the above embodiments, which will not be elaborated here.
[0138] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0139] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (such as TWS earphones, etc.) to execute the methods described in the various embodiments of the present invention.
[0141] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for suppressing high-frequency noise, characterized in that, The high-frequency noise suppression method includes the following steps: Acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal; When the noise type is burst noise, the target audio signal is subjected to multi-segment suppression processing based on the noise multi-segment suppression curve to obtain a noise-reduced audio signal. The noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve. The method further includes: The preset noise suppression default curve is segmented according to each preset frequency band to obtain the curve for each frequency band; Multiply the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band to obtain the target frequency band curve; The target frequency band curves are combined according to their respective frequency bands to obtain a multi-segment noise suppression curve; The method further includes: An audio signal training set is established based on multiple audio signals with known noise types; The target audio signal in the audio signal training set is input into a pre-constructed initial noise model; The speech features of the target audio signal are extracted by the feature extraction module of the initial noise model, and the noise signal in the target audio signal is separated. The initial noise model's spectrum analysis module performs spectrum analysis on the noise signal to determine the frequency range of the noise signal, and determines the noise type of the noise signal based on the frequency range. The model parameters of the initial noise model are adjusted according to the noise type to obtain the target noise model.
2. The high-frequency noise suppression method as described in claim 1, characterized in that, The higher the preset frequency band, the smaller the noise reduction coefficient corresponding to the preset frequency band.
3. The high-frequency noise suppression method as described in claim 1, characterized in that, The step of performing multi-segment suppression processing on the target audio signal based on the multi-segment noise suppression curve to obtain the noise-reduced audio signal includes: The noise multi-segment suppression curve and the target audio signal are multiplied to obtain the noise-reduced frequency signal.
4. A high-frequency noise suppression device, characterized in that, The high-frequency noise suppression device includes: The noise type module is used to acquire the target audio signal, input the target audio signal into the target noise model for classification, and obtain the noise type of the target audio signal; A multi-segment suppression module is used to perform multi-segment suppression processing on the target audio signal based on a noise multi-segment suppression curve when the noise type is burst noise, to obtain a noise-reduced audio signal, wherein the noise multi-segment suppression curve is obtained by segmenting a preset noise suppression default curve; The multi-segment suppression module is further configured to segment the preset noise suppression default curve according to each preset frequency band to obtain each frequency band curve; multiply the frequency band curve corresponding to the preset frequency band and the noise reduction coefficient corresponding to the preset frequency band to obtain the target frequency band curve; and combine each target frequency band curve according to its corresponding frequency band to obtain the noise multi-segment suppression curve. The high-frequency noise suppression device further includes a model training module, which is used to establish an audio signal training set based on multiple audio signals of known noise types; input the target audio signal in the audio signal training set into a pre-constructed initial noise model; extract the speech features of the target audio signal through the feature extraction module of the initial noise model, and separate the noise signal in the target audio signal; perform spectral analysis on the noise signal through the spectral analysis module of the initial noise model to determine the spectral range of the noise signal, and determine the noise type of the noise signal based on the spectral range; adjust the model parameters of the initial noise model according to the noise type to obtain the target noise model.
5. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and a high-frequency noise suppression program stored in the memory and executable on the processor. When the high-frequency noise suppression program is executed by the processor, it implements the steps of the high-frequency noise suppression method as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a high-frequency noise suppression program, which, when executed by a processor, implements the steps of the high-frequency noise suppression method as described in any one of claims 1 to 3.