A method, apparatus, encoding method, medium and device for detecting and eliminating noise

CN115762547BActive Publication Date: 2026-08-18BEIJING BAIRUI INTERNET TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211212786.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-08-18
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

[0003]针对在进行Hiss噪声的检测和消除时,流程复杂,功耗高,高延迟的问题,本申请提出一种检测和消除噪声的方法、装置、编码方法、介质及设备

Benefits of technology

[0014] The noise detection and elimination method of this application encodes audio data by utilizing the existing encoding process in the encoder, and at the same time uses pseudospectral flatness to set the noise update speed parameter to perform the transition between multiple frames. It has the characteristics of low power consumption and low latency, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762547B_ABST
    Figure CN115762547B_ABST
Patent Text Reader

Abstract

The application discloses a method, device, encoding method, medium and equipment for detecting and eliminating noise, and belongs to the technical field of Bluetooth. The method comprises the following steps: obtaining original frequency domain spectrum coefficients corresponding to current frame audio in the encoding or decoding process of an audio in a codec with an improved discrete cosine transform process; calculating pseudo-spectrum flatness according to the original frequency domain spectrum coefficients; setting a noise update speed parameter according to the pseudo-spectrum flatness and calculating noise energy of the current frame audio under the condition that the pseudo-spectrum flatness is greater than a preset threshold; and calculating updated spectrum coefficients of the current frame audio according to the original frequency domain spectrum coefficients and the noise energy, and performing a subsequent encoding or decoding process according to the updated spectrum coefficients. The application encodes audio data by using an existing encoding process in an encoder, sets a noise update speed parameter by using pseudo-spectrum flatness, performs transition between multiple frames, reduces power consumption and delay, and improves user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Bluetooth technology, and in particular to a method, apparatus, encoding method, medium and device for detecting and eliminating noise. Background Technology

[0002] Hiss noise is a stationary additive noise across the entire frequency range, producing a hissing sound. It's often introduced during the digitization process of older records, and it's also common for broadcasters to experience during live streams due to recording equipment limitations, negatively impacting the user experience. Current Hiss noise suppression techniques typically involve the following steps: 1. Inputting the audio signal (PCM format); 2. Windowing; 3. Fourier transform; 4. Identifying frame type; 5. Hiss noise estimation and update; 6. Spectral gain; 7. Spectral smoothing; 8. Multiplication; 9. Inverse Fourier transform; 10. Synthesis window; 11. Overlapping addition; 12. Outputting the Hiss noise-suppressed audio signal. In current Bluetooth applications, there are requirements for high sound quality, low power consumption, and low latency. Firstly, the aforementioned processing steps are numerous and complex, increasing the power consumption of Bluetooth devices. Secondly, while the overlapping addition method is used to ensure smooth inter-frame signals, this increases latency, failing to meet Bluetooth's low-latency goals and requirements, thus affecting the user experience. Summary of the Invention

[0003] To address the issues of complex processes, high power consumption, and high latency in the detection and elimination of Hiss noise, this application proposes a method, apparatus, encoding method, medium, and device for detecting and eliminating noise.

[0004] In a first aspect, this application proposes a method for detecting and eliminating noise, comprising: obtaining the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding or decoding process of an audio encoder or decoder with improved discrete cosine transform processing; calculating the pseudo-spectral flatness of the current frame audio based on the original frequency domain spectral coefficients; setting a noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold; calculating the noise energy of the current frame audio based on the noise update rate parameter; and calculating the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy, and performing subsequent encoding or decoding processes based on the updated spectral coefficients.

[0005] Optionally, during the encoding or decoding of audio by a codec with improved discrete cosine transform processing, the original frequency domain spectral coefficients corresponding to the current frame audio are obtained, including: during the standard encoding or decoding process of the codec, performing an improved discrete cosine transform on the current frame audio to obtain the original frequency domain spectral coefficients.

[0006] Optionally, the pseudospectral flatness of the current frame audio is calculated based on the original frequency domain spectral coefficients, including: calculating the pseudospectral corresponding to the current frame audio based on the original frequency domain spectral coefficients; calculating and obtaining the pseudospectral geometric mean and pseudospectral arithmetic mean based on the pseudospectral; and calculating the pseudospectral flatness based on the pseudospectral geometric mean and pseudospectral arithmetic mean.

[0007] Optionally, when the pseudo-spectral flatness is greater than a preset threshold, a noise update rate parameter is set based on the pseudo-spectral flatness, including: setting the noise update rate parameter based on the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness.

[0008] Optionally, under the condition that the pseudo-spectral flatness is greater than a preset threshold, the noise update speed parameter is set according to the pseudo-spectral flatness, which further includes: under the condition that the pseudo-spectral flatness is greater than a preset threshold, determining whether there is an audio frame in the preset number of audio frames before the current frame audio that has a pseudo-spectral flatness less than a preset threshold; if not, then setting the noise update speed parameter according to the magnitude of the pseudo-spectral flatness corresponding to the current frame audio.

[0009] Optionally, the noise energy of the current frame audio is calculated based on the noise update rate parameter, including: selecting the non-voice band of the current frame audio and calculating the pseudo-spectral median corresponding to the current frame audio; and calculating the current frame moving average energy corresponding to the current frame audio based on the pseudo-spectral median, the noise update rate parameter, and the moving average energy of the previous frame audio, and using it as the noise energy of the current frame audio.

[0010] Secondly, this application proposes an apparatus for detecting and eliminating noise, characterized by comprising: a frequency domain spectral coefficient acquisition module, which acquires the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding or decoding process of an audio encoder or decoder with improved discrete cosine transform processing; a pseudo-spectral flatness calculation module, which calculates the pseudo-spectral flatness of the current frame audio based on the original frequency domain spectral coefficients; a noise update rate parameter determination module, which sets a noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold; a noise energy calculation module, which calculates the noise energy of the current frame audio based on the noise update rate parameter; and a spectral coefficient update module, which calculates the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy, and performs subsequent encoding or decoding processes based on the updated spectral coefficients.

[0011] Thirdly, this application proposes an audio coding method, characterized by comprising: obtaining the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding process of an encoder with improved discrete cosine transform processing; calculating the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients; setting a noise update rate parameter based on the pseudospectral flatness when the pseudospectral flatness is greater than a preset threshold; calculating the noise energy of the current frame audio based on the noise update rate parameter; calculating the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy; updating other coding modules in the encoder based on the updated spectral coefficients, and continuing to encode the current frame audio using the updated coding modules.

[0012] Fourthly, this application provides a computer-readable storage medium storing a computer program, wherein the computer program is operated to perform the noise detection and cancellation method in Scheme 1 or the audio coding method in Scheme 3.

[0013] Fifthly, this application provides a computer device including a processor and a memory, the memory storing a computer program, wherein the processor operates the computer program to perform the noise detection and cancellation method in Scheme 1 or the audio coding method in Scheme 3.

[0014] The noise detection and elimination method of this application encodes audio data by utilizing the existing encoding process in the encoder, and at the same time uses pseudospectral flatness to set the noise update speed parameter to perform the transition between multiple frames. It has the characteristics of low power consumption and low latency, and improves the user experience. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description exemplarily illustrate some embodiments of this application.

[0016] Figure 1 A schematic diagram of one embodiment of the method for detecting and eliminating noise according to this application is shown;

[0017] Figure 2 A schematic diagram of a noise segment and its corresponding noise flatness is shown in this application;

[0018] Figure 3 A schematic diagram of a human voice segment and its corresponding spectral flatness is shown in this application;

[0019] Figure 4 A schematic diagram of one embodiment of the noise detection and elimination apparatus of this application is shown;

[0020] Figure 5A schematic diagram of one embodiment of the audio encoding method of this application is shown.

[0021] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0022] The preferred embodiments of this application will now be described in detail with reference to the accompanying drawings, so that the advantages and features of this application can be more easily understood by those skilled in the art, thereby providing a clearer and more definite definition of the scope of protection of this application.

[0023] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0024] Hiss noise is a stationary additive noise across the entire frequency range, producing a hissing sound. It's often introduced during the digitization process of older records, and it's also common for broadcasters to experience during live streams due to recording equipment limitations, negatively impacting the user experience. Current Hiss noise suppression techniques typically involve the following steps: 1. Inputting the audio signal (PCM format); 2. Windowing; 3. Fourier transform; 4. Identifying frame type; 5. Hiss noise estimation and update; 6. Spectral gain; 7. Spectral smoothing; 8. Multiplication; 9. Inverse Fourier transform; 10. Synthesis window; 11. Overlapping addition; 12. Outputting the Hiss noise-suppressed audio signal. In current Bluetooth applications, there are requirements for high sound quality, low power consumption, and low latency. Firstly, the aforementioned processing steps are numerous and complex, increasing the power consumption of Bluetooth devices. Secondly, while the overlapping addition method is used to ensure smooth inter-frame signals, this increases latency, failing to meet Bluetooth's low-latency goals and requirements, thus affecting the user experience.

[0025] To address the aforementioned problems, this application proposes a method, apparatus, encoding method, medium, and device for detecting and eliminating noise. The method includes: obtaining the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding or decoding process of an audio encoder / decoder with improved discrete cosine transform processing; calculating the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients; setting a noise update rate parameter based on the pseudospectral flatness when the pseudospectral flatness exceeds a preset threshold; calculating the noise energy of the current frame audio based on the noise update rate parameter; and calculating the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy, and performing subsequent encoding or decoding processes based on the updated spectral coefficients. The technical solution of this application mainly targets Hiss noise. There are many types of noise, including stationary and non-stationary, additive and multiplicative noise. This application is also effective for additive stationary noise.

[0026] The noise detection and elimination method of this application encodes or decodes audio using the existing encoding or decoding process in a codec with improved discrete cosine transform processing, obtains the original frequency domain spectral coefficients, and obtains pseudo-spectral flatness. At the same time, it uses pseudo-spectral flatness to set the noise update speed parameter and perform transitions between multiple frames. It features low power consumption and low latency, improving the user experience.

[0027] The technical solutions of this application and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. The specific embodiments described below can be combined with each other to form new embodiments. The same or similar ideas or processes described in one embodiment may not be repeated in other embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0028] Figure 1 A schematic diagram of one embodiment of the method for detecting and eliminating noise according to this application is shown.

[0029] exist Figure 1 In the embodiment shown, the method for detecting and eliminating noise in this application includes process S101, which involves obtaining the original frequency domain spectral coefficients corresponding to the audio of the current frame during the encoding or decoding process of an audio encoder or decoder with improved discrete cosine transform processing.

[0030] In this embodiment, the method for detecting and eliminating noise in this application utilizes the existing standard encoding or decoding process for audio in the codec, and obtains the original frequency domain spectral coefficients of the current frame audio during the encoding or decoding process through improved discrete cosine transform processing, and then uses the original frequency domain spectral coefficients to detect subsequent noise.

[0031] Optionally, during the encoding or decoding of audio by a codec with improved discrete cosine transform processing, the original frequency domain spectral coefficients corresponding to the current frame audio are obtained, including: during the standard encoding or decoding process of the codec, performing an improved discrete cosine transform on the current frame audio to obtain the original frequency domain spectral coefficients.

[0032] In this optional embodiment, for ease of description, this application uses the LC3 codec's audio encoding process as an example to illustrate the technical solution of this application for codecs with improved discrete cosine transform processing. The processing procedures for other codecs with improved discrete cosine transform processing are similar.

[0033] This paper utilizes the existing improved discrete cosine transform module in the LC3 encoder to encode the audio of the current frame, obtaining the original frequency domain spectral coefficients corresponding to the current frame audio. These original frequency domain spectral coefficients are then used to detect Hiss noise in the audio data. As mentioned in the background section, existing technologies for Hiss noise removal include processes such as analysis windowing, Fourier transform, and inverse Fourier transform, resulting in high power consumption, high latency, and a poor user experience. In the method of this application, since the existing encoding module in the encoder is used, the above processing steps can be omitted, thereby reducing power consumption and lowering the requirements for storage space and computing power of the entire encoding system. Furthermore, in existing solutions, processes such as Fourier transform and inverse transform, synthesis windowing, and analysis windowing introduce certain errors. This application omits these modules, thus avoiding the impact of high latency on sound quality. Since this application does not introduce algorithmic latency, the user experience is improved. In addition, the technical solution of this application has lower computing power requirements, ensuring the battery life of the Bluetooth embedded device.

[0034] Specifically, taking the encoding of 10ms frame-long, 48kHz sampling rate audio data by the LC3 encoder to obtain frequency domain spectral coefficients as an example, the process of obtaining the original frequency domain spectral coefficients of the audio is explained. According to the standard encoding process of the LC3 encoder, the input 10ms frame-long audio data is subjected to low-latency improved discrete cosine transform (LD-MDCT) calculation to obtain the corresponding original frequency domain spectral coefficients. Wherein, the audio data of the current frame, n = 0, 1, x... s (n)2,…,N F ,

[0035] t(n) = x s (ZN F +n), for n=0…2·N F -1-Z

[0036] t(2N F -Z+n)=0,for n=0…Z-1

[0037]

[0038] In the above formula, based on the LC3 standard specification, N F It is 480, Z is 180, w Nms_NF (n) is the low-latency MDCT window, and X(k) is the temporal audio data x of the current frame. s (n) corresponds to the frequency domain spectral coefficients.

[0039] exist Figure 1 In the embodiment shown, the method for detecting and eliminating noise in this application includes process S102, which calculates the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients.

[0040] Optionally, the pseudospectral flatness of the current frame audio data is calculated based on the original frequency domain spectral coefficients, including: calculating the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients, including: calculating the pseudospectral corresponding to the current frame audio based on the original frequency domain spectral coefficients; calculating and obtaining the pseudospectral geometric mean and pseudospectral arithmetic mean based on the pseudospectral; and calculating the pseudospectral flatness based on the pseudospectral geometric mean and pseudospectral arithmetic mean.

[0041] Specifically, the process of calculating the pseudospectral flatness of the current frame audio data based on the original frequency domain spectral coefficients is as follows:

[0042] Calculate the pseudospectrum:

[0043]

[0044] Where X(k) = 0, when k = -1 or N F hour

[0045] Calculate the geometric mean of the pseudospectral:

[0046]

[0047] Calculate the arithmetic mean of the pseudospectral:

[0048]

[0049] Calculate pseudospectral flatness:

[0050]

[0051] exist Figure 1 In the embodiment shown, the method for detecting and eliminating noise in this application includes process S103, which sets a noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold.

[0052] In this embodiment, when performing audio noise cancellation, since noise data within audio frames is being removed, the degree of noise cancellation may differ between consecutive frames, causing significant fluctuations in the originally smooth audio and thus degrading the user experience. To ensure smooth transitions between multiple frames after noise processing, this application sets a noise update speed parameter during the noise cancellation process to adjust the degree of noise cancellation between different audio frames, thereby ensuring audio quality and effectively eliminating noise.

[0053] Specifically, a preset threshold is set to 0.15. This threshold is used to distinguish between a normal audio frame (tone frame) and a voiced frame (voiced frame) containing noise. If the pseudospectral flatness is less than this threshold, the current frame is a tone signal and therefore a tone frame, requiring no noise removal. If the pseudospectral flatness is greater than or equal to the preset threshold, the audio frame is a voiced frame and requires noise removal.

[0054] Optionally, when the pseudo-spectral flatness is greater than a preset threshold, a noise update rate parameter is set based on the pseudo-spectral flatness, including: setting the noise update rate parameter based on the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness.

[0055] In this optional embodiment, the noise update rate parameter is set reasonably based on the magnitude of the pseudo-spectral flatness when determining the noise update rate parameter. The magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness. The larger the pseudo-spectral flatness of the current frame audio data, the larger the corresponding noise update rate parameter is set; conversely, the smaller the pseudo-spectral flatness of the current frame audio data, the smaller the corresponding noise update rate parameter is set.

[0056] Optionally, under the condition that the pseudo-spectral flatness is greater than a preset threshold, the noise update speed parameter is set according to the pseudo-spectral flatness, which further includes: under the condition that the pseudo-spectral flatness is greater than a preset threshold, determining whether there is an audio frame in the preset number of audio frames before the current frame audio that has a pseudo-spectral flatness less than a preset threshold; if not, then setting the noise update speed parameter according to the magnitude of the pseudo-spectral flatness corresponding to the current frame audio.

[0057] In this optional embodiment, when the pseudo-spectral flatness of the current frame audio is determined to be greater than or equal to the preset threshold, indicating a voiced frame containing noise, specifically for speech, the unvoiced sounds following a voiced frame generally have a spectral flatness similar to noise. To make the noise processing more accurate, a pitch frame delay factor is set. If the current frame is a pitch frame, then the subsequent frames are all considered pitch frames, regardless of their spectral flatness. If the current frame is a pitch frame, it means there is no Hiss noise or its energy is very small. In this case, spectral coefficient updates are not performed to avoid damage to sound quality, but noise calculations are still updated to facilitate a more reasonable noise estimate when spectral coefficients need to be updated in a future frame. Specifically, the preset number of frames can be set to 5 frames. It should be noted that the selection of the preset threshold and the preset number of frames can be reasonably adjusted according to actual processing requirements, and this application does not impose specific limitations.

[0058] In addition, the noise update speed α is set according to the spectral flatness. That is, the lower the spectral flatness, the stronger the tone component, and the slower the noise update speed, and vice versa. Specifically, the update speed α can be set to 0.05 when the spectral flatness is 0.1, and 0.25 when the spectral flatness is 0.8. The linear relationship can be obtained as follows, and other values ​​can be obtained according to the following formula.

[0059]

[0060] Figure 2 A schematic diagram of a noise segment and its corresponding noise flatness is shown in this application. Figure 3 A schematic diagram of a human voice segment and its corresponding spectral flatness is shown in this application.

[0061] exist Figure 1 In the embodiment shown, the method for detecting and eliminating noise in this application includes process S104, which calculates the noise energy of the current frame audio based on the noise update speed parameter.

[0062] Optionally, the noise energy of the current frame audio is calculated based on the noise update rate parameter, including: selecting the non-voice band of the current frame audio and calculating the pseudo-spectral median corresponding to the current frame audio; and calculating the current frame moving average energy corresponding to the current frame audio based on the pseudo-spectral median, the noise update rate parameter, and the moving average energy of the previous frame audio, and using it as the noise energy of the current frame audio.

[0063] Specifically, a non-voice frequency band is selected and the pseudo-spectral median is calculated. Typically, the main energy of the voice frequency band is concentrated in the range of 300Hz to 3400Hz. Taking the current sampling rate of 48Hz as an example, its effective bandwidth is 24kHz. The frequency band of 8kHz to 20kHz can be selected to calculate and determine Hiss noise. The corresponding spectral coefficient index k ranges from 160 to 400. In practice, the selection of k mainly avoids the voice frequency band and high-frequency attenuation, which is mainly caused by the microphone and A / D converter, etc. It can be measured in detail. This invention has no limitations.

[0064] Non-speech band pseudospectral median of the current frame:

[0065] X pseudo-Med =median(X) pseudo (k)), k = 160~400, the Median() operation means taking the median value.

[0066] The moving average energy of the current frame is:

[0067] P ma =α*P ma-la +(1-α)*X pseudo-Me , where P ma-last It is the moving average energy of the previous frame, and α is the moving average factor, which is used to control the speed of noise updates and avoid abrupt changes in noise estimation between frames. A typical value is 0.95.

[0068] exist Figure 1 In the embodiment shown, the method for detecting and eliminating noise in this application includes process S105, which calculates the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and noise energy, and performs subsequent encoding or decoding processes based on the updated spectral coefficients.

[0069] Specifically, as mentioned before, when the spectral flatness is greater than 0.15, it indicates that there is noise in the current frame of audio, and noise processing begins by updating the spectral coefficients to eliminate the noise.

[0070]

[0071] The spectral coefficients are updated as shown in the formula above. After obtaining the new spectral coefficients, the encoder continues the encoding process of the remaining modules. The remaining encoding modules are completed according to the LC3 standard specification, including transform domain noise shaping, time domain noise shaping, quantization and noise level estimation, arithmetic and residual coding, and bitstream encapsulation.

[0072] The noise detection and cancellation method of this application can be applied to the encoding process of an encoder or the decoding process of a decoder with improved discrete cosine transform processing to detect and cancel Hiss noise. Furthermore, the noise detection and cancellation method of this application encodes audio data using the existing encoding process in the encoder, while simultaneously using pseudospectral flatness to set the noise update rate parameter for transitions between multiple frames. This method features low power consumption and low latency, improving the user experience.

[0073] Figure 4 An embodiment of the noise detection and elimination apparatus of this application is shown.

[0074] exist Figure 4 In the illustrated embodiment, the noise detection and elimination apparatus of this application includes: a frequency domain spectral coefficient acquisition module 401, which acquires the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding or decoding process of an audio encoder or decoder with improved discrete cosine transform processing; a pseudo-spectral flatness calculation module 402, which calculates the pseudo-spectral flatness of the current frame audio based on the original frequency domain spectral coefficients; a noise update speed parameter determination module 403, which sets a noise update speed parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold; a noise energy calculation module 404, which calculates the noise energy of the current frame audio based on the noise update speed parameter; and a spectral coefficient update module 405, which calculates the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy, and performs subsequent encoding or decoding processes based on the updated spectral coefficients.

[0075] Optionally, in the frequency domain spectral coefficient acquisition module 401, during the standard encoding or decoding process of the codec, an improved discrete cosine transform is performed on the audio of the current frame to obtain the original frequency domain spectral coefficients.

[0076] Optionally, in the pseudospectral flatness calculation module 402, the pseudospectral corresponding to the current frame audio is calculated based on the original frequency domain spectral coefficients; based on the pseudospectral, the geometric mean and arithmetic mean of the pseudospectral corresponding to the current frame audio are calculated and obtained; and the pseudospectral flatness is calculated based on the geometric mean and arithmetic mean of the pseudospectral.

[0077] Optionally, in the noise update rate parameter determination module 403, the noise update rate parameter is set according to the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness.

[0078] Optionally, in the noise update speed parameter determination module 403, under the condition that the pseudo-spectral flatness is greater than a preset threshold, it is determined whether there is an audio frame in the preset number of audio frames before the current frame audio whose pseudo-spectral flatness is less than the preset threshold; if not, the noise update speed parameter is set according to the magnitude of the pseudo-spectral flatness corresponding to the current frame audio.

[0079] Optionally, in the noise energy calculation module 404, the non-voice band of the current frame audio is selected, and the pseudo-spectral median corresponding to the current frame audio is calculated; based on the pseudo-spectral median, the noise update rate parameter, and the moving average energy of the previous frame audio, the moving average energy of the current frame audio is calculated and used as the noise energy of the current frame audio.

[0080] The noise detection and elimination device of this application encodes audio data by utilizing the existing encoding process in the encoder, and at the same time uses pseudospectral flatness to set the noise update speed parameter to perform the transition between multiple frames. It has the characteristics of low power consumption and low latency, thus improving the user experience.

[0081] Figure 5 An embodiment of the audio encoding method of this application is shown.

[0082] exist Figure 5 In the illustrated embodiment, the audio encoding method of this application includes: process S501, obtaining the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding process of an encoder with improved discrete cosine transform processing; process S502, calculating the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients; process S503, setting a noise update rate parameter based on the pseudospectral flatness when the pseudospectral flatness is greater than a preset threshold; process S504, calculating the noise energy of the current frame audio based on the noise update rate parameter; process S505, calculating the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy; and process S506, updating other encoding modules in the encoder based on the updated spectral coefficients, and continuing to encode the current frame audio using the updated encoding modules.

[0083] Optionally, during the encoding or decoding of audio by a codec with improved discrete cosine transform processing, the original frequency domain spectral coefficients corresponding to the current frame audio are obtained, including: during the standard encoding or decoding process of the codec, performing an improved discrete cosine transform on the current frame audio to obtain the original frequency domain spectral coefficients.

[0084] Optionally, the pseudospectral flatness of the current frame audio is calculated based on the original frequency domain spectral coefficients, including: calculating the pseudospectral corresponding to the current frame audio based on the original frequency domain spectral coefficients; calculating and obtaining the pseudospectral geometric mean and pseudospectral arithmetic mean based on the pseudospectral; and calculating the pseudospectral flatness based on the pseudospectral geometric mean and pseudospectral arithmetic mean.

[0085] Optionally, when the pseudo-spectral flatness is greater than a preset threshold, a noise update rate parameter is set based on the pseudo-spectral flatness, including: setting the noise update rate parameter based on the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness.

[0086] Optionally, under the condition that the pseudo-spectral flatness is greater than a preset threshold, the noise update speed parameter is set according to the pseudo-spectral flatness, which further includes: under the condition that the pseudo-spectral flatness is greater than a preset threshold, determining whether there is an audio frame in the preset number of audio frames before the current frame audio that has a pseudo-spectral flatness less than a preset threshold; if not, then setting the noise update speed parameter according to the magnitude of the pseudo-spectral flatness corresponding to the current frame audio.

[0087] Optionally, the noise energy of the current frame audio is calculated based on the noise update rate parameter, including: selecting the non-voice band of the current frame audio and calculating the pseudo-spectral median corresponding to the current frame audio; and calculating the current frame moving average energy corresponding to the current frame audio based on the pseudo-spectral median, the noise update rate parameter, and the moving average energy of the previous frame audio, and using it as the noise energy of the current frame audio.

[0088] The audio encoding method of this application encodes audio data by utilizing the existing encoding process in the encoder, and at the same time uses pseudospectral flatness to set the noise update speed parameter to perform the transition between multiple frames. It has the characteristics of low power consumption and low latency, thus improving the user experience.

[0089] In one embodiment of this application, a computer-readable storage medium stores computer instructions, wherein the computer instructions are operated to perform the noise detection and cancellation method or audio coding method described in any embodiment. The storage medium may be located directly in hardware, in a software module executed by a processor, or in a combination of both.

[0090] Software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in this art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium.

[0091] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, but alternatively, it can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration. Alternatively, the storage medium can be integrated with the processor. The processor and storage medium can reside in an ASIC. The ASIC can reside in the user terminal. Alternatively, the processor and storage medium can reside as discrete components in the user terminal.

[0092] In one specific embodiment of this application, a computer device includes a processor and a memory, the memory storing computer instructions, wherein the processor operates the computer instructions to perform the method for detecting and eliminating noise or the audio coding method described in any embodiment.

[0093] In the embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0094] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0095] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A method of detecting and eliminating noise, characterized by, include: In the process of encoding or decoding audio using a codec with improved discrete cosine transform processing, the original frequency domain spectral coefficients corresponding to the audio of the current frame are obtained; The pseudospectral flatness of the current frame audio is calculated based on the original frequency domain spectral coefficients. Under the condition that the pseudo-spectral flatness is greater than a preset threshold, a noise update rate parameter is set according to the pseudo-spectral flatness. The step of setting the noise update rate parameter according to the pseudo-spectral flatness under the condition that the pseudo-spectral flatness is greater than the preset threshold includes: setting the noise update rate parameter according to the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness. The noise energy of the current frame audio is calculated based on the noise update rate parameter; and The updated spectral coefficients of the current frame audio are calculated based on the original frequency domain spectral coefficients and the noise energy, and subsequent encoding or decoding processes are performed based on the updated spectral coefficients.

2. The method of claim 1, wherein, In the process of encoding or decoding audio using a codec with improved discrete cosine transform processing, obtaining the original frequency domain spectral coefficients corresponding to the current frame audio includes: During the standard encoding or decoding process of the codec, an improved discrete cosine transform is performed on the current frame audio to obtain the original frequency domain spectral coefficients.

3. The method of claim 1, wherein, The step of calculating the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients includes: Calculate the pseudospectral corresponding to the current frame audio based on the original frequency domain spectral coefficients; Based on the pseudospectrum, calculate and obtain the geometric mean and arithmetic mean of the pseudospectrum corresponding to the audio of the current frame; The pseudospectral flatness is calculated based on the geometric mean and arithmetic mean of the pseudospectral.

4. The method of claim 1, wherein, The step of setting a noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold further includes: Under the condition that the pseudo-spectral flatness is greater than a preset threshold, determine whether there is an audio frame in the preset number of audio frames before the current frame whose pseudo-spectral flatness is less than the preset threshold. If not, the noise update speed parameter is set according to the magnitude of the pseudospectral flatness corresponding to the current frame audio.

5. The method of claim 1, wherein, The step of calculating the noise energy of the current frame audio based on the noise update rate parameter includes: Select the non-speech bands of the current frame audio and calculate the pseudospectral midpoint corresponding to the current frame audio; The current frame moving average energy corresponding to the current frame audio is calculated based on the pseudospectral median, the noise update rate parameter, and the moving average energy of the previous frame audio, and is used as the noise energy of the current frame audio.

6. An apparatus for detecting and eliminating noise, characterized by include: The frequency domain spectral coefficient acquisition module acquires the original frequency domain spectral coefficients corresponding to the current frame audio during the encoding or decoding process of the codec with improved discrete cosine transform processing. The pseudospectral flatness calculation module calculates the pseudospectral flatness of the current frame audio based on the original frequency domain spectral coefficients. A noise update rate parameter determination module, which sets a noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than a preset threshold, wherein setting the noise update rate parameter based on the pseudo-spectral flatness when the pseudo-spectral flatness is greater than the preset threshold includes: setting the noise update rate parameter based on the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness; A noise energy calculation module that calculates the noise energy of the current frame audio based on the noise update rate parameter; and The spectral coefficient update module calculates the updated spectral coefficients of the current frame audio based on the original frequency domain spectral coefficients and the noise energy, and performs subsequent encoding or decoding processes based on the updated spectral coefficients.

7. An audio encoding method, characterized by, include: In the process of encoding audio by an encoder with improved discrete cosine transform processing, the original frequency domain spectral coefficients corresponding to the audio of the current frame are obtained; The pseudospectral flatness of the current frame audio is calculated based on the original frequency domain spectral coefficients. Under the condition that the pseudo-spectral flatness is greater than a preset threshold, a noise update rate parameter is set according to the pseudo-spectral flatness. The step of setting the noise update rate parameter according to the pseudo-spectral flatness under the condition that the pseudo-spectral flatness is greater than the preset threshold includes: setting the noise update rate parameter according to the magnitude of the pseudo-spectral flatness, wherein the magnitude of the noise update rate parameter is directly proportional to the magnitude of the pseudo-spectral flatness. The noise energy of the current frame audio is calculated based on the noise update rate parameter; and The updated spectral coefficients of the current frame audio are calculated based on the original frequency domain spectral coefficients and the noise energy. The other encoding modules in the encoder are updated according to the updated spectral coefficients, and the updated encoding modules are used to continue encoding the audio of the current frame.

8. A computer-readable storage medium storing a computer program, wherein the computer program is operated to perform the method for detecting and eliminating noise as described in any one of claims 1-5 or the audio coding method as described in claim 7.

9. A computer device comprising a processor and a memory, the memory storing a computer program, wherein: The processor operates the computer program to perform the method for detecting and eliminating noise as described in any one of claims 1-5 or the audio coding method as described in claim 7.

Citation Information

Patent Citations

  • Noise signal processing method, noise signal generation method, encoder, decoder, and encoding and decoding system

    US20170018277A1

  • Method for detecting driving noise and improving speech recognition in a vehicle

    US20170186423A1