Detecting Environmental Noise in User-Generated Content

A two-stage noise classification system effectively distinguishes UGC and PGC noise in audio content, improving audio quality by applying tailored processing techniques, thus enhancing the user experience.

JP7735537B2Active Publication Date: 2025-09-08DOLBY LABORATORIES LICENSING CORP
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
JP2024508330
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-03
Filing Date
2022-08-23
Publication Date
2025-09-08
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

Existing audio processing systems struggle to differentiate between user-generated content (UGC) noise and professional-generated content (PGC) noise, leading to poor audio quality in UGC due to inappropriate noise reduction or enhancement techniques.

Method used

A two-stage noise classification system using machine learning models to detect and distinguish between UGC and PGC noise, applying appropriate audio processing techniques based on the classification results.

Benefits of technology

Improves audio quality in UGC by accurately identifying and managing UGC noise, enhancing the user experience by applying targeted noise reduction or volume adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007735537000012
    Figure 0007735537000012
  • Figure 0007735537000013
    Figure 0007735537000013
  • Figure 0007735537000014
    Figure 0007735537000014
Patent Text Reader

Abstract

The method of audio processing includes classifying an audio signal into noise or non-noise using a first model. In case of a noise signal, the audio signal is classified into user generated content (UGC) noise or professional generated content (PGC) noise using a second model. In case of a non-noise signal or PGC noise, the audio signal is processed using a first audio processing process. In case of UGC noise, the audio signal is processed using a second audio processing process.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of priority from PCT International Application No. PCT / CN2021 / 114746, filed August 26, 2021, U.S. Provisional Application No. 63 / 244,495, filed September 15, 2021, and European Patent Application No. 21206205.3, filed November 3, 2021, each of which is incorporated by reference herein in its entirety.

[0002] [Technical field] The present disclosure relates to audio processing, and in particular to noise reduction. [Background technology]

[0003] Unless otherwise indicated herein, the approaches described in this section are not prior art to the claims of this application and are not admitted to be prior art by inclusion in this section.

[0004] Multimedia content, including audio, video, and combined audio / video, has always been important in the entertainment industry. Professional-generated content (PGC), such as movies and television programs, was previously the dominant form of multimedia content. However, in recent years, user-generated content (UGC) has grown exponentially. This has benefited from the rapid development of capture devices, network platforms, and playback technologies. Portable devices, such as smartphones and tablets, have become commonplace for capture devices. Users can independently capture and create UGC using their built-in cameras and microphones. Additionally, various platforms, such as popular video websites and emerging mobile applications, have greatly accelerated the spread of UGC.

[0005] Numerous techniques for content playback have been developed with the aim of enhancing the visual and auditory experience. Audio processing techniques, such as Dolby™ Audio Processing, can be applied during playback to enhance sound quality. While audio processing systems have primarily focused on PGC, the growing popularity of UGC provides opportunities to apply audio processing to UGC as well. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] U.S. Patent No. 9,064,497 [Patent Document 2] U.S. Patent No. 9,107,010 [Patent Document 3] U.S. Patent No. 9,984,701 [Patent Document 4] U.S. Patent No. 8,682,250 [Patent Document 5] U.S. Patent No. 8,238,497 [Patent Document 6] U.S. Patent No. 6,859,773 [Patent Document 7] U.S. Patent No. 7,171,246 [Patent Document 8] U.S. Patent No. 7,684,982 [Patent Document 9] U.S. Patent No. 9,633,671 [Patent Document 10] U.S. Patent No. 9,769,564 [Patent Document 11] U.S. Patent No. 7,464,029 [Patent Document 12] U.S. Patent No. 9,711,130 [Patent Document 13] U.S. Patent No. 8,325,939 [Patent Document 14] U.S. Patent No. 9,609,416 [Patent Document 15] U.S. Patent No. 9,934,791 [Patent Document 16] U.S. Patent Application Publication No. 2016 / 0155434 [Patent Document 17] U.S. Patent Application Publication No. 2020 / 0125316 [Patent Document 18] U.S. Patent Application Publication No. 2020 / 0020312 [Non-patent literature]

[0007] [Non-Patent Document 1] Fatemeh Saki, Abhishek Sehgal, Issa Panahi and Nasser Kehtarnavaz, “Smartphone-Based Real-Time Classification of Noise Signals using Subband Features and Random Forest Classifier”, in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), DOI 10.1109 / ICASSP.2016.7472068. [Non-patent document 2] Vanishree Gopalakrishna, Nasser Kehtarnavaz, Taher S. Mirzahasanloo and Philipos C. Loizou, “Real-Time Automatic Tuning of Noise Suppression Algorithms for Cochlear Implant Applications”, in IEEE Transactions on Biomedical Engineering (Volume: 59, Issue: 6, June 2012), DOI 10.1109 / TBME.2012.2191968. [Non-patent document 3] Fatemeh Saki and Nasser Kehtarnavaz, “Background noise classification using random forest tree classifier for cochlear implant applications”, in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), DOI 10.1109 / ICASSP.2014.6854270. [Non-Patent Document 4] Fatemeh Saki, Taher Mirzahasanloo and Nasser Kehtarnavaz, “A multi-band environment-adaptive approach to noise suppression for cochlear implants”, in Annu Int Conf IEEE Eng Med Biol Soc 2014, DOI 10.1109 / EMBC.2014.6943934. [Non-Patent Document 5] Peipei Shen, Zhou Changjun and Xiong Chen, “Automatic Speech Emotion Recognition using Support Vector Machine”, in Proceedings of 2011 International Conference on Electronic & Mechanical Engineering and Information Technology, DOI 10.1109 / EMEIT.2011.6023178. [Non-Patent Document 6] Humaid Alshamsi, Veton Kepuska, Hazza Alshamsi and Hongying Meng, “Automated Speech Emotion Recognition on Smart Phones”, in 2018 9th IEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), DOI 10.1109 / UEMCON.2018.8796594. Summary of the Invention

[0008] One problem with existing audio processing systems is that the techniques used for PGC can differ from those used for UGC. In contrast to high-quality PGC, much UGC has low audio quality. This can be due to unprofessional recording devices, complex environments, and fewer editing procedures. Quality issues include, but are not limited to, poor speech intelligibility and strong reverberation. One of the most common issues is environmental noise contained in UGC (hereinafter referred to as UGC noise). UGC noise can be easily captured using a mobile phone in a real-world scene. Generally, UGC noise is background noise and is therefore meaningless or unnecessary. Therefore, UGC noise, especially for nearly stationary noise, should not be boosted by any volume adjustment techniques because boosting this type of noise is clearly perceived by listeners and adversely affects the user experience. On the other hand, if the audio processing system knows that UGC noise is present in the content, it can apply an appropriate noise reduction method to the UGC noise to enhance the audio quality.

[0009] However, PGC also contains near-stationary noise-like content (hereinafter referred to as PGC noise). PGC noise may generally include noise intervals, such as background sound intervals between adjacent dialogue intervals in a movie. This PGC noise is usually captured independently from the dialogue using professional recording devices and carefully processed by an audio mixer during the content creation stage. In contrast to UGC noise, PGC noise is part of the content and is usually desired from the artist and content creator's perspective. In such cases, noise reduction methods should not be applied, and techniques such as volume leveling can safely boost the PGC noise.

[0010] As a result, UGC noise and PGC noise should be treated differently. A method for detecting UGC stationary noise while distinguishing it from PGC noise is highly desirable. Such a method can further be used to manipulate post-processing techniques for audio content playback. Embodiments are directed to a two-stage noise classification system.

[0011] According to one embodiment, a computer-implemented method for audio processing includes receiving an audio signal and calculating a first confidence score for the audio signal using a first machine learning model. The method further includes generating a processed audio signal by processing the audio signal according to a first audio processing process when the first confidence score indicates the presence of non-noise. The method further includes calculating a second confidence score for the audio signal using a second machine learning model when the first confidence score indicates the presence of noise. The method further includes generating a processed audio signal by processing the audio signal according to a second audio processing process when the second confidence score indicates the presence of user-generated content (UGC) noise. The method further includes generating a processed audio signal by processing the audio signal according to the first audio processing process when the second confidence score indicates the presence of professional-generated content (PGC) noise.

[0012] The step of calculating the first confidence score may include the steps of extracting a first plurality of features from the audio signal, classifying the first plurality of features using a first machine learning model, calculating a noise confidence score based on the result of classifying the first plurality of features, and calculating a weight based on the noise confidence score.

[0013] The step of calculating the second confidence score may include the steps of extracting a second plurality of features from the audio signal, where the second plurality of features are extracted over a longer period of time than the first plurality of features are extracted; calculating a second plurality of statistics based on the second plurality of features, where the second plurality of statistics are weighted according to the weights; classifying the second plurality of features and the second plurality of statistics using a second machine learning model; and calculating the second confidence score based on the result of classifying the second plurality of features and the second plurality of statistics.

[0014] According to another embodiment, an apparatus includes a loudspeaker and a processor configured to control the apparatus to implement one or more of the methods described herein. The apparatus may additionally include details similar to the details of one or more of the methods described herein.

[0015] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls an apparatus to perform processes including one or more of the methods described herein.

[0016] The following detailed description and accompanying drawings provide a further understanding of the nature and advantages of various implementations. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a block diagram of a noise classification system 100. [Figure 2] FIG. 2 is a block diagram showing details of the noise detector 102 (see FIG. 1). [Figure 3] FIG. 2 is a block diagram showing details of a noise discriminator 104 (see FIG. 1). [Figure 4] 4 is a mobile device architecture 400 for implementing the features and processes described herein, according to one embodiment. [Figure 5] 5 is a flowchart of a method 500 of audio processing. DETAILED DESCRIPTION OF THE INVENTION

[0018] Techniques related to audio processing are described herein. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may further include variations and equivalents of the features and concepts described herein.

[0019] In the following description, various methods, processes, and procedures are detailed. While certain steps may be described in a particular order, such order is primarily for convenience and clarity. Certain steps may be repeated more than once, may occur before or after other steps, or may occur in parallel with other steps, even if described in a different order. A second step is required to follow a first step only if the first step must be completed before the second step can begin. Such situations will be specifically pointed out if not clear from the context.

[0020] The terms "and," "or," and "and / or" are used herein. Such terms should be read as having an inclusive meaning. For example, "A and B" may mean at least "both A and B," or "at least both A and B." As another example, "A or B" may mean at least "at least A," "at least B," "both A and B," or "at least both A and B." As another example, "A and / or B" may mean at least "A and B," "A or B." When an exclusive or is intended, it is specifically stated, for example, "either A or B," "at most one of A and B," etc.

[0021] This specification describes various processing functions associated with structures such as blocks, elements, components, circuits, etc. Generally, these structures may be implemented by a processor controlled by one or more computer programs.

[0022] 1 is a block diagram of a noise classification system 100. Generally, the noise classification system 100 receives an audio signal, performs noise classification, and generates a processed audio signal according to the noise classification. The components of the noise classification system 100 may be implemented as one or more computer programs executed by a processor. The noise classification system 100 includes a noise detector 102, a noise discriminator 104, a PGC audio processor 106, and a UGC processor 108.

[0023] The noise detector 102 receives the audio signal 120 and performs noise detection. When the noise detection indicates that the audio signal 120 contains noise, such as UGC noise, PGC noise, etc., the noise detector 102 provides the audio signal 120 to a noise discriminator 104 for further processing. When the noise detection indicates that the audio signal 120 contains non-noise, such as music, speech, etc., the noise detector 102 provides the audio signal 120 to a PGC audio processor 106 for further processing. Further details of the noise detector 102 are provided with reference to FIG. 2.

[0024] The noise discriminator 104 receives the audio signal 120, for example, from the noise detector 102, and performs noise discrimination. Recall that the audio signal 120 provided to the noise discriminator 104 has previously been classified as "noise" by the noise detector 102. When the noise discrimination 104 indicates that the audio signal 120 contains PGC noise, the noise discriminator 104 provides the audio signal 120 to the PGC audio processor 106 for further processing. When the noise discrimination indicates that the audio signal 120 contains UGC noise, the noise discriminator 104 provides the audio signal 120 to the UGC audio processor 108 for further processing. Further details of the noise discriminator 104 are provided in FIG. 3.

[0025] The PGC audio processor 106 receives the non-noise-indicating audio signal 120 from the noise detector 102 or the PGC noise-indicating audio signal 120 from the noise discriminator 104 and performs audio processing according to a first audio processing process to generate a processed audio signal 130a. The first audio processing process generally corresponds to an audio processing setting appropriate for use on non-noise audio, such as dialogue emphasis, volume adjustment, etc. For example, if a listener desires to perform volume adjustment when outputting the processed audio signal, for example, on their mobile phone implementing the noise classification system 100, using the first audio processing setting results in output audio with the volume adjustment applied. The first audio processing setting is also appropriate for use when PGC noise is detected because it is more likely to adapt the listener's experience to the audio experience intended by the content creator. For example, if volume adjustment is desired by the listener, it would be appropriate to apply the volume adjustment to the PGC.

[0026] The UGC audio processor 108 receives the audio signal 120 indicative of UGC noise from the noise discriminator 104 and performs audio processing according to a second audio processing process to generate a processed audio signal 130b. The second audio processing process generally corresponds to audio processing settings appropriate for use with UGC noise, e.g., stationary noise. For example, if a listener desires to perform noise reduction when outputting the processed audio signal, e.g., on their mobile phone implementing the noise classification system 100, using the second audio processing settings results in output audio with noise reduction applied. Because UGC often has stationary noise, performing noise reduction on this content can improve the listener experience.

[0027] Consider the following example. Assume that a user desires both volume adjustment and noise reduction to be applied to the audio output, and the content includes traffic noise, e.g., some type of stationary noise. If the listener is watching a Hollywood movie on their mobile phone, it is appropriate to perform volume adjustment but not noise reduction, because it is appropriate not to excessively reduce the traffic noise to preserve the artistic intent of the content creator. In this situation, the noise classification system 100 detects the UGC noise, applies a first audio processing setting, and outputs the audio output appropriately. If the listener is watching a video taken while walking outdoors on a sidewalk, it is appropriate to perform both volume adjustment and noise reduction, because reducing the traffic noise rather than boosting it when performing volume adjustment will improve the listener experience. In this situation, the noise classification system 100 detects the UGC noise, applies a second audio processing setting, and outputs the audio output appropriately.

[0028] Generally, the noise classification system 100 operates in real time by processing portions of the audio signal 120. Each portion is called a clip or audio clip, and the noise classification system 100 may perform a process on a clip-by-clip basis. Each audio clip contains a defined number of consecutive audio frames. Taking audio with a sampling rate of 48 kHz as an example, a typical duration of an audio frame may be 1024 samples, or approximately 21.34 milliseconds, and an audio clip may contain 48 non-overlapping frames, or approximately 1.024 seconds. Adjacent audio frames may contain overlapping samples. Adjacent audio clips may also contain overlapping frames. For a given input audio clip, the noise detector 102 determines whether it is a noise clip that should be sent to the noise discriminator 104, or whether it belongs to another type, such as speech, music, etc., that should not be sent to the noise discriminator 104.

[0029] The noise classification system 100 may be referred to as a two-stage noise classification system. The noise detector 102, also referred to as the first stage, operates to classify a given audio clip according to a defined number of stationary noise-like frames. Note that the stationary noise-like content can be either captured environmental noise within the UGC, also referred to as UGC noise, or noise-like content that is part of the PGC, also referred to as PGC noise. The noise discriminator 104, also referred to as the second stage, is triggered for a nearly stationary noise clip determined by the noise detector 102. The noise discriminator 104 operates to distinguish noise types in terms of UGC noise versus PGC noise. If the clip is determined to be UGC noise, it will trigger the UGC processing path, which will route through the UGC audio processor 108. Otherwise, it will remain in the original processing path adapted for PGC content processing, which will route through the PGC audio processor 106.

[0030] For real-world content, a given audio clip may contain some noise frames and other types of frames, such as speech, music, transient sound events, etc. In other words, pure stationary noise clips are uncommon. To address this issue, the noise classification system 100 implements an attention-like mechanism to focus on the noise portions of a given audio clip. Specifically, for each frame, the noise detector 102 calculates a frame weight to indicate the likelihood of stationary noise, as further detailed in FIG. 2. The noise detector 102 calculates the frame weight based on a data-driven method called classification. The noise detector 102 uses the frame weight to calculate a clip confidence, which indicates the likelihood of near-stationary noise. Additionally, if the audio clip is a near-stationary noise clip as determined by the noise detector 102, the frame weight is further used to steer feature calculation in the noise discriminator 104, as further detailed in FIG. 3.

[0031] 2 is a block diagram illustrating details of the noise detector 102 (see FIG. 1). Generally, the noise detector 102 receives an audio signal, performs noise classification, and generates a confidence score according to the noise classification. The components of the noise detector 102 may be implemented as one or more computer programs executed by a processor. The noise detector 102 includes a feature extractor 202, a classifier 204, a model 206, a determiner 208, and optionally a root-mean-square (RMS) calculator 210.

[0032] The feature extractor 202 receives the audio signal 120, extracts features 220 from the audio signal 120, and provides the features 220 to the classifier 204. As described above, the audio signal 120 corresponds to an audio clip of the input audio signal, e.g., 48 frames (which may not overlap), and the feature extractor 202 operates on a portion of the audio clip, e.g., less than 48 frames, called a “short clip.” Using the short clip rather than using only the current frame may improve generalization ability when calculating the noise confidence score. Here, “generalization” refers to the ability of the noise detector 102 to adapt to new data not previously seen. A short clip may include the current frame of the clip and several previous frames. For example, a short clip may include five consecutive frames, e.g., the current frame and four previous frames. The current short clip may overlap with a previous short clip. For example, the overlap may have a hop size of one frame. In other words, the short clip moves one frame per step. As an example, for a given clip having 48 frames, one short clip may be frames 1-5, another may be frames 2-6, another may be frames 3-7, and so on, and another may be frames 44-48. Each frame is generally represented by its context corresponding to the short clip, e.g., the frame itself and the four previous frames, and features are extracted from this context, and classification is also performed based on this context. In this manner, the noise detector 102 operates on a short clip-by-short clip basis.

[0033] The features 220 correspond to audio features of the short clip. These audio features include one or more of temporal features, spectral features, time-frequency features, etc. The temporal features may include one or more of autocorrection coefficients (ACC), linear predictive coding coefficients (LPCC), zero-crossing rates (ZCR), etc. The spectral features may include one or more of spectral centroids, spectral roll-off, spectral energy distribution, spectral flatness, spectral entropy, Mel-frequency cepstral coefficients (MFCC), etc. The time-frequency features may include one or more of spectral flux, chroma, etc. The features 220 may also include statistics of the other features described above. These statistics may include the mean, standard deviation, and higher-order statistics, such as skewness, kurtosis, etc. For example, the features 220 may include the mean and standard deviation of the spectral energy distribution.

[0034] The classifier 204 receives the features 220 and performs classification of the features 220 using the model 206 to generate a noise confidence score 222. For each frame in a given clip, the classifier 204 may calculate a noise confidence score 222 for each frame based on the context of that frame (e.g., the current frame) and four previous frames corresponding to the short clip. The noise confidence score 222 may range from 0 to 1, with a higher score meaning a more likely noise frame, and vice versa.

[0035] The classifier 204 may correspond to a machine learning system, and the model 206 may correspond to a machine learning model. The model 206 may be obtained by a data-driven method, and the model 206 may be trained offline using a set of training data including positive training data, e.g., “noise” data including both UGC noise and PGC noise, and negative training data, e.g., “non-noise” data such as music, speech, etc. In other words, the training data is tagged into various categories, e.g., UGC noise, PGC noise, and non-noise, and the model 206 is obtained as a result of the model 206 being exposed to the training data during a training process. As a result, one definition of “PGC” is “data tagged in the training process as being expert-generated content,” one definition of “PGC noise” is “data tagged in the training process as being expert-generated content and having noise,” one definition of “UGC” is “data tagged in the training process as being other than expert-generated content,” and one definition of “UGC noise” is “data tagged in the training process as being other than expert-generated content and having noise.” An amount of training data that provides acceptable results is 50 hours of training data. This amount can be varied as needed; decreasing the amount may decrease the model's accuracy, while increasing the amount may increase the model's accuracy. Generally, the features 220 extracted by the feature extractor 202 correspond to the features used in training the model 206, e.g., the features extracted from each short clip.

[0036] The classifier 204 may implement various machine learning systems, including an adaptive boosting (AdaBoost) system, a deep neural network (DNN) system, etc. Generally, an AdaBoost system combines the outputs of other learning algorithms, also called "weak learners," into a weighted sum, and is "adaptive" in the sense that subsequent weak learners are fine-tuned to favor instances misclassified by previous classifiers. Generally, a DNN system is a neural network with multiple layers between an input layer and an output layer. In the case of an AdaBoost classifier, the classifier 204 may convert the output noise confidence to the interval [0,1] by applying a sigmoid function. The sigmoid function σ(z) may be defined according to Equation (1):

number

[0037] In equation (1), the output noise confidence σ(z) is transformed into the interval [0,1] by applying an inverse exponential function to the input z, which is a combination of the outputs of the weak learners. The output noise confidence approaches 0 as the input decreases, is 0.5 when the input is 0, and approaches 1 as the input increases. As a result, the output noise confidence increases as z increases.

[0038] If a DNN is used, the classifier 204 may use a sigmoid function as the activation function of the output layer. In either case, the classifier 204 sends the obtained noise confidence score to the determiner 208 as the noise confidence score 222.

[0039] The determiner 208 receives the noise confidence scores 222 and generates weights 224 based on the noise confidence scores 222. The determiner 208 also determines clip confidences w cIf the clip confidence is greater than a defined threshold, the clip is classified as noise and sent to the noise discriminator 104. In other words, the noise confidence score 222 for each frame in a given clip is used to classify the clip confidence w c It is used to calculate the clip confidence w c Further details regarding are provided below with reference to equations (4) and (8).

[0040] The determiner 208 uses the noise confidence scores 222 to calculate weights 224, also called frame weights, according to equation (2):

number

[0041] In equation (2), i is the current frame number and w f (i) is the weight 224 of the current frame, n(i) is the noise confidence score 222 of the current frame, and σ′(·) is the modified sigmoid function defined according to equation (3):

number

[0042] In equation (3), C is a scale factor and θ is a threshold. Generally, C is a positive value. Typical values ​​of C are 16, e.g., in the range of 10 to 20, and typical values ​​of θ are 0.5, e.g., in the range of 0.45 to 0.55. Thus, weight 224 increases when noise confidence score 222 exceeds the threshold, and weight 224 decreases when noise confidence score 222 falls below the threshold. The scale factor C can be adjusted to adjust the balance between weight 224 and noise confidence score 222; decreasing the scale factor decreases the contribution of weight 224, and increasing the scale factor increases the contribution of weight 224. The threshold θ can be adjusted to adjust the sensitivity of noise detection; increasing the threshold decreases weight 224, and decreasing the threshold increases weight 224.

[0043] The determiner 208 uses the weights 224 not only to calculate the clip confidence, as described in more detail below, but also to determine whether the clip should be sent to the noise discriminator 104, as described in more detail with reference to FIG. 3. The clip confidence calculation uses the number of noise frames in the clip. To measure the number of noise frames in a clip, the determiner 208 determines the noisiness weight w using equation (4): noi Calculate:

number

[0044] In equation (4), c(i) is a constant coefficient applied to the i-th frame. Note that the coefficient c(i) may be the same for all frames or may vary from frame to frame. If the coefficient is the same, the weight w noiapproximates the proportion of noise-like frames in the entire clip. Alternatively, in other scenarios where low latency is desired, the coefficients can be increased over time. In essence, the current frame is assigned the largest coefficient, and previous frames are assigned smaller coefficients. This method helps the framework respond quickly to the current frame.

[0045] In other words, the noisiness weight w noi is the ratio between two components: (1) a constant coefficient and frame weight for each frame summed over all frames in the clip, and (2) a constant coefficient for each frame summed over all frames in the clip. In effect, the noisiness weight is a weighted combination of the frame weights.

[0046] Therefore, the clip reliability w c is the noisiness weight w noi The clip confidence is in the range of 0 to 1.

[0047] Clip reliability c If the clip confidence w exceeds a defined threshold, the clip and weights 224 are sent to the noise discriminator 104 for further processing. A typical value for the threshold is 0.5. c If does not exceed the defined threshold, the clip is processed by the PGC audio processor 106 .

[0048] The RMS calculator 210 is optional. If present, it operates to limit the audio RMS to a specific range. The motivation is that noise within a specific RMS / loudness range exacerbates artifacts caused by post-processing techniques, e.g., loudness adjustment methods, while in other ranges, the artifacts are not perceptually noticeable. A simple approach is to use hard decisions with a predefined threshold. That is, noisy clips whose RMS exceeds the threshold are not processed. However, this can cause instability, especially for clips whose RMS frequently fluctuates around the threshold.

[0049] To overcome the instability problem, the RMS calculator 210 calculates the average RMS gain 226 for each frame. The RMS gain is used to weight the noise confidence score (clip confidence w, discussed below). c Since the RMS of the i-th frame is p(n), it is beneficial to operate at the short clip level for consistency with the feature extraction performed by feature extractor 202. Specifically, if the RMS of the i-th frame is denoted by p(n), then the short clip level RMS for the i-th frame is calculated by equation (5):

number

[0050] In equation (5), L is the short clip length in frames, for example, 5 frames.

[0051] If the RMS calculator 210 is present, the determiner 208 also receives the average RMS gain 226. The determiner 208 determines the RMS-based weights w using equation (6). rms Calculate:

number

[0052] In equation (6),

number

number

[0053] In equation (7), P L is the lower limit of the RMS interval, and P U is the upper limit of the RMS interval, and a L is a number greater than zero, and a U is a number less than zero.

[0054] In other words, equation (6) is the RMS-based weight w rms corresponds to a function g applied to the ratio between two components: (1) the average RMS gain and frame weight of each frame summed over all frames in the clip, and (2) the weight of each frame summed over all frames in the clip. Equation (7) states that function g is applied such that when the ratio of equation (6) is less than or equal to the lower limit of the RMS interval, the lower limit of the RMS interval is applied to the ratio to generate the RMS-based weight; when the ratio of equation (6) is greater than the upper limit of the RMS interval, the upper limit of the RMS interval is applied to the ratio to generate the RMS-based weight; otherwise, the RMS-based weight is set to 1.

[0055] The determiner 208 determines the clip confidence w according to equation (8). c Calculate:

number

[0056] In equation (8), w noi is the noisiness weight according to equation (4), and w rmsis the RMS-based weight according to equation (6). In other words, the clip confidence is a combination of the noisiness weight and the RMS-based weight.

[0057] 3 is a block diagram illustrating details of the noise discriminator 104 (see FIG. 1). In general, the noise discriminator 104 receives an audio signal classified as noise by, for example, the noise detector 102, performs noise discrimination, and generates a confidence score according to the noise discrimination. The components of the noise discriminator 104 may be implemented as one or more computer programs executed by a processor. The noise discriminator 104 includes a feature extractor 302, a classifier 304, a model 306, and a determiner 308.

[0058] The feature extractor 302 receives the audio signal 120 and the weights 224 (see FIG. 2 ) and extracts features 320 from the audio signal 120 based on the weights 224. As described above, the audio signal 120 corresponds to an audio clip of the input audio signal, e.g., 48 frames (which may not overlap), and the feature extractor 302 operates on the audio clip. Recall that the noise detector 102 operates on short clips, whereas the noise discriminator 104 operates on a longer period than the noise detector 102. More specifically, for a given clip, the feature extractor 302 extracts various features for each frame in the clip, also referred to as frame features, and the feature extractor 302 calculates statistics of the frame features. The feature extractor 302 uses the weights 224 as weighting factors when calculating the statistics. Thus, the features 320 correspond to both the frame features extracted based on each frame and the statistics calculated based on the clip. In this manner, the noise discriminator 104 operates on a clip-by-clip basis.

[0059] The frame features extracted by the feature extractor 302 may include one or more of temporal features, spectral features, time-frequency features, etc. The temporal features may include one or more of autocorrection coefficients (ACC), linear predictive coding coefficients (LPCC), zero-crossing rates (ZCR), etc. The spectral features may include one or more of spectral centroids, spectral roll-off, spectral energy distribution, spectral flatness, spectral entropy, Mel-frequency cepstral coefficients (MFCC), etc. The time-frequency features may include one or more of spectral flux, chroma, etc. The frame features extracted by the feature extractor 302 may be the same type of features as the features 220 extracted by the feature extractor 202 (see FIG. 2).

[0060] The feature extractor 302 may calculate various weighted statistics, including a weighted mean, a weighted standard deviation, etc. The feature extractor 302 may calculate the weighted mean μ using equation (9):

number

[0061] In equation (9), v(i) corresponds to the frame feature v extracted in frame i, also called frame index i, and w f where m corresponds to the frame weight (see equation (2) and weights 224), and M corresponds to the total number of frames in the clip, e.g., 48 frames. In other words, the weighted average corresponds to the ratio between (1) the sum of the weighted frame features and (2) the sum of the weights for a given clip.

[0062] The feature extractor 302 may calculate the weighted standard deviation σ using equation (10):

number

[0063] The variables in equation (10) are as described above with respect to equation (9). In other words, the weighted standard deviation corresponds to the square root of the ratio between (1) the weighted sum of the squares of the frame features and (2) the sum of the weights for a given clip.

[0064] The classifier 304 receives the features 320 and performs classification of the features 320 using the model 306 to generate a noise confidence score 322. The noise confidence score 322 indicates the likelihood of PGC noise versus UGC noise. For example, the classifier 304 may implement a sigmoid function to transform the noise confidence score 322 into the interval [0, 1], with a score closer to 0 indicating a higher likelihood of one type, e.g., UGC noise, and a score closer to 1 indicating a higher likelihood of the other type, e.g., PGC noise.

[0065] The classifier 304 may correspond to a machine learning system, and the model 306 may correspond to a machine learning model. The model 306 may be obtained by a data-driven method or may be trained offline using a set of training data including PGC noise training data and UGC noise training data. In other words, the training data is tagged with various categories, such as UGC noise, PGC noise, etc., and the model 306 is obtained as a result of the model 306 being exposed to the training data during the training process. Viewing the PGC noise training data as “positive” training data and the UGC noise training data as “negative” training data, the noise confidence score 322 is larger (e.g., greater than 0.5) when the feature 320 corresponds to a PGC and smaller (e.g., less than 0.5) when the feature 320 corresponds to a UGC. An amount of training data that provides acceptable results is 50 hours of training data. This amount can be varied as needed; reducing the amount may decrease the model's accuracy, while increasing the amount may increase the model's accuracy. Generally, the features 320 extracted by the feature extractor 302 correspond to the features used in training the model 306, e.g., the features extracted from each frame in a given clip and the statistics calculated therefrom.

[0066] The classifier 304 may implement various machine learning systems, including an adaptive boosting (AdaBoost) system, a deep neural network (DNN) system, etc. In the case of an AdaBoost classifier, the classifier 304 may convert the noise confidence scores to the interval [0, 1] by applying a sigmoid function (see equation (1)). If a DNN is used, the classifier 304 may use the sigmoid function as the activation function of the output layer.

[0067] The determiner 308 receives the noise confidence score 322 and generates a classification result 324 based on the noise confidence score 322 and a threshold. When the noise confidence score 322 is greater than the threshold, e.g., when the “positive” training data corresponds to the PGC training data, the clip is classified as PGC, and the noise discriminator 104 controls the PGC audio processor 106 to process the clip according to a desired PGC audio processing technique, such as volume adjustment. When the noise confidence score 322 is less than the threshold, e.g., when the “negative” training data corresponds to the UGC training data, the clip is classified as UGC, and the noise discriminator 104 controls the UGC audio processor 108 to process the clip according to a desired UGC audio processing technique, such as stationary noise reduction.

[0068] 4 illustrates a mobile device architecture 400 for implementing the features and processes described herein, according to one embodiment. Architecture 400 may be implemented in any electronic device, including, but not limited to, mobile devices such as desktop computers, home audio / visual (AV) equipment, radio broadcasting equipment, smartphones, tablet computers, laptop computers, wearable devices, etc. In the illustrated exemplary embodiment, architecture 400 is for a laptop computer and includes processor(s) 401, peripherals interface 402, audio subsystem 403, loudspeaker 404, microphone 405, sensors 406 such as accelerometers, gyros, barometers, magnetometers, cameras, etc., location processor 407 such as a GNSS receiver, wireless communication subsystem 408 such as Wi-Fi, Bluetooth, cellular, etc., and I / O subsystem(s) 409 including touch controller 410 and other input controllers 411, touch surface 412 and other input / control devices 413. Other architectures having more or fewer components may also be used to implement the disclosed embodiments.

[0069] Memory interface 414 is coupled to processor 401, peripherals interface 402, and memory 415, such as Flash, RAM, or ROM. Memory 415 stores computer program instructions and data, including, but not limited to, operating system instructions 416, communications instructions 417, GUI instructions 418, sensor processing instructions 419, telephony instructions 420, electronic messaging instructions 421, web browsing instructions 422, audio processing instructions 423, GNSS / navigation instructions 424, and applications / data 425. Audio processing instructions 423 include instructions for performing the audio processing described herein.

[0070] According to one embodiment, architecture 400 may correspond to a mobile phone. A user may use the mobile phone to output PGC, in which case noise classification system 100 (see FIG. 1 ) controls the mobile phone to generate audio output using audio processing appropriate for the PGC, e.g., PGC audio processor 106. A user may use the mobile phone to capture and output UGC, in which case noise classification system controls the mobile phone to generate audio output using audio processing appropriate for the UGC, e.g., UGC audio processor 108.

[0071] 5 is a flowchart of a method of audio processing 500. Method 500 may be performed by a device, such as a laptop computer, a mobile phone, or the like, having the components of architecture 400 of FIG. 4 to implement functionality such as noise classification system 100 (see FIG. 1), for example, by executing one or more computer programs.

[0072] At 502, an audio signal is received. For example, the noise classification system 100 (see FIG. 1) may receive the audio signal 120. The audio signal 120 may include audio samples, which may be arranged as audio frames, which may be arranged into audio clips, and the audio signal 120 may be processed in real time, clip by clip.

[0073] At 504, a first confidence score for the audio signal is calculated using a first machine learning model. For example, the noise detector 102 (see FIG. 2) may calculate a clip confidence for a given clip using the model 206. The noise detector 102 may include a feature extractor 202 that operates on a portion of a clip, called a short clip.

[0074] At 506, when the first confidence score indicates the presence of non-noise, a processed audio signal is generated by processing the audio signal according to a first audio processing process. For example, when the noise detector 102 (see FIG. 1) indicates the absence of noise, the PGC audio processor 106 may perform PGC audio processing on the clip to generate a processed audio signal.

[0075] At 508, when the first confidence score indicates the presence of noise, a second confidence score of the audio signal is calculated using a second machine learning model. For example, the noise discriminator 104 (see FIG. 3) may calculate the noise confidence score 322 using the model 306 applied to the features 320. The second confidence score may be calculated based on the clip, for example, by extracting features from each frame in the clip. Recall that the first confidence score is calculated based on the short clip (see 504).

[0076] At 510, when the second confidence score indicates the presence of UGC noise, the audio signal is processed according to a second audio processing process to generate a processed audio signal. For example, when the noise discriminator 104 (see FIG. 1) indicates the presence of UGC noise, the UGC audio processor 108 may perform UGC audio processing on the clip to generate a processed audio signal.

[0077] At 512, when the second confidence score indicates the presence of PGC noise, a processed audio signal is generated by processing the audio signal according to the first audio processing process. For example, when the noise discriminator 104 (see FIG. 1) indicates the presence of PGC noise, the PGC audio processor 106 may perform PGC audio processing on the clip to generate a processed audio signal.

[0078] The processed audio signal may then be stored in the device's memory, e.g., solid-state memory, transmitted to another device, e.g., for cloud storage, or output as sound, e.g., using a loudspeaker.

[0079] The method 500 may include additional steps corresponding to other functions of the noise classification system 100 described herein, such as calculating a weight, where the weight is used in calculating the second confidence score. As another example, calculating the first confidence score may include calculating an average RMS of the audio signal and using the calculated average RMS when calculating the first confidence score. As another example, calculating the second confidence score may include using the weight when calculating the second plurality of statistics.

[0080] Implementation details

[0081] An embodiment may be implemented in hardware, an executable module stored on a computer-readable medium, or a combination of both, e.g., a programmable logic array. Unless otherwise specified, steps performed by an embodiment need not inherently relate to any particular computer or other apparatus, although in particular embodiments they may. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be convenient to construct more specialized apparatus, e.g., integrated circuits, to perform the required method steps. Thus, an embodiment may be implemented in one or more computer programs executing on one or more programmable computer systems, each comprising at least one processor, at least one data storage system including volatile and nonvolatile memory and / or storage elements, at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.

[0082] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device, such as a solid-state memory or medium, or a magnetic or optical medium, so as to configure and operate the computer when the storage medium or device is read by the computer system to perform the procedures described herein. The system of the present invention may also be considered to be implemented as a computer-readable storage medium configured with a computer program, the storage medium so configured causing the computer system to operate in a specific, predefined manner to perform the functions described herein. Software itself and intangible or transitory signals are excluded, insofar as they are non-patentable subject matter.

[0083] Aspects of the systems described herein can be implemented in a computer-based sound processing network environment suitable for processing digital or digitized audio files. Portions of the adaptive audio system can include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks can be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0084] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. Also, it should be noted that the various functions disclosed herein may be described as hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media in terms of their behavior, register transfers, logical components, and / or other characteristics, using any number of combinations. The computer-readable media on which such formatted data and / or instructions may be embodied include various forms of physical, non-transitory, non-volatile media, such as, but not limited to, optical, magnetic, or semiconductor storage media.

[0085] The above description illustrates various embodiments of the present disclosure, along with examples of how aspects of the disclosure may be implemented. The above examples and embodiments should not be considered to be the only embodiments, but are presented to illustrate the flexibility and advantages of the present disclosure, as defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be adopted without departing from the spirit and scope of the present disclosure, as defined by the claims.

[0086] Various aspects of the present invention can be appreciated from the following enumerated exemplary embodiments (EEE). [EEE1] 1. A computer-implemented method of audio processing, comprising: receiving an audio signal; calculating a first confidence score for the audio signal using a first machine learning model; When the first confidence score indicates the presence of non-noise, generating a processed audio signal by processing the audio signal according to a first audio processing process; When the first confidence score indicates the presence of noise, calculating a second confidence score for the audio signal using a second machine learning model; when the second confidence score indicates the presence of the first type of noise; generating a processed audio signal by processing the audio signal according to a second audio processing process; when the second confidence score indicates the presence of a second type of noise; generating a processed audio signal by processing the audio signal according to a first audio processing process; A method comprising: [EEE2] outputting the processed audio signal as sound via a loudspeaker; The computer-implemented method of EEE1 further comprising: [EEE3] the audio signal includes a plurality of samples, the plurality of samples being arranged in a plurality of frames; A first confidence score is calculated in real time for each short clip, A second confidence score is calculated in real time for each clip, The given short clip and the given clip each contain several frames of the audio signal, and the given short clip contains fewer frames than the given clip; EE1 or EE2. [EEE4] the first audio processing process includes audio processing other than noise reduction; The second audio processing process includes noise reduction. 3. The computer-implemented method of any one of claims 1 to 3. [EEE5] The computer-implemented method of any one of EEE1 to 4, wherein the first type of noise corresponds to user-generated content (UGC) noise and the second type of noise corresponds to professional-generated content (PGC) noise, PGC being audio content created by a professional and UGC being audio content created by a non-professional. [EEE6] the first machine learning model is trained offline using positive training data and negative training data; the positive training data includes training data corresponding to a first type of noise and training data corresponding to a second type of noise; The negative training data includes non-noisy training data. 8. The computer-implemented method of any one of EEE1 to 5. [EEE7] The step of calculating the first confidence score includes: extracting a first plurality of features from the audio signal; classifying the audio signal by inputting the first plurality of features into a first machine learning model; calculating a noise confidence score based on the result of classifying the audio signal; 8. The computer-implemented method of any one of EEE1 to 6, comprising: [EEE8] The computer-implemented method described in EEE7, wherein a first plurality of features are extracted from a short clip including a current frame and a plurality of historical frames, and a noise confidence score for the current frame is obtained by inputting the first plurality of features of the short clip into a first machine learning model. [EEE9] calculating noise confidence scores for a number of frames in a clip; calculating a noise confidence score for the clip as a weighted combination of the noise confidence scores for the frames; 9. The computer-implemented method of any one of EEE7 to 8, further comprising: [EEE10] The step of calculating the noise confidence score comprises: combining the outputs of the weak learners into a weighted sum; converting the weighted sum into a noise confidence score using an inverse exponential function; 10. The computer-implemented method of any one of EEE7 to 9, comprising: [EEE11] The step of calculating the first confidence score includes: Calculating the average root mean square gain of the audio signal further comprising and calculating the noise confidence score based on a result of classifying the audio signal and an average root-mean-square gain of the audio signal. 8. The computer-implemented method according to any one of claims 8 to 10. [EEE12] A noise confidence score and an average root-mean-square gain are associated with the current frame of the audio signal, and the step of calculating the average root-mean-square gain comprises: Calculating an average root-mean-square gain as an average of root-mean-square levels of a number of frames of the short clip including the current frame; calculating a root-mean-square based weight based on a ratio of a first coefficient and a second coefficient, the first coefficient being a product of the root-mean-square gain and a frame weight of the current frame, and the second coefficient being the frame weight of the current frame; Including, The method is: calculating a plurality of noise confidence scores for a plurality of frames in the clip; computing a noisiness weight for the clip as a weighted combination of the plurality of noise confidence scores; calculating a clip confidence score by multiplying the root-mean-square based weight and the noisiness weight; The computer-implemented method of claim 11, further comprising: [EEE13] the first plurality of features includes one or more of a plurality of temporal features, a plurality of spectral features, a plurality of time-frequency features, and a first plurality of statistics; and / or the first plurality of statistics includes one or more of a mean and a standard deviation, the mean being calculated based on one or more of the first plurality of features and the standard deviation being calculated based on one or more of the first plurality of features; 8. A computer-implemented method according to any one of claims 8 to 12. [EEE14] The method further includes calculating a weight based on the noise confidence score, wherein the step of calculating the second confidence score comprises: extracting a second plurality of features from the audio signal, the second plurality of features being extracted over a longer period of time than the first plurality of features; calculating a second plurality of statistics based on the second plurality of features, the second plurality of statistics being weighted according to the weights; classifying the audio signal by inputting the second plurality of features and the second plurality of statistics into a second machine learning model; calculating a second confidence score based on the result of classifying the audio signal; 16. The computer-implemented method of any one of EEE7 to 13, comprising: [EEE15] EEE14. The computer-implemented method of claim 1, wherein a first plurality of features is extracted from a first plurality of frames of a short clip of an audio signal, and a second plurality of features is extracted from a second plurality of frames of the clip of the audio signal. [EEE16] The weight is the frame weight of the current frame, and the step of calculating the frame weight includes: calculating a frame weight by applying a modified sigmoid function to the noise confidence score of the current frame, the frame weight increasing when the noise confidence score exceeds a threshold and the frame weight decreasing when the noise confidence score falls below the threshold; 16. The computer-implemented method of claim 15, comprising: [EEE17] The computer-implemented method of any one of EEE14 to 16, wherein the audio signal comprises a clip, the clip comprising a plurality of frames, the second plurality of features comprising a plurality of frame features and a plurality of statistics, the plurality of frame features being extracted for each frame, and the plurality of statistics being calculated for each clip based on the plurality of frame features. [EEE18] The second machine learning model is trained offline using the positive and negative training data; the positive training data includes training data corresponding to a second type of noise; The negative training data includes training data corresponding to a first type of noise. 18. The computer-implemented method of any one of EEE1 to 17. [EEE19] A non-transitory computer readable medium having stored thereon a computer program which, when executed by a processor, controls an apparatus to perform processes including the methods described in any one of EEE1 to EEE18. [EEE20] 1. An apparatus for audio processing, comprising: the apparatus includes a processor, the processor being configured to control the apparatus to perform processes including the method described in any one of EEE1 to 18; Device.

Claims

1. 1. A computer-implemented method of audio processing, comprising: receiving an audio signal; calculating a first confidence score for the audio signal using a first machine learning model trained to classify the audio signal as non-noise or noise; When the first confidence score indicates the presence of non-noise, generating a processed audio signal by processing the audio signal according to a first audio processing process; when the first confidence score indicates the presence of noise; calculating a second confidence score for the audio signal using a second machine learning model trained to distinguish between a first type of noise and a second type of noise; when the second confidence score indicates the presence of the first type of noise; generating the processed audio signal by processing the audio signal according to a second audio processing process; when the second confidence score indicates the presence of the second type of noise; generating the processed audio signal by processing the audio signal according to the first audio processing process; Including, The first type of noise corresponds to user-generated content (UGC) noise, and the second type of noise corresponds to professional-generated content (PGC) noise, where PGC is audio content created by a professional and UGC is audio content created by a non-professional. Computer-implemented methods.

2. outputting the processed audio signal as sound by a loudspeaker. The computer-implemented method of claim 1 , further comprising:

3. the audio signal includes a plurality of samples, the plurality of samples being arranged in a plurality of frames; the first confidence score is calculated in real time for each short clip; the second confidence score is calculated in real time for each clip; a given short clip and a given clip each include several frames of the audio signal, and the given short clip includes fewer frames than the given clip; The computer-implemented method of claim 1 .

4. the first audio processing process includes audio processing other than noise reduction; the second audio processing process includes noise reduction; The computer-implemented method of claim 1 .

5. the first machine learning model is trained offline using positive training data and negative training data; the positive training data includes training data corresponding to the first type of noise and training data corresponding to the second type of noise; the negative training data includes non-noisy training data; The computer-implemented method of claim 1 .

6. The step of calculating the first confidence score comprises: extracting a first plurality of features from the audio signal; classifying the audio signal by inputting the first plurality of features into the first machine learning model; calculating a noise confidence score based on the result of classifying the audio signal; The computer-implemented method of claim 1 , comprising:

7. 7. The computer-implemented method of claim 6, wherein a first plurality of features are extracted from a short clip including a current frame and a plurality of historical frames, and the noise confidence score of the current frame is obtained by inputting the first plurality of features of the short clip into the first machine learning model.

8. calculating noise confidence scores for a number of frames in a clip; calculating a noise confidence score for the clip as a weighted combination of the noise confidence scores for the plurality of frames; The computer-implemented method of claim 6 further comprising:

9. The step of calculating the noise confidence score comprises: combining the outputs of the weak learners into a weighted sum; converting the weighted sum to the noise confidence score using an inverse exponential function; The computer-implemented method of claim 6 , comprising:

10. The step of calculating the first confidence score comprises: calculating the average root mean square gain of the audio signal; further comprising calculating the noise confidence score includes calculating the noise confidence score based on the result of classifying the audio signal and the average root-mean-square gain of the audio signal. The computer-implemented method of claim 6.

11. The noise confidence score and the average root-mean-square gain are associated with a current frame of the audio signal, and the step of calculating the average root-mean-square gain comprises: calculating the average root-mean-square gain as an average of root-mean-square levels of a plurality of frames of a short clip including the current frame; calculating a root-mean-square based weight based on a ratio of a first coefficient and a second coefficient, the first coefficient being the product of the root-mean-square gain and a frame weight of the current frame, and the second coefficient being the frame weight of the current frame; Including, The method comprises: calculating a plurality of noise confidence scores for a plurality of frames in the clip; computing a noisiness weight for the clip as a weighted combination of the plurality of noise confidence scores; calculating a clip confidence score by multiplying the root-mean-square based weight and the noisiness weight; The computer-implemented method of claim 10 further comprising:

12. the first plurality of features includes one or more of a plurality of temporal features, a plurality of spectral features, a plurality of time-frequency features, and a first plurality of statistics; and / or the first plurality of statistics includes one or more of a mean and a standard deviation, the mean being calculated based on one or more of the first plurality of features, and the standard deviation being calculated based on one or more of the first plurality of features; The computer-implemented method of claim 6.

13. The method further includes calculating a weight based on the noise confidence score, wherein the step of calculating the second confidence score comprises: extracting a second plurality of features from the audio signal, the second plurality of features being extracted over a longer period of time than the first plurality of features; calculating a second plurality of statistics based on the second plurality of features, the second plurality of statistics being weighted according to the weights; classifying the audio signal by inputting the second plurality of features and the second plurality of statistics into the second machine learning model; calculating the second confidence score based on the result of classifying the audio signal; The computer-implemented method of claim 6 , comprising:

14. 14. The computer-implemented method of claim 13, wherein the first plurality of features are extracted from a first plurality of frames of a short clip of the audio signal and the second plurality of features are extracted from a second plurality of frames of the clip of the audio signal.

15. The weight is a frame weight of a current frame, and the step of calculating the frame weight comprises: calculating the frame weight by applying a modified sigmoid function to the noise confidence score of the current frame, wherein the frame weight is increased when the noise confidence score exceeds a threshold and the frame weight is decreased when the noise confidence score is below the threshold; The computer-implemented method of claim 13, comprising:

16. 14. The computer-implemented method of claim 13, wherein the audio signal includes a clip, the clip includes a plurality of frames, the second plurality of features includes a plurality of frame features and a plurality of statistics, the plurality of frame features are extracted for each frame, and the plurality of statistics are calculated for each clip based on the plurality of frame features.

17. the second machine learning model is trained offline using positive training data and negative training data; the positive training data includes training data corresponding to the second type of noise; the negative training data includes training data corresponding to the first type of noise; The computer-implemented method of claim 1 .

18. A non-transitory computer readable medium having stored thereon a computer program which, when executed by a processor, controls an apparatus to perform processes including the method of any one of claims 1 to 17.

19. 1. An apparatus for audio processing, comprising: The device includes a processor, the processor being configured to control the device to perform a process comprising the method of any one of claims 1 to 17. Device.

Citation Information

Patent Citations

  • Apparatus and method for audio classification and processing

    JP2016519784A

  • Voice activity detector (VAD)-based multiple-microphone acoustic noise suppression

    US20160155434A1

  • Ambient noise reduction arrangements

    US20200020312A1

  • Wearable Audio Recorder and Retrieval Software Applications

    US20200125316A1

  • Method and device for voice recognition in environments with fluctuating noise levels

    US6859773B2