Media type-based reverberation removal

By classifying audio signals into media types and selectively applying reverberation removal, the solution enhances speech clarity while preserving audio quality in mixed content, addressing echo suppression challenges in user-generated audio.

JP7877347B2Active Publication Date: 2026-06-22DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2022-03-10
Publication Date
2026-06-22

AI Technical Summary

Technical Problem

Existing audio devices struggle to effectively suppress echoes, particularly in user-generated content that contains a mixture of various media types, leading to reduced audio quality and speech intelligibility.

Method used

Classify input audio signals into media types such as speech, music, or speech over music, and selectively apply reverberation removal only to audio signals classified as speech, while suppressing it for other types to maintain audio quality.

Benefits of technology

Improves speech clarity by selectively removing reverberation from speech signals and avoids quality degradation in non-speech content, enhancing the overall audio experience, especially in user-generated content like podcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007877347000001
    Figure 0007877347000001
  • Figure 0007877347000002
    Figure 0007877347000002
  • Figure 0007877347000003
    Figure 0007877347000003
Patent Text Reader

Abstract

A method of dereverberation suppression can include receiving an input audio signal. The method can include classifying a media type of the input audio signal as one of a group including at least (1) speech, (2) music, or (3) speech over music. The method can include determining whether to perform dereverberation on the input audio signal based at least on a determination that the media type of the input audio signal is classified as speech. The method can include generating an output audio signal by performing dereverberation on the input audio signal in response to a determination that dereverberation should be performed on the input audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Related Applications] This application claims priority from the following prior applications. This application claims priority from International Patent Application No. PCT / CN20201 / 080314 filed on March 11, 2021, U.S. Provisional Patent Application No. 63 / 180,710 filed on April 28, 2021, and European Patent Application No. 21174289.5 filed on May 18, 2021.

[0002] [Technical Field] The present disclosure relates to systems, methods, and media for echo cancellation. The present disclosure further relates to systems, methods, and media for classifying input audio signals.

Background Art

[0003] Audio devices such as headphones and speakers are widely popular. People frequently listen to audio content (e.g., podcasts, radio shows, TV shows, music videos, etc.) that contains various types of media content mixed together, such as speech, music, and speech over music. Such audio content may contain echoes. It can be difficult to suppress echoes in audio content, especially in user-generated audio content that contains a mixture of various types of media content.

Summary of the Invention

[0004] [Annotations and Terms] Throughout this disclosure, including the claims, the terms “speaker,” “loudspeaker,” and “audio playback transducer” are used synonymously to refer to any sound-emitting transducer (or set of transducers) driven by a single speaker feed. A typical headphone set includes two speakers. The speakers are implemented to include multiple transducers (e.g., a woofer and a tweeter), which are driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds may undergo different processing in different circuit branches coupled to different transducers.

[0005] Throughout this disclosure, including the claims, the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to indicate that an operation is performed directly on the signal or data, or on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or post-processing before the operation is performed).

[0006] Throughout this disclosure, including the claims, the expression “system” is used broadly to describe a device, system, or subsystem. For example, a subsystem that implements a decoder may be called a decoder system, and a system containing such a subsystem (for example, a system that generates X output signals in response to multiple inputs, of which the subsystem generates M inputs and the other XM inputs are received from an external source) may also be called a decoder system.

[0007] Throughout this disclosure, including the claims, the term “processor” is used broadly to describe a system or device that is programmable (by software or firmware) or otherwise configurable to perform operations on data (e.g., audio or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other speech data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0008] Throughout this disclosure, including the claims, the terms “connect” or “connected” are used to mean direct or indirect connection. Therefore, when the first device connects to the second device, such connection may be through a direct connection or through an indirect connection via other devices and connections.

[0009] Throughout this disclosure, including the claims, the term “classifier” is used generally to refer to an algorithm that predicts the class of an input. For example, as used herein, an audio signal may be classified as relating to a particular media type, such as speech, music, or speech on music. It should be understood that various types of classifiers may be used to implement the techniques described herein, including decision trees, Ada-boost, XG-boost, RandomForest, Generalized Method of Moments (GMM), Hidden Markov Models (HMM), NaiveBayes, and / or various types of neural networks (e.g., convolutional neural networks (CNN), deep neural networks (DNN), recurrent neural networks (RNN), long-term short-term memory (LSTM), gated recurrent units (GRU), etc.).

[0010] [overview] At least some aspects of this disclosure may be implemented by methods. Some methods may include the step of receiving an input audio signal. Some such methods may include the step of classifying the media type of the input audio signal as one of the groups including at least (1) speech, (2) music, or (3) speech over music. Some such methods may include the step of deciding whether to perform de-reverberation on the input audio signal, at least based on the decision that the media type of the input audio signal is classified as speech. Some such methods may include the step of generating an output audio signal by performing de-reverberation on the input audio signal, in response to the decision that de-reverberation should be performed on the input audio signal.

[0011] In some examples, the method may include determining the degree of reverberation in an input audio signal, where the decision of whether or not to perform reverberation removal on the input audio signal can be made based on the degree of reverberation. In some examples, the degree of reverberation can be based on at least one of 1) reverberation time (RT60), or 2) direct-to-reverberation ratio (DRR), or an estimate of the degree of diffusion. In some examples, determining the degree of reverberation may include calculating the two-dimensional acoustic modulation frequency spectrum of the input audio signal, where the degree of reverberation may be based on the amount of energy in the high-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum. In some examples, determining the degree of reverberation may include calculating at least one of the following: 1) the ratio of the energy in the high-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy across all modulation frequencies of the two-dimensional acoustic modulation frequency spectrum, or 2) the ratio of the energy in the high-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy in the low-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum.

[0012] In some examples, the method may include a step of determining whether to perform reverberation removal on the input audio signal based on the determination that the degree of reverberation exceeds a threshold.

[0013] In some examples, the method may include a step of classifying the media type of the input audio signal by separating the input audio signal into two or more spatial components. According to some implementations, the two or more spatial components may include a center channel and side channels. In some examples, the method may further include a step of calculating the power of the side channels and classifying the side channels in response to determining that the power of the side channels exceeds a threshold. According to other implementations, the two or more spatial components may include a diffuse component and a direct component. In some examples, the step of classifying the media type of the input audio signal is: This may include the step of classifying each of two or more spatial components as one of (1) speech, (2) music, or (3) speech over music. The media type of the input audio signal is classified by combining the classifications of two or more spatial components. In some examples, the input audio signal may be separated into two or more spatial components in response to the determination that the input audio signal contains stereo audio.

[0014] In some examples, the method may include classifying the media type of the input audio signal by separating the input audio signal into vocal and non-vocal components. In some examples, the input audio signal may be separated into vocal and non-vocal components in response to determining that the input audio signal contains a single audio channel. In some examples, the method may further include classifying the vocal component as either 1) speech or 2) non-speech. The method may further include classifying the non-vocal component as either 1) music or 2) non-music. In some examples, the media type of the input audio signal may be classified by combining the classification of the vocal component and the classification of the non-vocal component.

[0015] In some cases, the step of deciding whether to perform reverberation removal on the input audio signal may be based on the classification of a second input audio signal that precedes the input audio signal.

[0016] In some examples, the method may include the step of receiving a third input audio signal. The method may further include determining that no de-reverberation should be performed on the third input audio signal. In response to determining that no de-reverberation should be performed on the third input audio signal, the method may include prohibiting the de-reverberation algorithm from being performed on the third input audio signal. In some examples, the determination that no de-reverberation should be performed on the third input audio signal may be based at least in part on the classification of the media type of the third input audio signal. In some examples, the classification of the third input audio signal may be one of 1) music, or 2) speech over music. In some examples, the determination that no de-reverberation should be performed on the third input audio signal may be based at least in part on the determination that the degree of reverberation in the third input audio signal is below a threshold.

[0017] According to another aspect of this disclosure, a method for classifying an input audio signal as one of at least two media types, The steps include receiving the input audio signal, The steps include: separating the input audio signal into two or more spatial components, A step of classifying each of two or more spatial components as one of at least two media types, wherein the media type of the input audio signal is classified by combining the classifications of each of the two or more spatial components. A method including this is provided.

[0018] In some examples, two or more spatial components include a center channel and side channels, and the method is Steps for calculating the power of the side channel, and Steps for classifying the side channel in response to determining that the power of the side channel exceeds a threshold, and further include.

[0019] In some examples, two or more spatial components include a diffuse component and a direct component.

[0020] In some examples, an input audio signal is separated into two or more spatial components in response to determining that the input audio signal includes stereo audio.

[0021] In some examples, classifying the media type of an input audio signal can include separating the input audio signal into a vocal component and a non-vocal component. In some examples, in response to determining that the input audio signal includes a single audio channel, the input audio signal is separated into a vocal component and a non-vocal component. In some examples, the step of classifying the media type of the input audio signal classifying the vocal component as one of (1) speech or (2) non-speech, and classifying the non-vocal component as one of (1) music or (2) non-music, and include, The media type of the input audio signal is classified by combining the classification of the vocal component and the classification of the non-vocal component.

[0022] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored in one or more non-transitory media. Such non-transitory media may include, but are not limited to, memory devices as described herein, including random access memory (RAM), read-only memory (ROM), etc. Accordingly, various novel aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory media storing software.

[0023] At least some aspects of the present disclosure may be implemented via a device. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the device is or includes an audio processing system having an interface system and a control system. The control system may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic element, discrete gate or transistor logic, discrete hardware component, or a combination thereof.

[0024] The present disclosure provides various technical advantages. For example, the clarity of speech can be improved by selectively performing reverberation removal on a specific type of input audio signal (e.g., an input audio signal classified as speech). Further, by prohibiting reverberation removal on other types of input audio signals (e.g., input audio signals classified as music, speech over music, etc.), adverse effects of reverberation removal such as a reduction in audio quality can be avoided for audio signals that do not require improvement in speech clarity. The technical advantages of the present disclosure can be particularly useful for user-generated content such as podcasts.

[0025] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. The relative dimensions in the following drawings may not be drawn to scale.

Brief Description of the Drawings

[0026] [Figure 1A] shows an example of an audio signal including reverberation. [Figure 1B] shows an example of an audio signal including reverberation.

[0027] [Figure 2] The following are block diagrams of example systems that perform reverberation removal based on media type, using several implementations.

[0028] [Figure 3] This section presents examples of processes that perform reverberation removal based on media type, using several implementations.

[0029] [Figure 4] This section presents examples of spatial separation of input audio signals using several implementations.

[0030] [Figure 5] This section shows examples of source separation processes for input audio signals using several implementations.

[0031] [Figure 6] This section presents examples of processes that determine the degree of reverberation, using several implementations.

[0032] [Figure 7A] An exemplary graph of the two-dimensional acoustic modulation frequency spectrum of an exemplary audio signal is shown. [Figure 7B] An exemplary graph of the two-dimensional acoustic modulation frequency spectrum of an exemplary audio signal is shown. [Figure 7C] An exemplary graph of the two-dimensional acoustic modulation frequency spectrum of an exemplary audio signal is shown. [Figure 7D] An exemplary graph of the two-dimensional acoustic modulation frequency spectrum of an exemplary audio signal is shown.

[0033] [Figure 8] This block diagram shows examples of components for devices that can implement various aspects of this disclosure.

[0034] Similar reference numerals and symbols in various drawings indicate the same elements. [Modes for carrying out the invention]

[0035] Reverberation occurs when an audio signal is distorted by various reflections from different surfaces (e.g., walls, ceilings, floors, furniture, etc.). Reverberation can significantly affect sound quality and speech intelligibility. Therefore, reverberation removal from audio signals, including speech, can improve speech intelligibility.

[0036] Sound reaching a receiver (e.g., a human listener, a microphone, etc.) includes direct sound, which is the sound directly from a sound source without reflections, and reverberation, which is the sound reflected from various surfaces in the environment. Reverberation consists of early and late reflections. Early reflections arrive at the receiver immediately after or simultaneously with the direct sound and may therefore be partially incorporated into the direct sound. The integration of early and direct reflections produces a spectral coloration effect that contributes to the perceived sound quality. Late reflections arrive at the receiver after the early reflections (e.g., 50-80 milliseconds or more after the direct sound). Late reflections can negatively affect speech clarity. Therefore, de-reverberation can be applied to audio signals to reduce the impact of late reflections present in the audio signal and improve speech clarity.

[0037] Figure 1A shows an example of an acoustic impulse response in a reverberant environment. As illustrated, the early reflection 102 may arrive at the receiver simultaneously with or immediately after the direct sound. In contrast, the late reflection 104 may arrive at the receiver after the early reflection 102.

[0038] Figure 1B shows an example of a time-domain input audio signal 152 and the corresponding spectrogram 154. As shown in spectrogram 154, early reflections can cause changes in spectrogram 154, as shown by the spectral coloration 156.

[0039] Reverb removal can reduce audio quality, for example, by reducing perceived loudness or altering spectral color effects. This reduced audio quality can be particularly detrimental when applied to audio signals that primarily contain music or speech over music. For example, the audio quality of an audio signal that primarily contains music or speech over music may decrease without improving the intelligibility of the speech. As a more specific example, reverb removal may be suitable for processing low-quality speech content, such as user-generated content captured in long-distance use cases. Continuing this specific example, user-generated content, such as podcasts, may contain both low-quality speech content and professionally produced music content. In some cases, professionally produced music content may contain artificial reverb. In such cases, applying reverb removal to mixed media content (e.g., including low-quality speech content and professionally produced music content with artificial reverb) may result in excessive reverb suppression and a decrease in audio quality.

[0040] In some implementations, reverberation removal can be performed on an input audio signal based on the identification of the media type associated with the input audio signal. For example, the input audio signal can be analyzed to determine whether it is 1) speech, 2) music, 3) speech over music, or 4) other. Examples of speech over music content include podcast intros and outros, and television show intros and outros.

[0041] In some embodiments, de-reverberation may be performed on input audio signals identified as speech or primarily speech. Conversely, de-reverberation may be suppressed on input audio signals identified as music, primarily music, speech over music, or speech over music. By suppressing de-reverberation for media types that are not speech or primarily speech, it is possible to perform de-reverberation on input audio signals that would substantially benefit from it (for example, because the input audio signal primarily contains speech), and to prevent the degradation of sound quality caused by de-reverberation when such reverberation is not necessary to improve the clarity of speech.

[0042] In some implementations, an input audio signal can be classified into one of the following categories: 1) speech, 2) music, 3) speech on music, or 4) other various techniques. As used herein, "other" can refer to noise, sound effects, speech on sound effects, etc. For example, in some implementations, an input audio signal can be classified by separating it into two or more spatial components and classifying each spatial component as one of the following: 1) speech, 2) music, 3) speech on music, or 4) other. Continuing this example, in some implementations, the classifications of each spatial component can be combined to generate an aggregate classification of the input audio signal. As another example, in some implementations, an input audio signal can be classified by separating it into vocal and non-vocal components. The vocal component can be classified into either 1) speech or 2) non-speech. The non-vocal component can be classified into either 1) music or 2) non-music. Continuing this example, in some implementations, the classifications of the vocal and non-vocal components can be combined to generate an aggregate classification of the input audio signal. This disclosure describes several methods for classification in the context of reverberation suppression methods, but the methods of classification of the present invention can be used in other contexts. In particular, this disclosure describes a method for classifying a relevant input audio signal as one of at least two media types, The steps include receiving the input audio signal, The steps include separating the input audio signal into two or more spatial components, A step of classifying each of the two or more spatial components as at least one of two media types, wherein the media type of the input audio signal is classified by combining the classifications of each of the two or more spatial components, This relates to methods that include [specific methods].

[0043] In some implementations, the input audio signal classified as speech can be further analyzed to determine the amount of reverberation present in the input audio signal. In some such implementations, reverberation removal can be performed on input audio signals identified as having reverberation exceeding a threshold amount. The amount of reverberation can be identified using the Direct to Reverberant Ratio (DRR), and / or the Reverberation Time (RT) up to 60 dB (e.g., RT60), and / or diffusion measurements and / or other appropriate measurements of reverberation. Note that the amount of reverberation is a function of DRR, where the amount of reverberation increases as the DRR value decreases and decreases as the DRR value increases.

[0044] As an addition or alternative, some implementations allow de-reverberation to be performed on the input audio signal based on the classification of the media type of the preceding audio signal. In some implementations, the preceding audio signal can be a preceding frame or part of audio content preceding the input audio signal. In some implementations, the classification of the input audio signal can be adjusted based on the classification of the preceding audio signal so that the classification of adjacent audio signals is effectively smoothed. The adjustment can be performed based on the confidence level of each classification. Deciding whether to perform de-reverberation on the input audio signal based at least partially on the classification of the preceding audio signal can prevent the reverberation from being applied intermittently, thereby improving the overall audio quality.

[0045] In some implementations, reverberation removal can be performed on an input audio signal using various techniques. For example, in some implementations, reverberation removal can be performed based on the amplitude modulation of the input audio signal in various frequency bands. In a more specific example, in some embodiments, a time-domain audio signal can be converted to a frequency-domain signal. Continuing this more specific example, the frequency-domain signal can be divided into multiple subbands, for example, by applying a filter bank to the frequency-domain signal. Continuing this more specific example, an amplitude modulation value can be determined for each subband, and a band-pass filter can be applied to the amplitude modulation value. In some implementations, the band-pass filter value can be selected based on the rhythm of human speech, for example, such that the center frequency of the band-pass filter exceeds the rhythm of human speech (e.g., in the range of 10-20 Hz, approximately 15 Hz, etc.). Further continuing this specific example, the gain of each subband can be determined based on a function of the amplitude-modulated signal value and the band-pass filtered amplitude-modulated value. The gain can then be applied to each subband. In some embodiments, reverberation removal can be performed using the techniques described in U.S. Patent No. 9,520,140, ​​which is incorporated herein by reference in whole.

[0046] As another example, some implementations can perform de-reverberation by estimating the de-reverberated signal using deep neural networks, weighted prediction error methods, variance-normalized delayed linear prediction methods, single-channel linear filters, or multi-channel linear filters. As yet another example, some implementations can perform de-reverberation by estimating the room response and performing a deconvolution operation on the input audio signal based on the room response.

[0047] It should be noted that the techniques described herein for media type-based reverberation removal can be applied to a variety of types or forms of audio content, including but not limited to podcasts, radio programs, audio content related to video conferencing, and audio content related to television programs or films. The audio content may be live or pre-recorded.

[0048] Figure 2 shows a block diagram of an example of system 200, which performs reverberation removal based on identified media types associated with the input audio signal, with several implementations.

[0049] As shown in the figure, system 200 may include a media type classifier 202. The media type classifier 202 can receive an input audio signal. In some implementations, the media type classifier 202 can classify the input audio signal as 1) voice, 2) music, 3) speech with music, or 4) other.

[0050] In some implementations, in response to determining that the input audio signal is not speech, or primarily not speech (for example, determining that the input audio signal is music, musical speech, or something else), the media type classifier 202 can allow the input audio signal to pass through without leading it to the reverberation analyzer 204. Conversely, in response to determining that the input audio signal is speech, or primarily speech, the media type classifier 202 can pass the input audio signal to the reverberation analyzer 204.

[0051] In some embodiments, the reverberation analyzer 204 can determine the degree of reverberation present in the input audio signal. In some embodiments, the reverberation analyzer 204 can determine that reverberation removal should be performed on the input audio signal in response to determining that the degree of reverberation exceeds a threshold. That is, in some embodiments, the reverberation analyzer 204 can further guide the input audio signal to the reverberation removal component 206 in response to determining that the input audio signal is sufficiently reverberant. In contrast, in response to determining that the input audio signal is sufficiently non-reverberant (for example, the input audio signal contains a relatively "dry" sound), the reverberation analyzer 204 can allow the input audio signal to pass through without guiding it to the reverberation removal component 206, effectively suppressing reverberation removal from the input audio signal.

[0052] The reverberation removal component 206 can take an input audio signal that has been determined to have reverberation exceeding a threshold as input and generate a reverberation-removed audio signal. It should be understood that the reverberation removal component 206 can perform any appropriate reverberation suppression technique.

[0053] In some implementations, the media type classifier 202 classifies the media type of the input audio signal based on either the spatial separation of the components of the input audio signal or the music source separation of the components of the input audio signal, or both.

[0054] For example, in some implementations, the media type classifier 202 may include a spatial information separator 208. The spatial information separator 208 can separate an input audio signal into two or more spatial components. Examples of two or more spatial components include a direct component and a diffuse component, a side channel and a center channel, etc. In some embodiments, the spatial information separator 208 can classify the media type of the input audio signal by classifying each of the two or more spatial components separately. In some embodiments, the spatial information separator 208 can then generate a classification of the input audio signal by combining the classifications of each of the two or more components, for example, by using a decision fusion algorithm. Examples of decision fusion algorithms that may be used to combine the classifications of each of the two or more components include Bayesian analysis, the Dempster-Shafer algorithm, and fuzzy logic algorithms. The technique for classifying media types based on spatial source separation is shown in Figure 4 and will be described below in relation to Figure 4.

[0055] As another example, in some implementations, the media type classifier 202 may include a music source separator 210. The music source separator 210 can separate the input audio signal into vocal and non-vocal components. In some implementations, the music source separator 210 can classify the vocal components as either 1) speech or 2) non-speech. In some implementations, the music source separator 210 can classify the vocal components as either 1) music or 2) non-music. In some implementations, based on the classification of vocal and non-vocal components, the music source separator 210 can generate the input audio signal as either 1) speech, 2) music, 3) speech over music, or 4) other: for example, in some implementations, the music source separator 210 can combine the classification of vocal and non-vocal components (e.g., using a decision fusion algorithm). Examples of decision fusion algorithms that can be used to combine the classifications of two or more components include Bayesian analysis, the Dempster-Shafer algorithm, and fuzzy logic algorithms.

[0056] In some implementations, the media type classifier 202 can determine whether to use the spatial information separator 208 or the music source separator 210 to classify the media type of the input audio signal. For example, in response to a determination that the input audio signal is a stereo audio signal, the media type classifier 202 may determine that the media type should be classified using the spatial information separator 208. In another example, in response to a determination that the input audio signal is a monaural channel audio signal, the media type classifier 202 may determine that the media type should be classified using the music source separator 210.

[0057] In the example in Figure 2, the media type classifier 202 is used in the context of system 200 for performing reverberation removal. It is emphasized that the media type classifier 202 may be used as a standalone system or in other audio processing systems.

[0058] Figure 3 shows an example of a process 300 for performing reverberation removal on an input audio signal based on media type classification, according to several implementations. In some implementations, blocks of process 300 may be performed by a device or instrument (e.g., instrument 200 in Figure 2). Note that in some implementations, blocks of process 300 may be performed in an order not shown in Figure 3, and / or one or more blocks of process 300 may be performed substantially in parallel. Furthermore, note that in some implementations, one or more blocks of process 300 may be omitted.

[0059] In 302, the process 300 can receive an input audio signal. The input audio signal may be recorded or live content. The input audio signal may include various types of audio content, such as speech, music, or speech over music. Exemplary types of audio content may include podcasts, radio programs, television programs, or audio content related to movies.

[0060] In 304, the process 300 can classify the media type of the input audio signal. For example, in some implementations, the process 300 can classify the input audio signal as one of the following: 1) speech, 2) music, 3) speech over music, or 4) other.

[0061] In some implementations, the process 300 can classify the media type of the input audio signal based on the separation of its spatial components. For example, in some implementations, the process 300 can separate the input audio signal into two or more spatial components, such as direct and diffuse components, or side channels and center channels. In some implementations, the process 300 can then classify the media type of the audio content within each spatial component. In some implementations, the process 300 can then classify the input audio signal by combining the classifications of each spatial component. Note that more detailed techniques for classifying the media type of the input audio signal based on spatial separation are shown and explained below in relation to Figure 4.

[0062] As an addition or alternative, in some implementations, the process 300 can classify the media type of the input audio signal based on music source separation of the input audio signal. For example, in some embodiments, the process 300 can separate the input audio signal into vocal and non-vocal components. In some implementations, the process 300 can then classify the media type of the audio content in each of the vocal and non-vocal components. In some implementations, the process 300 can then classify the input audio signal by combining the classifications of the vocal and non-vocal components. Note that more detailed techniques for classifying the media type of the input audio signal based on music source separation are shown and described below in relation to Figure 5.

[0063] In block 306, process 300 can decide whether or not to analyze the reverberation characteristics of the input audio signal. In some implementations, process 300 can decide whether or not to analyze the reverberation characteristics based on the media type classification of the input audio signal determined in block 304. For example, in some implementations, process 300 can decide to analyze the reverberation characteristics in response to the determination that the media type classification of the input audio signal is speech (YES in 306). Conversely, in some implementations, process 300 can decide not to analyze the reverberation characteristics in response to the determination that the media type classification of the input audio signal is not speech (for example, the media type classification is music, musical speech, or something else) (NO in 306).

[0064] If, in step 306, process 300 decides not to analyze the reverberation characteristics ("NO" in step 306), process 300 can be terminated in step 314.

[0065] Conversely, if in 306 the process 300 determines that the reverberation characteristics should be analyzed (YES in 306), then in 308 the process 300 can determine the degree of reverberation in the input audio signal.

[0066] In some implementations, the degree of reverberation can be calculated using the RT60 metric and / or DRR metric associated with the input audio signal.

[0067] As an addition or alternative, in some implementations, the process 300 can determine the degree of reverberation in the input audio signal based on spectrogram information. For example, in some implementations, the process 300 can determine the degree of reverberation based on the energy of the input audio signal at various modulation frequencies. In particular, since non-reverberating speech tends to have a peak in modulation frequency at relatively low modulation frequencies (e.g., 3 Hz, 4 Hz, etc.), and reverberating speech tends to have substantial energy at higher modulation frequencies (e.g., 10 Hz, 20 Hz, 50 Hz, etc.), the process 300 can determine the degree of reverberation in the input audio signal based on the energy of the input audio signal at relatively high modulation frequencies (e.g., above 10 Hz, above 20 Hz, etc.).

[0068] Further detailed techniques for determining the degree of reverberation based on spectrogram information are shown and explained below in relation to Figure 7.

[0069] At 310, process 300 can decide whether or not to perform reverberation removal on the input audio signal. In some implementations, process 300 can decide whether or not to perform reverberation removal based on the degree of reverberation determined in block 308. For example, in some implementations, process 300 can decide to perform reverberation removal in response to the determination that the degree of reverberation exceeds a threshold (YES at 310). As another example, in some implementations, process 300 can decide not to perform reverberation removal in response to the determination that the degree of reverberation falls below a threshold (NO at 310).

[0070] In some embodiments, the process 300 may, additionally or alternatively, determine whether to perform reverberation removal on the input audio signal based on the media type classification of the preceding audio signal. The preceding audio signal may correspond to a preceding frame or portion of audio content that precedes the input audio signal. Note that the frame or portion of audio content may have any appropriate duration, such as 10 milliseconds or 20 milliseconds.

[0071] In some implementations, process 300 can determine whether to perform reverberation removal on the input audio signal based on the media type classification of the preceding audio signal by adjusting the media type classification (determined in block 304, for example) based on the classification of the preceding audio signal. For example, in some implementations, the media type classification of the input audio signal can be adjusted based on the confidence level of the media type classification of the input audio signal and / or the confidence level of the media type classification of the preceding audio signal. As a more specific example, if the media type classification of the preceding audio signal is associated with a relatively high confidence level (e.g., over 70%, over 80%) and the media type classification of the input audio signal is associated with a relatively low confidence level (e.g., less than 30%, less than 20%), the media type classification of the input audio signal can be adjusted or changed to match the media type classification of the preceding audio signal. Note that the adjustment of the media type classification of the input audio signal can be done once or multiple times. For example, the media type classification can be adjusted before analyzing the reverberation characteristics in block 306. As another example, the media type classification can be adjusted after determining the degree of reverberation in block 308.

[0072] At 310, process 300 decides that reverberation characteristics will not be performed (NO at 306), and process 300 can terminate at 314.

[0073] Conversely, if in 310 the process 300 determines that de-reverberation should be performed (YES in 310), the process 300 can generate an output audio signal by performing de-reverberation on the input audio signal. For example, in some implementations, de-reverberation can be performed based on amplitude modulation of the input audio signal in various frequency bands. As a more specific example, de-reverberation can be performed using the technique described in U.S. Patent No. 9,520,140, ​​which is incorporated herein by reference in its entirety. As another example, in some implementations, de-reverberation can be performed by estimating the de-reverberated signal using a deep neural network, a multi-channel linear filter, etc. As yet another example, in some implementations, de-reverberation can be performed by estimating the room response and performing a deconvolution operation on the input audio signal based on the room response.

[0074] After that, process 300 can be terminated at 314.

[0075] It should be noted that after termination in 314, the output audio signal can be presented through, for example, speakers, headphones, etc. In some implementations, if de-reverbing in block 312 is not performed (for example, because the input audio signal is classified as music, musical speech, or other non-speech content), the output audio signal may be the original input audio signal. Alternatively, in some implementations, if de-reverbing in block 312 is not performed (for example, because the input audio signal is classified as speech, musical speech, or other non-speech content), a different de-reverbing technique than that applied in 312 may be applied to the original input audio signal.

[0076] In some implementations, if de-reverbance is performed in block 312, the output audio signal may correspond to the de-reverbanced input audio signal.

[0077] In some implementations, the media type of an input audio signal may be classified based on the spatial separation of the components of the input audio signal. Examples of components include direct and diffuse components, center channel and side channels, etc. In some implementations, each spatial component can be classified into one of the following: 1) speech, 2) music, 3) speech over music, or 4) other. In some implementations, the input audio signal may be classified based on a combination of the classifications of each spatial component. In some implementations, two or more spatial components may be identified based on the upmix of the input audio signal. In some implementations, the media type classification of an input audio signal based on the spatial separation of its components may be performed in response to the determination that the input audio signal is a multi-channel audio signal (e.g., stereo audio signal, 5.1 audio signal, 7.1 audio signal, etc.).

[0078] Figure 4 shows an example of a process 400 for media type classification of an input audio signal based on the spatial separation of components of the input audio signal, according to several implementations. Note that the blocks of process 400 may be executed in various orders not shown in Figure 4, and / or in some implementations, two or more blocks of process 400 may be executed substantially in parallel. Note that, as an addition or alternative, in some implementations, one or more blocks of process 400 may be omitted.

[0079] Process 400 can be initiated by receiving the input audio signal at 402. In some implementations, the input audio signal may contain two or more audio channels.

[0080] In 404, the processor 400 can upmix the input audio signal to increase the number of audio channels associated with the input audio signal. The processor 400 can use various types of upmixing. For example, in some implementations, the processor 400 can perform upmixing techniques such as shuffling from left / right to middle / side. As another example, in some implementations, the processor 400 can perform upmixing techniques that convert a stereo audio input into multi-channel content such as 5.1 or 7.1.

[0081] In some implementations, an input audio signal can be separated into direct and diffuse components. For example, some implementations can distinguish between direct and diffuse components based on inter-channel coherence. More specifically, some implementations can distinguish between direct and diffuse components based on coherence matrix analysis.

[0082] In 406, the process 400 can obtain side and center channels from the upmixed input audio signal. For example, if the upmixed input audio signal corresponds to shuffled middle / side channels, the side channels can correspond to shuffled side channels and the center channel can correspond to shuffled middle channels. As another example, if the upmixed input audio signal corresponds to a multi-channel upmix (e.g., 5.1, 7.1, etc.), the center channel can be extracted directly from the upmixed audio signal, and the side channels can be obtained by downmixing the left / right pairs (e.g., Left / Right, Left Surround / Right Surround, etc.).

[0083] When an input audio signal is split into direct and diffuse components, the center channel may correspond to the direct component, and the side channels may correspond to the diffuse component.

[0084] In 408, process 400 can determine whether the power of the side channel exceeds a threshold. Examples of thresholds may be -65 dB, -68 dBFS, -70 dBFS, -72 dBFS, etc., relative to full scale (dBFS).

[0085] If it is determined in 408 that the power of the side channel does not exceed the threshold (NO in 408), then process 400 can proceed to block 412.

[0086] Conversely, if in 408 it is determined that the power of the side channel exceeds a threshold ("YES" in 408), then in 410 the process 400 can classify the side channel as one of the following: 1) speech, 2) music, 3) speech over music, or 4) other. In some implementations, the classification of the side channel may be associated with a confidence level. Examples of classifiers that can be used to classify side channels include k nearest neighbors, case-based inference, decision trees, NaiveBayes, and / or various types of neural networks (e.g., convolutional neural networks (CNNs)).

[0087] In 412, process 400 can classify the center channel as one of the following: 1) speech, 2) music, 3) speech over music, or 4) other. In some implementations, the classification of the center channel may be associated with a confidence level. Examples of classifiers that can be used to classify the center channel include k nearest neighbors, case-based inference, decision trees, NaiveBayes, and / or various types of neural networks (e.g., convolutional neural networks (CNNs)).

[0088] In 414, the process 400 can classify the input audio signal as one of the following: 1) speech, 2) music, 3) speech over music, or 4) other, by combining side channel classification (if present) and center channel classification.

[0089] For example, in some implementations, side-channel classification and center-channel classification can be combined using decision fusion algorithms. Examples of decision fusion algorithms that can be used to combine the classifications of two or more components include Bayesian analysis, the Dempster-Shafer algorithm, and fuzzy logic algorithms.

[0090] As another example, in some implementations, in response to the side channels being classified as music, musical speech, or something else, the input audio signal can be classified as "non-speech," regardless of the center channel's classification. More specifically, if the center channel is classified as "speech" and the side channels are classified as "music," the input audio signal can be classified as musical speech.

[0091] As yet another example, in some implementations, side-channel and center-channel classifications can be combined based on the confidence levels associated with each classification. More specifically, in some implementations, side-channel and center-channel classifications can be combined such that classifications of spatial components associated with higher confidence levels are weighted more heavily in the combination. For example, if the center channel is classified as "speech" with a relatively high confidence level (e.g., 70% or higher, 80% or higher) and the side channels are classified as "music," "musical speech," or "other" with a relatively low confidence level (e.g., less than 30%, less than 20%), the input audio signal may be classified as speech. As yet another specific example, if the center channel is classified as "speech" with a relatively low confidence level (e.g., less than 30%, less than 20%) and the side channels are classified as "music," "musical speech," or "other" with a relatively high confidence level (e.g., 70% or higher, 80% or higher), the input audio signal may be classified as "musical speech" or "other."

[0092] Note that if the side channels are not classified (for example, because the power of the side channels is below the threshold determined in block 408), the classification of the input audio signal may correspond to the classification of the center channel.

[0093] In some embodiments, the input audio signal may be classified based on music source separation into vocal and non-vocal components. The vocal component can be classified as speech or non-speech. The non-vocal component can be classified as music or non-music. In some implementations, the input audio signal can then be generated as 1) speech, 2) music, 3) speech over music, or 4) other, based on a combination of the vocal and non-vocal component classifications. In some implementations, the input audio signal can be classified using music source separation of the input audio signal in response to determining that the input audio signal is a monaural channel audio signal. Alternatively, in some implementations, the input audio signal can be classified using music source separation in addition to the classification of the input audio signal based on spatial separation of components.

[0094] Figure 5 shows an example of a process 500 for classifying an input audio signal based on music source separation, according to several implementations. Note that the blocks of process 500 may be executed in various orders not shown in Figure 5, and / or in some implementations, two or more blocks of process 500 may be executed substantially in parallel. Note that, as an addition or alternative, in some implementations, one or more blocks of process 500 may be omitted.

[0095] Process 500 can be started by receiving an input audio signal in 502. In some implementations, the input audio signal may be a single-channel audio signal.

[0096] In step 504, process 500 can separate the input audio signal into vocal and non-vocal components. In some implementations, the vocal and non-vocal components can be identified using one or more trained machine learning models. Examples of machine learning models that can be used to separate the input audio signal into vocal and non-vocal components include deep neural networks (DNNs), convolutional neural networks (CNNs), long-term short-term memory (LSTM) networks, convolutional recurrent neural networks (CRNNs), gated recurrent units (GRUs), and convolutional gated recurrent units (CGRUs).

[0097] In some examples, process 500 can classify vocal components as either 1) speech or 2) non-speech. In some implementations, the classification of vocal components may be associated with a confidence level. Examples of classifiers that can be used to classify vocal components include k-nearest neighbors, case-based inference, decision trees, NaiveBayes, and / or various types of neural networks (e.g., convolutional neural networks (CNNs)).

[0098] In some examples, process 500 can classify non-vocal components as either 1) music or 2) non-music. In some implementations, the classification of non-vocal components may be associated with a confidence level. Examples of classifiers that can be used to classify non-vocal components include k-nearest neighbors, case-based inference, decision trees, NaiveBayes, and / or various types of neural networks (e.g., convolutional neural networks (CNNs)).

[0099] In 510, the process 500 can classify the input audio signal as one of the following: 1) speech, 2) music, 3) speech over music, or 4) other, by combining the classification of vocal components with the classification of non-vocal components. For example, in some embodiments, the classification of vocal components can be combined with the classification of non-vocal components using any suitable decision fusion algorithm that combines the classifications from two classifiers to generate an aggregated classification of the input audio signal. Examples of decision fusion algorithms that may be used to combine the classifications of two or more components include Bayesian, Dempster-Shafer, and fuzzy logic algorithms.

[0100] As another example, in some implementations, the classification of vocal components can be combined with the classification of non-vocal components based on the confidence levels of each of the vocal and non-vocal component classifications. As a more specific example, in some implementations, the classification of vocal and non-vocal components can be combined such that components associated with higher confidence levels are weighted more heavily in the combination.

[0101] In some implementations, the amount of reverberation present in an input audio signal can be determined. In some implementations, the amount of reverberation can be calculated using DRR. For example, in some implementations, the amount of reverberation can be inversely correlated with DRR, such that a decrease in DRR value increases the amount of reverberation, and an increase in DRR value decreases the amount of reverberation. In some implementations, the amount of reverberation can be calculated using the time period required for the sound pressure level to decrease by a certain amount (e.g., 60 dB). For example, the amount of reverberation can be calculated using RT60, which indicates the time it takes for the sound pressure level to decrease by 60 dB. In some implementations, the DRR or RT60 related to the input audio signal can be estimated using various algorithms or techniques, which may be signal processing-based and / or machine learning model-based.

[0102] In some implementations, the reverberation of an input audio signal can be calculated by estimating the diffusion of the input audio signal. Figure 6 shows an example of process 600 for estimating the diffusion of an input audio signal according to some implementations. Note that the blocks of process 600 may be executed in various orders not shown in Figure 6, and / or in some implementations, two or more blocks of process 600 may be executed substantially in parallel. Note that, as an addition or alternative, in some implementations, one or more blocks of process 600 may be omitted.

[0103] It should be noted that in some implementations, the amount of reverberation may be determined based on a combination of multiple metrics. These metrics may include, for example, DRR, RT60, and diffusion estimation. In some implementations, multiple metrics can be combined using various techniques such as weighted averaging. In some implementations, one or more metrics can be scaled or normalized.

[0104] Process 600 can be started by receiving the input audio signal at 602.

[0105] In step 604, process 600 can calculate the two-dimensional acoustic modulation frequency spectrum of the input audio signal. The two-dimensional acoustic modulation frequency spectrum can show the energy present in the input audio signal as a function of acoustic frequency and modulation frequency.

[0106] In 606, process 600 can determine the degree of spread of the input audio signal based on the energy in the high-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum (e.g., modulation frequencies above 6 Hz, modulation frequencies above 10 Hz, etc.). For example, in some implementations, process 600 can calculate the ratio of the energy in the high-modulation-frequency portion to the energy across all modulation frequencies. As another example, in some implementations, process 600 can calculate the ratio of the energy in the high-modulation-frequency portion to the energy in the low-modulation-frequency portion (e.g., modulation frequencies below 10 Hz, below 20 Hz, etc.).

[0107] Figures 7A, 7B, 7C, and 7D show examples of two-dimensional acoustic modulation frequency spectra for various types of input speech signals. As illustrated, each two-dimensional acoustic modulation frequency represents the energy present in the input signal as a function of acoustic frequency (shown on the y-axis of each spectrum shown in Figures 7A, 7B, 7C, and 7D) and modulation frequency (shown on the x-axis of each spectrum shown in Figures 7A, 7B, 7C, and 7D).

[0108] As shown in Figure 7A, clean speech with little or no reverberation may have a two-dimensional acoustic modulation frequency spectrum in which most of the energy is concentrated at relatively low modulation frequencies (e.g., below 5 Hz, below 10 Hz, etc.).

[0109] As shown in Figure 7B, an input signal containing both clean speech and early and late reverberation reflections may have a two-dimensional acoustic modulation frequency spectrum in which the energy is dispersed across all modulation frequencies.

[0110] As shown in Figure 7C, an input signal containing both clean speech and early reverberation reflections may have a two-dimensional acoustic modulation frequency spectrum in which the energy is generally concentrated at relatively low modulation frequencies (e.g., below 5 Hz, below 10 Hz). In other words, the two-dimensional acoustic modulation frequency of an input signal containing clean speech and early reverberation reflections (but not late reverberation reflections) may be substantially similar to the two-dimensional acoustic modulation frequency spectrum of clean speech alone.

[0111] As shown in Figure 7D, an input signal that includes late reverberation reflections and does not include clean speech or early reverberation reflections may have a two-dimensional acoustic modulation frequency spectrum in which the energy is dispersed across all modulation frequencies.

[0112] Therefore, as shown in Figures 7A, 7B, 7C, and 7D, diffusion estimation may be calculated based on the ratio of energy at relatively high modulation frequencies to the total energy, or based on the relative ratio of energy at relatively high modulation frequencies to energy at relatively low modulation frequencies.

[0113] Figure 8 is a block diagram showing examples of components of a device that can implement various aspects of the present disclosure. As with other figures provided in this specification, the number and types of elements shown in Figure 8 are merely examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 800 may be configured to perform at least some of the methods disclosed in this specification. In some embodiments, device 800 may be, or include, a television, one or more components of an audio system, a mobile device (such as a cell phone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0114] In some alternative implementations, device 800 may be a server or include one. In some such implementations, device 800 may be an encoder or include one. Thus, in some examples, device 800 may be a device configured for use in an audio environment, such as a home audio environment, and in other examples, device 800 may be a device configured for use in a “cloud,” such as a server.

[0115] In this example, device 800 includes an interface system 805 and a control system 810. In some implementations, the interface system 805 may be configured to communicate with one or more other devices in an audio environment. In some examples, the audio environment may be a home audio environment. In other examples, the audio environment may be a different type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, or a park environment. In some implementations, the interface system 805 may be configured to exchange control information and related data with audio devices in the audio environment. In some examples, the control information and related data may be related to one or more software applications running on device 800.

[0116] In some implementations, the interface system 805 may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, audio signals. In some cases, the audio data may include spatial data such as channel data and / or spatial metadata. In some examples, the content stream may include video data and corresponding audio data for the video data.

[0117] The interface system 805 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more USB (universal serial bus) interfaces). According to some implementations, the interface system 805 may include one or more wireless interfaces. The interface system 805 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 805 may include one or more interfaces between the control system 810 and a memory system, for example, an arbitrary memory system 815 shown in Figure 8. However, in some cases, the control system 810 may include a memory system. In some implementations, the interface system 805 may be configured to receive input from one or more microphones in the environment.

[0118] The control system 810 may include, for example, a general-purpose single or multiple chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic elements, discrete gates or transistor logic, and / or discrete hardware components.

[0119] In some implementations, the functionality of the control system 810 may reside in more than one device. For example, part of the control system 810 may reside in a device within one of the environments described herein, while another part of the control system 810 may reside in a device outside the environment, such as a server or a mobile device (e.g., a smartphone or tablet computer). In other examples, part of the control system 810 may reside in a device within one environment, while another part of the control system 810 may reside in one or more other devices within the environment. For example, the functionality of the control system may be distributed among multiple smart audio devices within the environment, or shared between an orchestration device (such as what is called a smart home hub) and one or more other devices within the environment. In other examples, part of the control system 810 may reside in a device implementing cloud-based services, such as a server, while another part of the control system 810 may reside in another device implementing cloud-based services, such as another server or memory device. The interface system 805 may also reside in multiple devices in some examples.

[0120] In some implementations, the control system 810 may be configured to perform at least partially the methods disclosed in this specification. According to some examples, the control system 810 may be configured to perform a method of reverberation removal based on media type classification.

[0121] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as RAM (random access memory) devices, ROM (read-only memory) devices, etc., as described herein. One or more non-temporary media may reside, for example, in an arbitrary memory system 815 and / or control system 810 shown in Figure 8. Accordingly, various novel aspects of the subject matter described herein may be implemented on one or more non-temporary media storing software. The software may include, for example, instructions for classifying the media type of audio content, determining the degree of reverberation, determining whether reverberation removal is performed, and controlling at least one device to perform reverberation removal on an audio signal. The software may be executable by one or more components of a control system, for example, control system 810 in Figure 8.

[0122] In some examples, device 800 may include an optional microphone system 820, as shown in Figure 8. The optional microphone system 820 may include one or more microphones. In some implementations, one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc. In some examples, device 800 does not have to include a microphone system 820. However, in some such implementations, device 800 may nevertheless be configured to receive microphone data from one or more microphones in the audio environment via the interface system 810. In some such implementations, a cloud-based implementation of device 800 may be configured to receive microphone data, or noise metrics at least partially corresponding to microphone data, from one or more microphones in the audio environment via the interface system 810.

[0123] In some implementations, device 800 may include an optional microphone system 825 as shown in Figure 8. The optional speaker system 825 may include one or more speakers, which may here be referred to as “speakers” or more generally “audio playback transducers.” In some examples (e.g., cloud-based implementations), device 800 does not need to include the speaker system 825. In some implementations, device 800 may include headphones. The headphones can be connected to or coupled to device 800 via a headphone jack or wireless connection (e.g., BLUETOOTH®).

[0124] In some examples, device 800 may include an optional sensor system 830, as shown in Figure 8. The optional sensor system 830 may include one or more touch sensors, gesture sensors, motion detectors, etc. According to some implementations, the optional sensor system 830 may include one or more cameras. In some implementations, the cameras may be standalone cameras. In some examples, one or more cameras of the optional sensor system 830 may be located in an audio device, which may be a single-purpose audio device or a virtual assistant. In some examples, one or more cameras of the optional sensor system 830 may be located in a television, mobile phone, or smart speaker. In some examples, device 800 does not have to include a sensor system 830. However, in some such implementations, device 800 may nevertheless be configured to receive sensor data from one or more sensors in the audio environment via an interface system 810.

[0125] In some examples, the device 800 may include an optional sensor system 835, as shown in Figure 8. The optional display system 835 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, the optional display system 835 may include one or more organic light-emitting diode (OLED) displays. In some cases, the optional display system 835 may include one or more displays of a television. In other examples, the optional display system 835 may include a laptop display, a mobile device display, or another type of display. In some examples where the device 800 includes the display system 835, the sensor system 830 may include a touch sensor system and / or a gesture sensor system adjacent to one or more displays of the display system 835. According to some such implementations, the control system 810 may be configured to control the display system 835 to present one or more graphical user interfaces (GUIs).

[0126] In some such examples, device 800 may be or include a smart audio device. In some such implementations, device 800 may be or include a startup word detector. In some examples, device 800 may be or include a virtual assistant.

[0127] Some aspects of this disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media (e.g., disks) for storing code to implement one or more examples or steps of the disclosed methods. For example, some of the disclosed systems are or may include programmable general-purpose processors, digital signal processors, or microprocessors programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments or steps of the disclosed methods. Such a general-purpose processor is or may include a computer system that includes input devices, memory, and processing subsystems programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0128] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and other configured) to perform necessary processing on an audio signal, including the execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor (e.g., a personal computer (PC), other computer system, or microprocessor (which may include input devices and memory)) programmed in software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the present invention may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to execute one or more examples of the disclosed methods, and the system may also include other elements (e.g., one or more speakers and / or one or more microphones). A general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled with input devices (e.g., a mouse and / or keyboard), memory, and a display device.

[0129] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code for performing one or more examples of the disclosed method or its steps (e.g., a Coder executable file to perform the execution).

[0130] While specific embodiments of the disclosed invention and applications of the disclosed invention are described in this specification, as will be apparent to those skilled in the art, many modifications to the embodiments and applications described in this specification are possible without departing from the scope of the disclosed invention as described and claimed in this specification. It should be understood that while specific forms of the disclosed invention are shown and described, the disclosed invention is not limited to the specific embodiments or methods described and shown.

[0131] Various aspects of the present invention may be apparent from the enumerated example embodiments (EEE) listed below.

[0132] (EEE1) A method for suppressing reverberation, The steps include receiving the input audio signal, The steps include classifying the media type of the input audio signal as one of the groups including at least (1) speech, (2) music, or (3) speech over music, A step of determining whether to perform reverberation removal on the input audio signal, at least based on the determination that the media type of the input audio signal is classified as speech, In response to a decision to perform reverberation removal on the input audio signal, the steps include: performing reverberation removal on the input audio signal to generate an output audio signal; A method that includes this.

[0133] (EEE2) The method according to EEE1, further comprising the step of determining the degree of reverberation in the input audio signal, wherein the decision of whether or not to perform reverberation removal on the input audio signal is made based on the degree of reverberation.

[0134] (EEE3) The method described in EEE2, wherein the degree of reverberation is based on reverberation time (RT60), direct-to-reverberant ratio (DRR), diffusion estimation, or any combination thereof.

[0135] (EEE4) The method according to EEE3, wherein the step of determining the degree of reverberation includes the step of calculating the two-dimensional acoustic modulation frequency spectrum of the input audio signal, and the degree of reverberation is based on the amount of energy in the high-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum.

[0136] (EEE5) The method according to EEE4, wherein the step of determining the degree of reverberation includes calculating at least one of the following: 1) the ratio of the energy of the high-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy of the two-dimensional acoustic modulation frequency spectrum over all modulation frequencies of the two-dimensional acoustic modulation frequency spectrum, or 2) the ratio of the energy of the high-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy of the low-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum.

[0137] (EEE6) The method according to EEE4 or 5, wherein the step of determining whether to perform reverberation removal on the input audio signal is based on the determination that the degree of reverberation exceeds a threshold.

[0138] (EEE7) The method according to any one of EEE1 to 6, wherein the step of classifying the media type of the input audio signal includes the step of separating the input audio signal into two or more spatial components.

[0139] (EEE8) The method according to EEE7, wherein the two or more spatial components include a center channel and a side channel.

[0140] (EEE9) The method according to EEE8, further comprising the steps of calculating the power of a side channel and classifying the side channel in response to the determination that the power of the side channel exceeds a threshold.

[0141] (EEE10) The method according to EEE7, wherein the two or more spatial components include a diffusive component and a direct component.

[0142] (EEE11) The step of classifying the media type of the input audio signal is: The step includes classifying each of the two or more spatial components as one of (1) speech, (2) music, or (3) speech over music, The method according to any one of EEE7 to 10, wherein the media type of the input audio signal is classified by combining the classifications of each of the two or more spatial components.

[0143] (EEE12) The method according to any one of EEE7 to 11, wherein the input audio signal is separated into two or more spatial components in response to the determination that the input audio signal includes stereo audio.

[0144] (EEE13) The method according to any one of EEE1 to 6, wherein the step of classifying the media type of the input audio signal includes the step of separating the input audio signal into vocal components and non-vocal components.

[0145] (EEE14) The method according to EEE13, wherein, in response to determining that the input audio signal includes a single audio channel, the input audio signal is separated into the vocal component and the non-vocal component.

[0146] (EEE15) The step of classifying the media type of the input audio signal is: The steps include classifying the aforementioned vocal component as either (1) speech or (2) non-speech, The steps include classifying the aforementioned non-vocal component as either (1) music or (2) non-music, Includes, The media type of the input audio signal is classified by combining the classification of the vocal component and the classification of the non-vocal component, according to the method of EEE13 or 14.

[0147] (EEE16) The method according to any one of EEE1 to 15, wherein the step of determining whether to perform reverberation removal on the input audio signal is based on the classification of a second input audio signal preceding the input audio signal.

[0148] (EEE17) Step of receiving a third input audio signal, The steps include deciding not to perform reverberation removal on the third input audio signal, In response to deciding not to perform reverberation removal on the third input audio signal, the steps include: prohibiting the execution of the reverberation removal algorithm on the third input audio signal; One of the methods EEE1-16, which further includes the above.

[0149] (EEE18) The method according to EEE17, wherein the step of determining whether reverberation de-emission is performed on the third input audio signal is at least in part based on the classification of the media type of the third input audio signal.

[0150] (EEE19) The method according to EEE18, wherein the classification of the media type of the third input speech signal is one of 1) music or 2) speech over music.

[0151] (EEE20) The method according to any one of EEE17-19, wherein the step of determining that reverberation removal is not performed on the third input audio signal is at least in part based on the determination that the degree of reverberation in the third input audio signal is below a threshold.

[0152] (EEE21) Equipment configured to carry out the method described in any one of the EEE1 to EEE20 items.

[0153] (EEE22) A system configured to implement the method described in any one of the items EEE1 to EEE20.

[0154] (EEE23) One or more non-temporary media storing software, wherein the software includes instructions for controlling one or more devices to perform the method described in any one of EEE1 to 20.

[0155] (EEE24) A method for classifying an input audio signal as one of at least two media types, The steps include receiving the input audio signal, The steps include separating the input audio signal into two or more spatial components, A step of classifying each of the two or more spatial components as at least one of two media types, wherein the media type of the input audio signal is classified by combining the classifications of each of the two or more spatial components, A method that includes this.

[0156] (EEE25) The two or more spatial components include a center channel and a side channel, and the method is The steps include calculating the power of the side channel, A step of classifying the side channel in response to determining that the power of the side channel exceeds a threshold, The method described in EEE24, further including the above.

[0157] (EEE26) The method according to EEE24, wherein the two or more spatial components include a diffusing component and a direct component.

[0158] (EEE27) The method according to any one of EEE24 to 26, wherein the input audio signal is separated into two or more spatial components in response to the determination that the input audio signal includes stereo audio.

[0159] (EEE28) The method according to any one of EEE24 to 26, wherein the step of classifying the media type of the input audio signal includes the step of separating the input audio signal into vocal components and non-vocal components.

[0160] (EEE29) The method according to EEE28, wherein, in response to determining that the input audio signal includes a single audio channel, the input audio signal is separated into the vocal component and the non-vocal component.

[0161] (EEE30) The step of classifying the media type of the input audio signal is: The steps include classifying the aforementioned vocal component as either (1) speech or (2) non-speech, The steps include classifying the aforementioned non-vocal component as either (1) music or (2) non-music, Includes, The media type of the input audio signal is classified by combining the classification of the vocal component and the classification of the non-vocal component, according to the method of EEE28 or 29.

[0162] A system configured to implement the method described in any one of the items of EEE24-30.

[0163] A non-temporary medium storing software, wherein the software includes instructions for controlling one or more devices to perform the method described in any one of the EEE1 to 30.

Claims

1. A method for suppressing reverberation, The steps include receiving the input audio signal, The steps include classifying the media type of the input audio signal as one of the groups including at least (1) speech, (2) music, or (3) speech over music, A step of determining the degree of reverberation of the input audio signal, wherein the step of determining the degree of reverberation includes calculating the two-dimensional acoustic modulation frequency spectrum of the input audio signal, and the degree of reverberation is based on the amount of energy in the high-modulation frequency portion of the two-dimensional acoustic modulation frequency spectrum, A step of determining whether to perform reverberation removal on the input audio signal, based at least on the determination that the media type of the input audio signal is classified as speech, and the degree of reverberation, In response to a decision to perform reverberation removal on the input audio signal, the steps include generating an output audio signal by performing reverberation removal on the input audio signal, A method that includes this.

2. The method according to claim 1, wherein the degree of reverberation is based on reverberation time (RT60), direct-to-reverberation ratio (DRR), estimation of diffusion, or any combination thereof.

3. The method according to claim 2, wherein the step of determining the degree of reverberation includes calculating at least one of (1) the ratio of the energy of the high-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy of the two-dimensional acoustic modulation frequency spectrum over all modulation frequencies of the two-dimensional acoustic modulation frequency spectrum, or (2) the ratio of the energy of the high-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum to the energy of the low-modulation-frequency portion of the two-dimensional acoustic modulation frequency spectrum.

4. The method according to claim 2 or 3, wherein the step of determining whether to perform reverberation removal on the input audio signal is based on the determination that the degree of reverberation exceeds a threshold.

5. The step of classifying the media type of the input audio signal includes the step of separating the input audio signal into two or more spatial components. The method according to any one of claims 1 to 4, wherein, in response to the determination that the input audio signal includes stereo audio, the input audio signal is separated into two or more spatial components.

6. The two or more spatial components include a center channel and a side channel, and the method is The steps include calculating the power of the side channel, A step of classifying the side channel in response to determining that the power of the side channel exceeds a threshold, The method according to claim 5, further comprising:

7. The method according to claim 5, wherein the two or more spatial components include a diffusion component and a direct component.

8. The step of classifying the media type of the input audio signal is: The step includes classifying each of the two or more spatial components as one of (1) speech, (2) music, or (3) speech over music, The method according to any one of claims 5 to 7, wherein the media type of the input audio signal is classified by combining the classifications of each of the two or more spatial components.

9. The step of classifying the media type of the input audio signal includes the step of separating the input audio signal into vocal components and non-vocal components. The method according to any one of claims 1 to 4, wherein, in response to determining that the input audio signal includes a single audio channel, the input audio signal is separated into the vocal component and the non-vocal component.

10. The step of classifying the media type of the input audio signal is: The steps include classifying the aforementioned vocal component as either (1) speech or (2) non-speech, The steps include classifying the aforementioned non-vocal component as either (1) music or (2) non-music, Includes, The method according to claim 9, wherein the media type of the input audio signal is classified by combining the classification of the vocal component and the classification of the non-vocal component.

11. The method according to any one of claims 1 to 10, wherein the step of determining whether to perform reverberation removal on the input audio signal is based on the classification of a second input audio signal preceding the input audio signal.

12. The steps include receiving a third input audio signal, The step of deciding not to perform reverberation removal on the third input audio signal, In response to deciding not to perform reverberation removal on the third input audio signal, the steps include: prohibiting the execution of the reverberation removal algorithm on the third input audio signal; It further includes, The method according to any one of claims 1 to 11, wherein the step of deciding not to perform reverberation removal on the third input audio signal is at least in part based on (a) classifying the media type of the third input audio signal, or (b) determining that the degree of reverberation in the third input audio signal is below a threshold, and the classification of the media type of the third input audio signal is one of (1) music, or (2) speech over music.

13. A device configured to carry out the method described in any one of claims 1 to 12.

14. One or more non-temporary media storing software, wherein the software includes instructions for controlling one or more devices to perform the method according to any one of claims 1 to 12.