Content enhancement based on extracted, inferred, and / or supplemented audio characteristics
The method and system optimize audio content using machine-learning classifiers and dynamic range control to adapt audio for diverse platforms, addressing inefficiencies in content creation and enhancing user experience.
Patent Information
- Application Number
- PCT/US2025/029627
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-05-15
- Publication Date
- 2025-11-20
AI Technical Summary
Content creators face challenges in optimizing audio content for different platforms, requiring complex workflows and multiple versions due to varying audience preferences and distribution channels, which can be inefficient and bandwidth-intensive.
A method and system for enhancing audio content using machine-learning-based classifiers for type identification, denoising, and dynamic range control, with low computational complexity, enabling content optimization for various platforms and devices.
Enables efficient creation of a single master copy that adapts to multiple platforms, improving audio quality and personalization, reducing the need for multiple versions, and enhancing consumer experience.
Smart Images

Figure US2025029627_20112025_PF_FP_ABST
Abstract
Description
CONTENT ENHANCEMENT BASED ON EXTRACTED, INFERRED, AND / OR SUPPLEMENTED AUDIO CHARACTERISTICS 1. Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from PCT Application No. PCT / CN2024 / 094050 filed on 17 May 2024, U.S. Provisional Application No. 63 / 658,827, filed on 11 June 2024, and European Application No. 24183184.1 filed on 19 June 2024, each of which is incorporated by reference herein in its entirety. 2. Field of the Disclosure
[0002] Various example embodiments relate generally, but not exclusively, to content enhancement using audio recognition and / or sound classification based on a plurality of extracted, inferred, and / or supplemented audio characteristics. 3. Background
[0003] For content creators in the audio space, it is often important to edit their content for different platforms. For example, for podcasting, narrating, or producing music, a content creator typically needs to optimize audio quality, format, length, and key features for the target audience, audio-rendering environment, and distribution channel. Accordingly, practices and tools for editing audio content for different platforms are being actively developed by manufactures of audio and imaging products for the cinema, television, broadcast, and entertainment industries. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS
[0004] Disclosed herein are various examples, embodiments, features, and aspects of a method and system for enhancing audio content. In some examples, the method is based on time-frequency distribution, narrative importance, and context information in capture. Operations of the method may include percussive / harmonic component separation, music detection, beat extraction, independent dynamic range control (DRC), and environment type analysis in capture. In at least some examples, the method can beneficially be implemented with relatively low computational complexity.
[0005] In one example, a method of enhancing audio content includes: analyzing the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; applying respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assigning respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and performing content enhancement on the audio content based on the type identifier and the respective levels of importance.
[0006] In another example, an audio system for enhancing audio content includes: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: analyze the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; apply respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assign respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and perform content enhancement on the audio content based on the type identifier and the respective levels of importance.
[0007] According to yet another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method of enhancing audio content. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:
[0009] FIG. 1 is a block diagram illustrating an audio system in which various embodiments can be practiced.
[0010] FIG. 2 is a block diagram illustrating a content-enhancement workflow implemented in a component of the audio system of FIG. 1 according to some examples.
[0011] FIG. 3 is a flowchart illustrating a method of training an audio and visual classifier used in the content-enhancement workflow of FIG. 2 according to some examples.
[0012] FIG. 4 is a flowchart illustrating a beat extraction and rhythm detection method used in the content-enhancement workflow of FIG. 2 according to some examples.
[0013] FIG. 5 graphically illustrates frequency filters used in the beat extraction and rhythm detection method of FIG. 4 according to some examples.
[0014] FIGS. 6-13 graphically illustrate various operations performed in the beat extraction and rhythm detection method of FIG. 4 according to some examples.
[0015] FIGS. 14A-14F graphically illustrate the loudness statistics corresponding to different sources of audio signals according to some examples.
[0016] FIG. 15 is a block diagram illustrating a dynamic range control (DRC) method used in the content-enhancement workflow of FIG. 2 according to some examples.
[0017] FIGS. 16A-16H graphically illustrate the root mean square (RMS) and peak statistics of speech and music content according to some examples.
[0018] FIGS. 17A-17C graphically illustrate gain calculations performed in the DRC method of FIG. 15 according to one example.
[0019] FIGS. 18A-18C graphically illustrate gain calculations performed in the DRC method of FIG. 15 according to another example.
[0020] FIGS. 19A-19C graphically illustrate gain calculations performed in the DRC method of FIG. 15 according to yet another example.
[0021] FIGS. 20A-20F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a first embodiment of the DRC method of FIG. 15 according to some examples.
[0022] FIGS. 21A-21F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a second embodiment of the DRC method of FIG. 15 according to some examples.
[0023] FIGS. 22A-22F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a third embodiment of the DRC method of FIG. 15 according to some examples.
[0024] FIG. 23 is a flowchart illustrating a content-enhancement method implemented using the audio system of FIG. 1 according to some examples.
[0025] FIG. 24 is a block diagram of an example computing device one or more instances of which can be used in the audio system of FIG. 1 according to various examples. DETAILED DESCRIPTION
[0026] Next Generation Audio (NGA) offers significant benefits to audiences, content creators, and broadcasters, such as immersive audio, personalization for listening preferences, and long-term efficiencies in the production and distribution. Although immersive audio is already highly valued, many in the industry believe that personalization has the potential to be even more widely appreciated. For example, the ability to adjust dialogue levels relative to other sounds will address one of the most frequent complaints received by broadcasters. NGA also offers a potential for realizing long-term efficiencies in the production and distribution, especially when content revisions are needed.
[0027] In some examples, NGA tools benefit content creators and broadcasters by enabling one master copy or stream to provide optimum experience for a range of platforms and user devices. This approach is not just bandwidth efficient but can also create process efficiencies for broadcasters and producers by simplifying complex workflows. For example, multiple versions can be created from the master copy in an expeditious and straightforward manner with content enhancement tools, thereby potentially removing the need for creating multiple master versions targeted for differentplatforms. These capabilities enable content creators to tell their stories better to the benefit of the consumer. Consumers also gain the ability to personalize the audio for better experience and improved content accessibility.
[0028] FIG. 1 is a block diagram illustrating an audio system 100 in which various embodiments can be practiced. The audio system 100 includes an audio encoder 120 and an audio decoder 140 connected via a communication channel 130. The encoder 120 receives input signals 112 and processes the received input signals to generate an encoded bitstream 132. The decoder 140 receives the encoded bitstream 132 via the communication channel 130 and decodes the received bitstream to generate output audio signals 142. The output audio signals 142 are applied to an audio rendering component 150 that operates to render and playback the audio content represented by the audio signals 142.
[0029] The input signals 112 received by the encoder 120 are generated with a content enhancer 110 in response to input audio signals 102. The content enhancer 110 applies processing to the input audio signals 102 based on configuration parameters 104 to generate the corresponding enhanced signals 112, which are then directed to the encoder 120. In some examples, the processing applied by the content enhancer 110 to the input audio signals 102 is based on time-frequency distribution, narrative importance, and context information in capture. Various examples of the processing implemented in the content enhancer 110 are described in more detail below in reference to FIGS. 2-23.
[0030] In the example shown, the content enhancer 110 is located at the transmitter side (e.g., at the server). However, embodiments are not so limited. In some examples, the content enhancer 110 can be placed at the receiver side (e.g., at the player). In the latter examples, the content enhancer 110 may be inserted between the decoder 130 and the audio rendering component 150.
[0031] In some examples, the audio rendering component 150 may include any professional or consumer-grade audio system, such as a home theater (e.g., including an A / V receiver, a soundbar, a Blu-ray player, etc.), one or more E-media devices (e.g., a computer, a tablet, a mobile phone equipped with headphones and / or speakers, etc.), a TV set, and a sound reproduction system. In some examples, the audio rendering component 150 provides an audio environment for playback of audio or audio / visual content using a plurality of speakers and suitable playback devices. In some examples, the audio rendering component 150 may represent any environment in which a listener isexperiencing playback of the audio content, such as a cinema, a concert hall, an outdoor theater, a home theater or room, a listening booth, a car, a game console, a headset device, a public address (PA) system, or other audio playback environment.
[0032] FIG. 2 is a block diagram illustrating a content-enhancement workflow 200 implemented in the content enhancer 110 according to some examples. The workflow 200 includes, but is not limited to, percussive / harmonic component separation, music detection, beat extraction, independent dynamic range control (DRC), and environment-type analysis in capture. In the example shown, the workflow 200 is implemented using a plurality of blocks, which are labeled using the reference numerals 210-260. The functionalities of the blocks 210-260 of the workflow 200 are briefly described in reference to FIG. 2 and are further detailed and illustrated in reference to FIGS. 3-22.
[0033] The content analysis block 210 of the workflow 200 is configured to understand the types and / or objects of the audio content received via the input audio signals 102 and to identify content types using audio information and / or corresponding visual information (when available). In some examples, the content types can be selected from the group of content types associated with the following short content descriptors: person, animal, urban, vehicle(s), nature landscape, sport(s), music, cooking, and food sharing. In some cases, the content analysis block 210 is configured to use a contrastive loss to reduce the occurrence of large prediction errors. An illustrative nonlimiting example of such large prediction error is a case of identifying the sport type content as a cooking type content. For example, the content analysis block 210 may employ a music detection module to determine if the dominant signal is music (or not). The content analysis block 210 is preferably capable of operating with low latency. In some examples, the content analysis block 210 may employ a selected conventional segment-feature-based method (such as AdaBoost or a convolutional neural network, CNN) and be configured to use relatively few look-ahead frames, with the extent of this limitation depending on the specific use case. For a frame-feature-based method, such as a deep neural network (DNN) or a recurrent neural network (RNN), a smoothing mechanism can be applied to reduce or avoid undesired fluctuations. Other suitable content analysis algorithms can also be used in the content analysis block 210, depending on the specific use case.
[0034] The denoising blocks 2201and 2202of the workflow 200 are configured to use different respective denoising methods for a music signal and a non-music signal. For example, for the music signal, the music denoising block 2201 is configured to remove stationary noise. For the non-musicsignal, the non-music denoising block 2202is configured to treat all signal components as noise, with the exception of speech.
[0035] The components separation blocks 2301 and 2302 of the workflow 200 are configured to estimate percussive, harmonic, and noise-floor time-frequency (T-F) components present in their respective input signals received from the denoising blocks 2201and 2202, respectively. Herein, harmonic components are represented by T-F bins that are relatively smooth in the time dimension. Percussive components are represented by T-F bins that are relatively smooth in the frequency dimension. These smoothness characteristics are referred to as “anisotropic smoothness” of the signal. The remaining component of the signal after the percussive and harmonic components are separated out represents the noise floor.
[0036] The blocks 2401and 2402of the workflow 200 are configured to perform beat extraction / rhythm detection and narrative importance ranking, respectively. The results of beat extraction and rhythm detection generated with the block 2401 are used to enhance the music content. The audio components determined with the block 2402 provide a basis for narrative- importance-based remixing for each pre-defined content type in the non-music content.
[0037] The dynamic range control (DRC) blocks 2501 and 2502 of the workflow 200 are configured to use different respective DRC signal transfer curves for different content types (e.g., speech, music, noise) and / or different signal components (e.g., percussive components, harmonic components, etc.).
[0038] The content enhancement blocks 2601 and 2602 of the workflow 200 are configured to perform content enhancement and optimization based on the previously obtained information, such as the audio and visual objects or types determined in the content analysis block 210. In some examples, the content enhancement blocks 2601 and 2602 use a different respective profile for each content type to improve the audio and video quality accordingly. Content Analysis
[0039] In some examples, the content analysis block 210 includes a neural-network-based content classifier. A challenging aspect of training the content classifier is to obtain adequate and sufficient training data. In some examples, AudioSet can be used as the audio training material for training the content classifier used in the content analysis block 210. AudioSet is publicly availablefrom the Sound Understanding group in the Machine Perception Research organization at Google and includes an expansive ontology of 632 audio event classes and a collection of 2,084,320 human- labeled 10-second sound clips drawn from YouTube videos. The ontology is specified as a hierarchical graph of event categories, covering a wide range of human and animal sounds, musical instruments, and genres, and common everyday environmental sounds.
[0040] In additional examples, YouTube-8M can be used as the visual training material for training the content classifier used in the content analysis block 210. YouTube-8M is also publicly available from the Machine Perception Research organization at Google and includes approximately 8 million videos (500,000 hours of video), annotated with a vocabulary of 4,800 visual entities. The video labels are obtained with the YouTube video-annotation system, which labels videos with their main topics. While these labels are machine-generated, they are derived from a variety of human- based signals including metadata and query click signals. The video labels (Knowledge Graph entities) are filtered using both automated and manual curation strategies, including asking human raters if the labels are visually recognizable. Each video is decoded at one-frame-per-second, and a deep CNN pre-trained on ImageNet is used to extract the hidden representation prior to the classification layer. The frame features are compressed, and both the features and video-level labels are available for download.
[0041] FIG. 3 is a flowchart illustrating a method 300 of training an audio and visual classifier used in the content analysis block 210 according to some examples. The method 300 is directed at filtering the AudioSet and YouTube-8M datasets and improving the original labeling therein for the content-enhancement purposes of the workflow 200. The method 300 also includes finally training the classifier(s) used in the content analysis block 210.
[0042] The method 300 includes grouping related classes in AudioSet and YouTube-8M into fewer predefined content types (in a block 302). For the visual data, images are extracted from the video per every 30 frames. Example content types include person, animal, urban, vehicle(s), nature landscape, sport(s), music, cooking, food, etc.
[0043] The method 300 also includes performing initial training of the audio and visual classifiers (in a block 304). The training is performed using the grouped dataset obtained in the block 302. The accuracy of these two content classifiers may be initially relatively low because some of the training data obtained in the block 302 may be insufficiently accurately labeled. Theaccuracy typically gradually improves in subsequent instances of the block 304 as the method 300 proceeds through multiple loops 302-308.
[0044] The method 300 also includes filtering the training dataset with the preliminarily trained classifier based on a selected confidence threshold (in a block 306). Example operations of the block 306 include excluding wrongly predicted data from the training dataset based on the confidence threshold. The confidence threshold is a hyperparameter of the method that can be set iteratively or empirically, e.g., based on prior experience.
[0045] The method 300 also includes determining whether a sufficient volume of data is correctly predicted with the preliminarily trained classifier (in a decision block 308). Operations of the block 308 include comparing the volume of correctly predicted data with a selected threshold value. When the volume is below the threshold value (“No” at the decision block 308), the processing of the method 300 is looped back to the block 302. When the volume is at or above the threshold value (“Yes” at the decision block 308), the processing of the method 300 is directed to a block 310.
[0046] Operations of the block 310 include finally training the audio and visual classifiers using the latest filtered training dataset obtained in the last occurrence of the block 306. Upon completion of the final training, the method 300 is terminated.
[0047] To avoid relatively large prediction errors in the classifier, a contrastive loss is used in the training operations of the method 300. The contrastive loss includes categorical cross-entropy and a margin weight. The weight is obtained by calculating the id distance between the true value y_true and the predicted value y_pred, e.g., as follows: ^^^^^^^^^^^^(y^^^^, y^^^^) = |^^^^^^(y^^^^) − ^^^^^^(y_pred)| !^ (1)(2)Based on the definition provided by Eqs. (1)-(2), one can control the performance of the classifier by sorting the class based on a preference. For example, one can accept that a cooking content is recognized as a “person” type because there is typically a person present who is making the food. However, it will constitute large error when a cooking content is recognized as a sport type. To address this type of issues, one can set the id distance between cooking and sport to be as large as possible. Based on the margin weight, the final contrastive loss, #!^$%&&, can be expressed as:#!^$%&& = categorical.^ / 00^1^^ / ^2324567,28579: ∗ [^^^^^^^^^^^^3y^^^^, y^^^^: + 1] (3)model choice, and AdaBoost, CNN, DNN, RNN, or various combinations thereof can be used. More specifically, a music detection block of the content analysis block 210 is configured to detect whether the dominant signal is music (or not) and is capable to make such detection with relatively low latency. For conventional segment-feature-based methods, such as AdaBoost or CNN, the segment preferably has as few look-ahead frames as possible, depending on the specific use case. For frame-feature- based methods, such as DNN or RNN, some smoothing mechanism can be applied to make the result relatively smooth and to avoid undesired fluctuations. Other suitable content analysis algorithms are segment-based or field-based, with the made selection depending on the specific use case as well. Denoising
[0049] The denoising blocks 2201and 2202of the workflow 200 are configured to use different respective denoising methods for a music signal and a non-music signal. For example, for a music signal, only stationary noise is removed. For a non-music signal, all components are regarded as noise, with the exception of speech. In some examples, the denoising block 2201is implemented using RNNv3 (e.g., as described in BCAP 1.6). The denoising block 2202 is implemented using LensNet. By using these two denoising model, we are able to get two types of noise. The first type is stationary noise, which is removed in most cases. The second type is the ambient noise, which can provide an immersive feeling. In some examples, one can choose different strategies with respect to the ambient sound based on the content analysis results of the content analysis block 210. Such strategies are described in more detail below in the subsection entitled “Content Enhancement and Optimization.” Percussive and Harmonic Components Separation
[0050] In some examples, a concept implemented in the blocks 2301and 2302of the workflow 200 is based on the observation that that the dimensions of smoothness in the spectrograms corresponding to typical percussive and harmonic musical instruments are different. More specifically, spectrograms of harmonic instruments, such as a guitar, are typically “smooth” in the time dimension, owing to their quasi-stationarity. In contrast, spectrograms of percussiveinstruments are typically “smooth” in the frequency dimension, owing to their impulse-like nature. Based on these concepts, in various embodiments, the harmonic components are defined as the T-F bins that are smooth in the time dimension, whereas the percussive components are defined as the T- F bins that are smooth in the frequency dimension. These smoothness characteristics represent “anisotropic smoothness” of the signal. The remaining component, after identification of the percussive and harmonic components, represents the noise floor. In some examples, implementation of the blocks 2301 and 2302 may benefit from certain features of the denoising method described in K. Miyamoto, M. Tatezono, J. Le Roux, H. Kameoka, N. Ono, and S.Sagayama, “Separation of harmonic and non-harmonic sounds based on 2D-filtering of the spectrogram,” Proc. Acoustic Soc. Jpn. Autumn Conf. (in Japanese), 2007, pp. 825–826, which is incorporated herein by reference in its entirety. In such examples, the implementation complexity for the blocks 2301and 2302is relatively low. Beat Extraction and Rhythm Detection
[0051] FIG. 4 is a flowchart illustrating a beat extraction and rhythm detection method 400 used in the content-enhancement workflow 200 according to some examples. The method 400 can be used to implement parts of the blocks 2301 and 2302 of the content-enhancement workflow 200.
[0052] The method 400 includes applying a T-F transform (in a block 402). In some examples, operations of the block 402 include applying a 1024-point fast Fourier transform (FFT) with 512 overlaps and applying frequency filtering using one or more filters. The operations of the block 402 also include outputting selected filtered signals produced in the filtering as extracted beat signals 404.
[0053] FIG. 5 graphically illustrates frequency filters used in the block 402 of the method 400 according to some examples. In the example shown, the frequency filters include: (i) a low- frequency filter characterized by a first frequency band 502; (ii) an intermediate-frequency filter characterized by a second frequency band 504; and (iii) a high-frequency filter characterized by a third frequency band 506. The filtered signals produced with the low-frequency filter (represented by the frequency band 502) are used as the extracted beat signals 404.
[0054] The method 400 also includes banding (in a block 406) the frequency bins produced in the block 402. In some examples, the bins are grouped into 18 octave bands. Four bins are used as a basis. There are two sub-bands in each octave scale.
[0055] The method 400 also includes beat detection in the frequency domain (in a block 408). In some examples, operations of the block 408 include calculating the first order delta energy between bands. Operations of the block 408 further include removing negative values.
[0056] The method 400 also includes smoothing operations (in a block 410). In some examples, the smoothing operations include median filtering with a kernel of seven.
[0057] The method 400 also includes beat detection in the time domain (in a block 412). In some examples, operations of the block 412 include calculating the first order delta energy between frames. Operations of the block 412 further include removing negative values.
[0058] The method 400 also includes normalization operations (in a block 414). In some examples, operations of the block 414 include performing general (global) normalization or performing local normalizations, e.g., in each of the frequency bands.
[0059] The method 400 also includes Gaussian filtering (in a block 416). In some examples, the parameters of the Gaussian filtering are selected as follows: (i) interval width = 100ms; (ii) amplitude p=1; and (iii) sigma=5.
[0060] The method 400 also includes low peak filtering (in a block 418). In some examples, the filtering threshold for the low peak filtering in the block 418 is selected to be 0.1. All peaks with amplitudes lower than this threshold are filtered out (removed).
[0061] The method 400 also includes beat period calculation (in a block 420). In some examples, operations of the block 420 include calculate the median beat period among a plurality of intervals.
[0062] The method 400 also includes beat modeling or approximation (in a block 422). In some examples, operations of the block 422 are implemented in accordance with the following expressions: tt = {0,1,2,…, ⌈ABC^ ∗ D^⌉} (4)HII ^= AA ∗ F^G IJKL∗MLNO ∗ cos (2 ∗ G^ ∗ ^^&R∗ STU^G) / 2 (5)where tdur = of the block 422, the
[0063] FIGS. 6-13 graphically illustrate various operations performed in the beat extraction and rhythm detection method 400 according to some examples. The time scale used in FIGS. 6, 7, 11, and 13 is given in terms of the sample number. The frequency scale in FIGS. 8-10 is given in terms of the bin number. FIGS. 6-13 are described in more detail below in continued reference to FIG. 4.
[0064] FIG. 6 is a graph illustrating an input audio signal 602 received at the block 402 of the beat extraction and rhythm detection method 400 according to one example. The input audio signal 602 contains sounds from several musical instruments. One of those instruments is a bass drum.
[0065] FIG. 7 graphically illustrates a result of applying the filtering of the block 402 in the method 400 to the waveform 602. The filtering is performed with the above-described low frequency filter characterized by the frequency response curve 502 (FIG. 5). A resulting filtered waveform 702 represents an example of the extracted signals of the block 404 in the method 400.
[0066] FIG. 8 graphically illustrates a result of beat detection in the frequency domain performed in the block 408 of the method 400. A resulting spectrum 802 corresponds to the filtered waveform 702 shown in FIG. 7. Note that the negative values have been removed.
[0067] FIG. 9 graphically illustrates a combined result of the processing performed in the blocks 410-414 of the method 400. A resulting spectrum 902 is normalized and contains spectral data corresponding to several time segments.
[0068] FIG. 10 graphically illustrates a combined result of the processing performed in the blocks 416 and 418 of the method 400. Note that a resulting spectrum 1002 has fewer peaks than the spectrum 902 due to the removal of low amplitude peaks in the block 418.
[0069] FIG. 11 graphically illustrates a waveform 1102 corresponding to one beat. The waveform 1102 is obtained by applying an inverse Fourier transform to the spectrum 1002.
[0070] FIGS. 12-13 graphically illustrate operations performed in the block 422 of the method 400. More specifically, FIG. 12 graphically shows an estimated beat position 1202 of the beat 1102in the waveform 702 as a function of time. FIG. 13 graphically shows a modeled waveform 1302 corresponding to the waveform 702. The waveform 1302 is obtained by combining multiple instances of the single-beat waveform 1102 using the estimated beat position 1202. Narrative Importance Ranking
[0071] Various embodiments are directed to content enhancement based on narrative importance. For each content type, the audio objects that may exist in that content are classified into the following categories: • Highest Importance (EI): the highest level, containing speech, bird chirping, etc. • High Importance (HI): important non-speech sounds, such as the animal / pets / babble sounds. • Medium Importance (MI): background sounds (which can be somewhat attenuated), such as wind noise, rain, thunder, water sound, etc. • Low Importance (LI): ambient sounds (which can be strongly attenuated to improve intelligibility). In this manner, we do not need to separate and process all audio objects. Excluding some of the audio objects may be beneficial due to lower power consumption and a complexity reduction for profile settings.
[0072] As an illustration, an example importance classification is provided below for the “Cooking / Food Sharing” type of content: EI: speech; HI: background music; MI: Percussive sounds, such as chewing and cutting; and LI: background noise, such as a kitchen ventilator noise. For other content types, we similarly group different objects and background sounds according to the EI, HI, MI, and LI levels. Low-Complexity Dynamic Range Control (DRC)
[0073] The signal loudness captured by different devices (e.g., mobile devices, earbuds, etc.) typically has different ranges. In addition, the loudness range may vary for different captured signals (e.g., speech, music, etc.). To make the signal level match human’s best listening experience, the signal is typically boosted or compressed to the most satisfying levels. The DRC isa principle that can be used to control the range of the signal. In practice, this type of control means that soft sounds are mapped to higher gains to boost, whereas loud sounds are mapped to lower gains to compress. Therefore, the range of covered loudness levels becomes smaller. Most of conventional DRC algorithms rely on a long history or look-ahead information. Additionally, the algorithm complexity is another important issue when the targeted deployment includes portable devices. Thus, developing a real-time DRC algorithm characterized by relatively low complexity is an important need.
[0074] As used herein, the term “real time” refers to a computer-based process that controls a corresponding environment by receiving data, processing the received data, and generating a response sufficiently quickly to affect the environment without significant delay. Real-time responses are often understood to be on the order of milliseconds, or sometimes microseconds. In the context of DRC, “real-time” updates mean that the device sufficiently accurately adjusts the gain(s) at any point in time.
[0075] FIGS. 14A-14F graphically illustrate the loudness statistics corresponding to different sources of audio signals according to some examples. The shown examples demonstrate representative distributions in the collection of one hundred speech-dominant vlogs and fifty music dominant vlogs recorded using an earbud. The acquired statistics indicate that the Loudness Units relative to Full Scale (LUFS) range of the recorded speech content is in the approximate dB range of [−60, −15] whereas the LUFS range of the recorded music content is in the approximate dB range of [−25, −5].
[0076] FIG. 15 is a block diagram illustrating a DRC method 1500 implemented in the DRC blocks 2501 and 2502 of the workflow 200 according to some examples. In the example shown, the DRC method 1500 includes: (i) a downmixing module 1510; (ii) an energy distribution module 1520; (iii) a frame loudness calculation module 1530; (iv) an output loudness calculation module 1540; (v) a gain calculation module 1550; (vi) a postprocessing module 1560; and (vii) a two- channel gain mapping module 1570. Each of these modules is described in more detail below.
[0077] The downmixing module 1510 receives a two-channel (e.g., stereo) input 1502 and operates to convert the received two-channel input into a mono output 1512. In one example, the conversion is performed in accordance with Eq. (6): mono = 0.5 L + 0.5 R (6)where L and R denote the two channels of the two-channel input 1502. While working reasonably well in most cases, the conversion performed in accordance with Eq. (6) may cause temporal inconsistency when the respective peaks of the L and R channels are relatively delayed by a relatively large delay time, which may occur for an audio object with a large direction change. Such situations are handled better when the conversion performed in accordance with Eq. (7): mono = 0.5 E(L) + 0.5 E(R) (7) where E denotes energy.
[0078] The energy distribution module 1520 is configured to calculate the energy difference (in dB) between the peak and mean energy, hereafter denoted as pmd. The energy distribution module 1520 is further configured to calculate the percentage of bins whose energy is higher than the alpha×(peak energy). The value of the parameter alpha (which is smaller than 1) can be selected based on the specific application. The energy distribution module 1520 is further configured to adjust the pmd value based on the calculated percentage. In some examples, the applied adjustment is as follows: if the percentage > 0.1: pmd = delta_E_dB × 0.9 elif the percentage > 0.5: pmd = delta_E_dB × 0.8 else: pmd = delta_E_dB × 0.6 In other examples, other adjustment schemes can also be used.
[0079] FIGS. 16A-16H graphically illustrate the RMS and peak statistics of speech and music content according to some examples. From the results illustrated in FIGS. 16A-16H, it can be concluded that the delta (dB) between the RMS and peak levels is similar for the speech and music content. In some examples, based on these results, the pmd value for the energy distribution module 1520 can be fixed, e.g., at 10dB. In some other examples, a 2nd-order fitting curve can be used to fit and approximate the delta values, e.g., as indicated in FIGS. 16D and 16H.
[0080] FIGS. 17A-17C graphically illustrate gain calculations performed in the gain calculation module 1550 of the method 1500 according to one example. More specifically, FIG. 17A graphically shows a signal transfer curve 1702. For the signal transfer curve 1702, the output energy can be represented as: output_E = max_EdB. FIG. 17B graphically shows a dB-gain curve 1704 corresponding to the signal transfer curve 1702. FIG. 17C graphically shows a linear-gain curve 1706 corresponding to the signal transfer curve 1702.
[0081] FIGS. 18A-18C graphically illustrate gain calculations performed in the gain calculation module 1550 of the method 1500 according to another example. More specifically, FIG. 18A graphically shows a signal transfer curve 1802. For the signal transfer curve 1802, the output energy can be represented as: output_E = max_EdB + (E_data_dB + pmd − max_EdB) / 5 (8) where E_data_dB is the energy of a frame; and the factor 5 represents the compression rate. In some examples, the compression rate can be modified by user. The quantity max_EdB is obtained as follows: max_EdB = EdB_limit + target_dB – pmd (9) where target_dB is the value set by the user. In some examples, the value of target_dB equals to the loudness to which the user wants to boost the signal after the DRC processing. In the example shown, target_dB is set to 15 dB. EdB_limit is a limiting value that enables the system to avoid excessively loud sounds. In the example shown, EdB_limit is set to –3dB.
[0082] FIG. 18B graphically shows a dB-gain curve 1804 corresponding to the signal transfer curve 1802. Based on the above description of the signal transfer curve 1802, then dB-gain curve 1804 can be calculated as follows: delta_gain_dB = output_E – E_data_dB (10) FIG. 18C graphically shows a linear-gain curve 1806 corresponding to the signal transfer curve 1802. Based on Eq. (10), linear-gain curve 1806 can be calculated as follows: gain_t = pow(10, delta_gain_dB / 20) (11)
[0083] FIGS. 19A-19C graphically illustrate gain calculations performed in the gain calculation module 1550 of the method 1500 according to yet another example. More specifically, FIG. 19A graphically shows a signal transfer curve 1902. The signal transfer curve 1902, consists of a linear transfer portion 1902a and a compression curve portion 1902b. The linear transfer portion 1902a can be described with y=x, where y and x represent the output and input signals, respectively. The compression curve portion 1902b can be described with y=ax2. The portions 1902a and 1902b are connected such that that the first derivative of the signal transfer curve 1902 is continuous. FIG. 19B graphically shows a dB-gain curve 1904 corresponding to the signal transfer curve 1902. FIG. 19C graphically shows a linear-gain curve 1906 corresponding to the signal transfer curve 1902.
[0084] In some examples, the postprocessing module 1560 is configured to apply the attack and release control to the gain to smooth gain fluctuations. The attack time (Ta) is the time it takes the compressor gain to rise from 10% to 90% of its final value when the input goes above the threshold. The release time (Tr) is the time it takes the compressor gain to drop from 90% to 10% of its final value when the input goes below the threshold. In one example, the attack and release control is implemented in the postprocessing module 1560 in accordance with Eqs. (12)-(13): ML∗XL ]R^_^M`abc^%dM`abrelease. / 1^^ / W = F HYZ[(\)(12) The resulting gain is^^^^^[^] = (1 − ∗ [^] + ∗ [^ c j] ^S ^ >0 (14a)^^^^^[^] = (1 − ^FUF^DFh%i^R%$) ∗ ^^^^^[^] + ^FUF^DFh%i^R%$ ∗ ^^^^^[^ c j]^S ^ >0 (14b)In some examples, Ta = 0.05 and Tr = 0.2. In some examples, to avoid excessively aggressive compression, one can set maximum compression by imposing a limit on the amount of compression, for example, as follows:^^^^_A[^^^^_A < 0.3162] = 0.3162 # DFA AℎF ^^^^^C^ ^^^^ S!^ #!^G^FDD^!^ 10Bs ^!t(15)
[0085] In some examples, the two-channel gain mapping module 1570 is configured to duplicate the mono gain received from the postprocessing module 1560 to obtain two respective gains for the L and R channels. One benefit of this particular approach is that the gain is relatively smooth and does not exhibit inter-channel inconsistencies. A different mono-to-stereo gain mapping approach may need to be applied in cases where there is a relatively large energy difference between the channels of the two-channel input 1502.
[0086] FIGS. 20A-20F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a first embodiment of the DRC method 1500 according to some examples. The loudness statistics of the corresponding input signals is shown in FIGS. 14A-14F, respectively. The first embodiment of the DRC method 1500 is configured to use gain calculations described above in reference to FIGS. 17A-17C.
[0087] FIGS. 21A-21F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a second embodiment of the DRC method 1500 according to some examples. The loudness statistics of the corresponding input signals is shown in FIGS. 14A-14F, respectively. The second embodiment of the DRC method 1500 is configured to use gain calculations described above in reference to FIGS. 18A-18C.
[0088] FIGS. 22A-22F graphically illustrate the loudness statistics corresponding to different sources of audio signals after being processed with a third embodiment of the DRC method 1500 according to some examples. The loudness statistics of the corresponding input signals is shown in FIGS. 14A-14F, respectively. The third embodiment of the DRC method 1500 is configured to use gain calculations described above in reference to FIGS. 19A-19C. Content Enhancement and Optimization
[0089] In some examples, the content enhancement blocks 2601and 2602of the workflow 200 are configured to use a different respective profile for each content type to improve the audio and video quality accordingly. Provided below are several nonlimiting examples of such different profiles for content processing. Based on the provided description, a person of ordinary skill in the pertinent art will be able to make and use various additional content-type profiles without any undue experimentation.
[0090] An example “music” profile is configured to emphasize the following content- enhancement operations: • Remove stationary noise. • Optimize audio for the music content by rebalancing the levels of percussive and harmonic components. • Perform DRC for percussive and harmonic components separately to avoid clipping while protecting the original relative level difference between those components. • Perform beat enhancement.
[0091] An example “person” profile is configured to emphasize some or all of the following content-enhancement operations: • Perform speech enhancement. • Spatialize speech sources according to face detection results in video (if any).• Increase volume of speech when zooming in. • Reduce background noise. • Attenuate percussive sound to avoid high energy impulse sound.
[0092] An example “nature landscape” profile is configured to emphasize some or all of the following content-enhancement operations: • Use less aggressive denoising to protect the sounds of water, birds, waterfalls, thunderstorms, and other nature sound effects according to the content type. • Reduce background noise only. • Rebalance the speech level and natural sound if speech is present and make the speech more intelligible while also keeping the natural sounds to preserve immersive effects.
[0093] An example “sports” profile is configured to emphasize some or all of the following content-enhancement operations: • Suppress percussive sounds (hits on or of the ball, etc.) if such sounds are deemed too loud. • Remove stationary noise. • Suppress babble / crowded noise accordingly when speech is present.
[0094] An example “urban and / or traffic” profile is configured to emphasize some or all of the following content-enhancement operations: • Keep city sounds (such as cars, vehicles, ambulance, police, etc.). • Remove stationary noise only. • Perform wind noise management.
[0095] An example “cooking and / or food sharing” profile is configured to emphasize some or all of the following content-enhancement operations: • Suppress percussive sound if such sound is deemed too loud. • Reduce background noise.
[0096] An example “animal” profile is configured to emphasize some or all of the following content-enhancement operations: • Enhance animal sounds. • Spatialize animal sounds. • Perform sound level management when zooming in.• Remove background noise.
[0097] FIG. 23 is a flowchart illustrating a method 2300 of enhancing audio content according to some examples. In various examples, the method 2300 can be implemented using the audio system 100 and / or the workflow 200.
[0098] The method 2300 includes analyzing the audio content with a machine-learning-based classifier (in a block 2302). The classifier is configured to assign to the audio content a type identifier selected from a plurality of type identifiers. In some examples, the classifier is configured to consider metadata corresponding to the audio content, with the metadata providing additional (supplemental) information about the audio content. In some examples, the classifier is configured to consider a video corresponding to the audio content, e.g., when the audio content is the soundtrack of a video clip. In some examples, the classifier includes a neural network trained using a contrastive loss function. In some examples, the classifier includes a neural network trained using a contrastive loss function. In various examples, the classifier is configured to use a segment- feature-based classification method or a frame-feature-based classification method. In some examples, the classifier is selected from the group consisting of an AdaBoost-type classifier, a convolutional-neural-network-based classifier, a deep-neural-network-based classifier, a recurrent- neural-network-based classifier, and a random forest classifier. In some examples, the type identifier is selected from a plurality of type identifiers, each of which is associated with a respective content descriptor selected from the group consisting of: person, animal, urban, vehicle, nature landscape, sport, music, cooking, and food sharing.
[0099] The method 2300 also includes applying respective denoising processing to each of one or more signal components of the audio content (in a block 2304). The respective denoising processing performed in the block 2304 is based on a respective denoising method selected based on the type identifier from a plurality of denoising methods. In some examples, the plurality of denoising methods includes at least one non-music denoising method and at least one music denoising method. Example denoising methods that can be used in the block 2304 are described in more detail above, e.g., in the subsection entitled “Denoising.”
[0100] In some examples, for a non-music signal component, the respective denoising processing performed in the block 2304 includes preserving only a speech component part and removing all other component parts of the signal component. For a music signal component, therespective denoising processing performed in the block 2304 includes removing only stationary noise.
[0101] The method 2300 also includes separating percussive, harmonic, and noise floor parts of a music signal component, if any (in a block 2306). In some examples, operations of the block 2306 include: (i) applying a Fourier transform to segments of the music signal component to obtain a time sequence of frequency bins; (ii) obtaining a time sequence of frequency bands by combining sets of the frequency bins; and (iii) classifying different ones of the frequency bands as harmonic bands or percussive bands based on anisotropic smoothness characteristics of the frequency bands in the time and frequency dimensions. Operations of the block 2306 may further include generating an estimate of the harmonic part of the music signal component using the harmonic bands and generating an estimate of the percussive part of the music signal component using the percussive bands.
[0102] The method 2300 also includes assigning respective levels of importance to audio objects of the one or more signal components (in a block 2308). The respective levels of importance are selected in the block 2308 from a plurality of predefined levels. In some examples, the plurality of levels includes first, second, third, and fourth levels of importance, such as the above-described EI, HI, MI, and LI levels. In some examples, a same sound present in different types of audio content may be assigned different levels of importance in the respective different instances of the block 2308.
[0103] The method 2300 also includes applying respective dynamic range controls to separated parts of the one or more signal components of the audio content (in a block 2310). The respective dynamic range controls are selected in the block 2310 based on the type identifier of each of the one or more signal components and further based on classification of the separated parts. In some examples, the respective dynamic range controls are selected in the block 2310 from the group consisting of a first dynamic range control characterized by a first signal transfer curve, a second dynamic range control characterized by a second signal transfer curve, and a third dynamic range control characterized by a third signal transfer curve. Examples of such first, second, and third signal transfer curves are illustrated in FIGS. 17-19, respectively.
[0104] The method 2300 also includes performing content enhancement on the audio content (in a block 2312). The performed content enhancement is based on the type identifier of the audio content assigned in the block 2302 and the respective levels of importance assigned in the block2308. Operations of the block 2312 include applying a different respective enhancement profile to each audio content type. Each of such different respective enhancement profiles is selected in accordance with the type identifier. In various examples, the different respective enhancement profiles specify different respective sets of content enhancement operations. Nonlimiting examples of such sets are described above in the subsection entitled “Content Enhancement and Optimization.” In some examples, the content enhancement is performed using the estimates of the harmonic and percussive parts of the music signal generated in the block 2306.
[0105] The method 2300 also includes playing the enhanced audio content (in a block 2314). In some examples, the enhanced audio content is played using the audio rendering component 150 of the audio system 100. In configurations of the audio system 100 in which the content enhances is located at the transmitter side, operations of the block 2314 also include: (i) generating the bitstream 132 by encoding the enhanced audio content with the encoder 120; (ii) transmitting the bitstream 132 through the communication channel 130 to the decoder 140; and (iii) reconstructing the enhanced audio content from the bitstream 132 at the decoder 140. Example Hardware
[0106] FIG. 24 is a block diagram of an example computing device 2400 according to various examples. In some examples, the computing device 2400 is configured to perform at least some operations of the methods 2300, 1500 and the workflow 200. In some examples, two or more instances of the computing device 2400 are used in the audio system 100.
[0107] The computing device 2400 of FIG. 24 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 2400 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 2402 and one or more storage devices 2404). Additionally, in various embodiments, the computing device 2400 may not include one or more of the components illustrated in FIG. 24, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, awireless interface, or any other appropriate interface). For example, the computing device 2400 may not include a display device 2410, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 2410 may be coupled.
[0108] The computing device 2400 includes a processing device 2402 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 2402 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.
[0109] The computing device 2400 also includes a storage device 2404 (e.g., one or more storage devices). In various embodiments, the storage device 2404 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 2404 may include memory that shares a die with the processing device 2402. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 2404 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 2402), cause the computing device 2400 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.
[0110] The computing device 2400 further includes an interface device 2406 (e.g., one or more interface devices 2406). In various embodiments, the interface device 2406 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 2400 and other computing devices. For example, the interface device 2406 may include circuitry for managing wireless communications for thetransfer of data to and from the computing device 2400. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 2406 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 2406 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 2406 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 2406 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 2406 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.
[0111] In some embodiments, the interface device 2406 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 2406 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 2406 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 2406 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 2406 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 2406 may be dedicated to wireless communications, and a second set of circuitry of the interface device 2406 may be dedicated to wired communications.
[0112] The computing device 2400 also includes battery / power circuitry 2408. In various embodiments, the battery / power circuitry 2408 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 2400 to an energy source separate from the computing device 2400 (e.g., to AC line power).
[0113] The computing device 2400 also includes a display device 2410 (e.g., one or multiple individual display devices). In various embodiments, the display device 2410 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0114] The computing device 2400 also includes additional input / output (I / O) devices 2412. In various embodiments, the I / O devices 2412 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.
[0115] Depending on the specific embodiment, various components of the interface devices 2406 and / or I / O devices 2412 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 2406 and / or I / O devices 2412 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 2402 and / or the storage device 2404. In some additional examples, the interface devices 2406 and / or I / O devices 2412 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 2402 and / or the storage device 2404 into an analog form suitable for being transmitted through a communication channel.
[0116] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs. 1-24, provided is an audio system for enhancing audio content comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: analyze the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; apply respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assign respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and perform content enhancement on the audio content based on the type identifier and the respective levels of importance.
[0117] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs. 1-24, provided is a method of enhancing audio content comprising: analyzing the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; applying respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assigning respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and performing content enhancement on the audio content based on the type identifier and the respective levels of importance.
[0118] In some embodiments of the above method, the analyzing includes considering metadata corresponding to the audio content.
[0119] In some embodiments of any of the above methods, the analyzing includes considering a video corresponding to the audio content.
[0120] In some embodiments of any of the above methods, the machine-learning-based classifier includes a neural network trained using a contrastive loss function.
[0121] In some embodiments of any of the above methods, the machine-learning-based classifier is configured to use a segment-feature-based classification method or a frame-feature- based classification method.
[0122] In some embodiments of any of the above methods, the machine-learning-based classifier is selected from the group consisting of an AdaBoost-type classifier, a convolutional- neural-network-based classifier, a deep-neural-network-based classifier, a recurrent-neural-network- based classifier, and a random forest classifier.
[0123] In some embodiments of any of the above methods, the type identifier is selected from a plurality of type identifiers, each of which is associated with a respective content descriptor selected from the group consisting of: person, animal, urban, vehicle, nature landscape, sport, music, cooking, and food sharing.
[0124] In some embodiments of any of the above methods, the plurality of denoising methods includes at least one non-music denoising method and at least one music denoising method.
[0125] In some embodiments of any of the above methods, for a non-music signal component, the respective denoising processing includes preserving only a speech component part and removing all other component parts.
[0126] In some embodiments of any of the above methods, for a music signal component, the respective denoising processing includes removing only stationary noise.
[0127] In some embodiments of any of the above methods, the method further comprises separating percussive, harmonic, and noise floor parts of the music signal component.
[0128] In some embodiments of any of the above methods, the separating comprises: applying a Fourier transform to segments of the music signal component to obtain a time sequence of frequency bins; obtaining a time sequence of frequency bands by combining sets of the frequency bins; and classifying different ones of the frequency bands as harmonic bands or percussive bands based on anisotropic smoothness characteristics of the frequency bands in time and frequency dimensions.
[0129] In some embodiments of any of the above methods, the separating further comprises: generating an estimate of the harmonic part of the music signal component using the harmonicbands; and generating an estimate of the percussive part of the music signal component using the percussive bands.
[0130] In some embodiments of any of the above methods, the content enhancement is performed using the estimates.
[0131] In some embodiments of any of the above methods, the plurality of levels includes first, second, third, and fourth levels of importance.
[0132] In some embodiments of any of the above methods, the method further comprises applying respective dynamic range controls to separated parts of the one or more signal components of the audio content, the respective dynamic range controls being selected based on the type identifier of each of the one or more signal components and further based on classification of the separated parts.
[0133] In some embodiments of any of the above methods, the respective dynamic range controls are selected from the group consisting of a first dynamic range control characterized by a first signal transfer curve, a second dynamic range control characterized by a second signal transfer curve, and a third dynamic range control characterized by a third signal transfer curve; wherein the first signal transfer curve is different from each the second signal transfer curve and the third signal transfer curve; and wherein the second signal transfer curve is different from the third signal transfer curve.
[0134] In some embodiments of any of the above methods, the performing comprises applying a different respective enhancement profile to each audio content type, said different respective enhancement profile being selected in accordance with the type identifier.
[0135] In some embodiments of any of the above methods, the different respective enhancement profiles specify different respective sets of content enhancement operations.
[0136] In some embodiments of any of the above methods, the method further comprises playing the enhanced audio content using an audio rendering system component.
[0137] In some embodiments of any of the above methods, the method further comprises: generating a bitstream by encoding the enhanced audio content with an encoder; transmitting thebitstream through a communication channel to a decoder; reconstructing the enhanced audio content from the bitstream received by the decoder; and playing the reconstructed enhanced audio content using an audio rendering system component.
[0138] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.
[0139] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0140] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0141] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0142] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used tointerpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
[0143] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.
[0144] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0145] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non- transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.
[0146] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.
[0147] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0148] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0149] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”
[0150] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.
[0151] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”
[0152] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition ofone or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
[0153] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.
[0154] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0155] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) foroperation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0156] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0157] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
CLAIMS What is claimed is:
1. A method of enhancing audio content, comprising: analyzing the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; applying respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assigning respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and performing content enhancement on the audio content based on the type identifier and the respective levels of importance.
2. The method of claim 1, wherein the analyzing includes considering metadata corresponding to the audio content.
3. The method of claim 1 or 2, wherein the analyzing includes considering a video corresponding to the audio content.
4. The method of any preceding claim, wherein the machine-learning-based classifier includes a neural network trained using a contrastive loss function.
5. The method of any preceding claim, wherein the machine-learning-based classifier is configured to use a segment-feature-based classification method or a frame-feature-based classification method.
6. The method of any preceding claim, wherein the machine-learning-based classifier is selected from the group consisting of an AdaBoost-type classifier, a convolutional-neural-network- based classifier, a deep-neural-network-based classifier, a recurrent-neural-network-based classifier, and a random forest classifier.
7. The method of any preceding claim, wherein the type identifier is selected from a plurality of type identifiers, each of which is associated with a respective content descriptor selected from the group consisting of: person, animal, urban, vehicle, nature landscape, sport, music, cooking, and food sharing.
8. The method of any preceding claim, wherein the plurality of denoising methods includes at least one non-music denoising method and at least one music denoising method.
9. The method of any preceding claim, wherein, for a non-music signal component, the respective denoising processing includes preserving only a speech component part and removing all other component parts.
10. The method of any preceding claim, wherein, for a music signal component, the respective denoising processing includes removing only stationary noise.
11. The method of claim 10, further comprising separating percussive, harmonic, and noise floor parts of the music signal component.
12. The method of claim 11, wherein the separating comprises: applying a Fourier transform to segments of the music signal component to obtain a time sequence of frequency bins; obtaining a time sequence of frequency bands by combining sets of the frequency bins; and classifying different ones of the frequency bands as harmonic bands or percussive bands based on anisotropic smoothness characteristics of the frequency bands in time and frequency dimensions.
13. The method of claim 12, wherein the separating further comprises: generating an estimate of the harmonic part of the music signal component using the harmonic bands; and generating an estimate of the percussive part of the music signal component using the percussive bands.
14. The method of any one of claims 10 to 13, being configured to perform at least one of beat extraction and rhythm detection of the music signal component.
15. The method of claim 14, wherein the beat extraction and rhythm detection further comprises: applying a Fourier transform to the music signal component; and applying frequency filtering to the transformed music signal component to output selected filtered signals as extracted beat signals.
16. The method of claim 15, further comprising: combining frequency bins into bands; performing one or more of: beat detection in the frequency domain; signal smoothing operations; beat detection in the time domain; normalization operations; Gaussian filtering operations; low peak filtering operations; beat period calculations; and beat modeling or approximation.
17. The method of claim 13 or 14, wherein the content enhancement is performed using the estimates.
18. The method of any preceding claim, wherein the plurality of levels includes first, second, third, and fourth levels of importance.
19. The method of any preceding claim, further comprising applying respective dynamic range controls to separated parts of the one or more signal components of the audio content, the respective dynamic range controls being selected based on the type identifier of each of the one or more signal components and further based on classification of the separated parts.
20. The method of claim 19, wherein the respective dynamic range controls are selected from the group consisting of a first dynamic range control characterized by a first signal transfer curve, a second dynamic range control characterized by a second signal transfer curve, and a third dynamic range control characterized by a third signal transfer curve; wherein the first signal transfer curve is different from each the second signal transfer curve and the third signal transfer curve; and wherein the second signal transfer curve is different from the third signal transfer curve.
21. The method of claim 20, wherein each of the signal transfer curves are obtained by: receiving a two-channel input signal and converting the received two-channel input signal into a mono signal, calculating a mono gain based on an energy distribution of the mono signal and a target loudness, and duplicating the mono gain to obtain two respective gains for L and R channels.
22. The method of claim 20, wherein each of the signal transfer curves are obtained by: receiving a two-channel input signal and converting the received two-channel input signal into a mono signal (mono) according to either: mono = 0.5L + 0.5R for L and R channels, or mono = 0.5E(L) + 0.5E(R), wherein E(L) and E(R) denote energy for the L and R channels, respectively.
23. The method of any preceding claim, wherein the performing comprises applying a different respective enhancement profile to each audio content type, said different respective enhancement profile being selected in accordance with the type identifier.
24. The method of claim 23, wherein the different respective enhancement profiles specify different respective sets of content enhancement operations.
25. The method of any preceding claim, further comprising playing the enhanced audio content using an audio rendering system component.
26. The method of any preceding claim, further comprising: generating a bitstream by encoding the enhanced audio content with an encoder; transmitting the bitstream through a communication channel to a decoder;reconstructing the enhanced audio content from the bitstream received by the decoder; and playing the reconstructed enhanced audio content using an audio rendering system component.
27. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1-26.
28. An audio system for enhancing audio content, the audio system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: analyze the audio content with a machine-learning-based classifier to assign to the audio content a type identifier selected from a plurality of type identifiers; apply respective denoising processing to each of one or more signal components of the audio content, the respective denoising processing being based on a respective denoising method selected based on the type identifier from a plurality of denoising methods; assign respective levels of importance to audio objects of the one or more signal components, the respective levels of importance being selected from a plurality of levels; and perform content enhancement on the audio content based on the type identifier and the respective levels of importance.
Citation Information
Patent Citations
Equalizer controller and controlling method
EP3232567A1
Data driven audio enhancement
US20210217436A1
Multi-source audio processing systems and methods
US20230115674A1
Cited By
Eating event detection and food intake identification method and device for ear-mounted equipment and storage medium
CN121415817A
Real-time content-adaptive headphone amplifier with dynamic power consumption and psychoacoustic correction
US20260136135A1