METHOD AND DEVICES FOR USER-GENERATED CONTENT CAPTURE AND ADAPTIVE PLAYBACK

DE602023013028T2Active Publication Date: 2026-03-04DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602023013028
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-29
Filing Date
2023-04-03
Publication Date
2026-03-04
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

Existing user-generated content (UGC) captured by mobile devices often suffers from sound quality issues due to hardware limitations and diverse recording environments, and real-time enhancements may not be compatible with end-to-end content processing, leading to suboptimal user experiences.

Method used

A method and apparatus for UGC processing that applies frame-wise audio enhancements during capture, storing metadata for further enhancements, enabling adaptive rendering based on device capabilities and long-term statistics, allowing for improved audio quality on playback devices with or without additional software support.

Benefits of technology

Enhances audio quality of UGC by providing context metadata for adaptive rendering, ensuring optimal listening experiences across various devices and environments, and enabling end-to-end content processing.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 336,700, filed April 29, 2022, and International Application No. PCT / CN2022 / 085777, filed April 8, 2022.TECHNICAL FIELD

[0002] The present document relates to methods, apparatus, and systems for capture and adaptive rendering of user generated content (UGC). The present document particularly relates to UGC content creation on mobile devices that enables adaptive rendering during playback, and to adaptive rendering during playback.BACKGROUND

[0003] Recently, UGC has become a trend of personal moment sharing in variable environments. UGC is mostly recorded by mobile devices. Most of this content will have sound artifacts due to consumer hardware limitation, system performance requirements, diversity of capture practices, and playback environment.

[0004] To overcome sound quality issues introduced by hardware limitations and recording environment, UGC audio may be enhanced for better listening experience. Certain audio enhancements could be applied in real-time during or immediately after capture, with the information available at that time. Such enhancements can be applied to the audio stream directly and generate enhanced audio streams in real-time. The enhanced audio can then be rendered without specific software support on playback devices. Thereby, UGC content creators could improve audio quality of their content without additional effort and make sure that such enhancement would be available to their content consumers to the greatest extent.

[0005] However, there are also audio enhancements that rely on additional information beyond what is available in real-time, for further enhanced audio quality. Moreover, the real-time enhancement following capture may not be compatible with end-to-end content processing and user experience.

[0006] Thus, there is a need for improved techniques for UGC capture and adaptive rendering.

[0007] US 2013 / 246077 provides techniques for adaptive processing of media data based on separate data specifying a state of the media data. A device in a media processing chain may determine whether a type of media processing has already been performed on an input version of media data. If so, the device may adapt its processing of the media data to disable performing the type of media processing. If not, the device performs the type of media processing. The device may create a state of the media data specifying the type of media processing. The device may communicate the state of the media data and an output version of the media data to a recipient device in the media processing chain, for the purpose of supporting the recipient device's adaptive processing of the media data.

[0008] US 2017 / 309286 A1 discloses the reconstruction, on the basis of a bitstream (P), of an n-channel audio signal (X) by deriving an m-channel core signal (Y) and multichannel coding parameters (α) from the bitstream, where 1≦m < n. Also derived from the bitstream are pre-processing dynamic range control, DRC, parameters (DRC2) quantifying an encoder-side dynamic range limiting of the core signal. The n-channel audio signal is obtained by parametric synthesis in accordance with the multichannel coding parameters and while cancelling any encoder-side dynamic range limiting based on the pre-processing DRC parameters. In particular embodiments, the reconstruction further includes use of compensated post-processing DRC parameters quantifying a potential decoder-side dynamic range compression. Cancellation of an encoder-side range limitation and range compression are preferably performed by different decoder-side components. Cancellation and compression may be coordinated by a DRC pre-processor.SUMMARY

[0009] According to an aspect, a method of processing audio data relating to user generated content is provided according to claim 1.

[0010] Configured as referred above, the proposed method can provide enhanced audio data that is suitable for direct playback by a playback device, without further audio processing by the playback device. On the other hand, the method also provides context metadata for the enhanced audio data. This context metadata allows to restore the raw audio for additional / alternative audio enhancement by a playback device with different (e.g., better) processing capabilities, or for audio editing with an editing tool. Thereby, rendering at the playback device can be performed in an adaptive manner, depending on the device's hardware capabilities, the playback environment, user-specific settings, etc.. In other words, providing the context metadata allows for end-to-end content processing from capture to playback, taking into account characteristics of the specific capture and rendering hardware, specific environments, user preferences, etc., thereby enabling optimal enhancement of the audio data and listening experience.

[0011] Further preferable aspects are defined by the claims depending on claim 1. As such, the UGC generated by the proposed method is particularly applicable to being consumed by mobile devices with typically limited processing capabilities, for example in a streaming framework for devices without specific software support for reading metadata. On the other hand, if the device in the streaming framework has specific software support for reading the metadata, the metadata and enhanced audio data may be read, raw audio may be generated / restored from the enhanced audio data using the metadata, and further enhanced audio may be generated based on the raw audio.

[0012] According to another aspect, a method of processing audio data relating to user generated content according to claim 6 is provided.

[0013] By performing the referred method, restoring the raw audio data, a replay / editing device applies audio enhancement or audio editing depending on long term statistics of the (non enhanced) audio data. Thereby, end-to-end content processing and optimum user experience can be achieved. On the other hand (in an aspect not encompassed by the claimed invention), if processing capabilities should not be sufficient for audio enhancement, the received enhanced audio data can be directly rendered without additional processing.

[0014] Preferable aspects are defined by the claims depending on claim 6.

[0015] According to another aspect, an apparatus for processing audio data relating to user generated content according to claim 11 is provided.

[0016] According to another aspect, an apparatus for processing audio data relating to user generated content according to claim 12 is provided.

[0017] According to another aspect, an apparatus for processing audio data relating to user generated content accordingt to claim 13 is provided.

[0018] According to a further aspect, a computer program according to claim 14 is provided.

[0019] According to another aspect, a computer-readable storage medium according to claim 15 is provided.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein Fig. 1 illustrates a conceptual diagram of an example apparatus for UGC processing during / after capture according to embodiments of the disclosure; Fig. 2 is a flowchart illustrating an example method of UGC processing during / after capture according to embodiments of the disclosure; Fig. 3 illustrates a conceptual diagram of an example apparatus for UGC processing for rendering; Fig. 4 illustrates a conceptual diagram of an example apparatus for UGC processing for rendering according to embodiments of the disclosure; Figs. 5 is a flowchart illustrating an example method of UGC processing for rendering according to embodiments of the disclosure; and Fig. 6 illustrates a conceptual diagram of an example computing apparatus for performing techniques according to embodiments of the disclosure. DETAILED DESCRIPTION

[0021] Broadly speaking, the present disclosure relates to methods, apparatus, and systems for UGC content creation, for example on mobile devices, enabling adaptive rendering based on information available at a playback device, and to methods, apparatus, and systems for UGC adaptive rendering.

[0022] Real-time audio enhancement at a capture side can yield enhanced audio content that could be rendered without specific support at a playback device. On the other hand, also more sophisticated audio enhancements exist that rely on additional information beyond what is available in real-time, for further enhanced audio quality. According to techniques described herein, such further audio enhancements can be applied to the audio stream, usually stored as metadata along with the audio stream, after the audio capture and real-time enhancement process is finished. With a playback device capable of reading this metadata, the further audio enhancements could be applied in audio content rendering, or in audio editing. Accordingly, techniques described herein can further improve audio quality of UGC for certain content consumers with specific playback devices capable of reading the metadata, or for all content consumers after editing the content with software tools capable of reading the metadata.

[0023] At a conceptual level, a capture and rendering ecosystem according to embodiments of the disclosure may be composed of or characterized by some or all of the following elements: A binaural capture device that can record at least two channel recordings and a playback device that can render the at least two channel recordings. The recording device and the playback device can be the same device, two connected devices, or two separate devices. The capture device comprises a processing module for enhancing the captured audio in real time. The processing may comprise at least one of level adjustment, dynamic range control, noise management, and timbre management. The capture device comprises an analysis module for providing long-term statistics, possibly further providing file-based features and context information, from the audio recording. The analysis results will be stored as context metadata, alongside the enhanced audio content generated by the processing module. The metadata comprises frame-by-frame analysis results, which may include at least the band gains or full-band gains applied by one or more components of the processing module, as well as file-based global results, which include at least one of the loudness, content type, etc. of the audio, and the context information. During playback, the rendering is adaptive based on the availability of the context metadata. In one case, the playback device only has access to the enhanced audio, thus during playback it will render the enhanced audio directly without processing, or process it without the help of context metadata. In another case, the playback device has access to both the enhanced audio and the context metadata. During playback, the playback device will further process the enhanced audio based on the context metadata, for improved listening experience. The capture device and / or the playback device may also feature an editing tool. When the editing tool has access to the context metadata, the editing of the enhanced audio would generate comparable results as compared to the editing results of the raw audio.

[0024] Fig. 1 schematically illustrates an apparatus (e.g., device, system) 100 for processing audio data 105 relating to UGC. Apparatus 100 relates to a capture side for UGC, and as such may correspond to or be included in a mobile device (e.g., mobile phone, tablet computer, PDA, laptop computer, etc.). Processing performed by apparatus 100 may enable adaptive rendering at a rendering or playback device. The apparatus 100 comprises a processing module 110 and an analysis module 120. Optionally, the apparatus 100 may further comprise a capturing module for capturing the audio data 105 (not shown). The capturing module (or capturing device) may be a binaural capturing device, for example, that can record at least two channel recordings.

[0025] The processing module 110 is adapted to apply frame-wise audio enhancement to the audio data 105. This frame-wise audio enhancement is applied during or immediately following capture of the UGC. As a result of the frame-wise audio enhancement, enhanced audio data 115 is obtained and output by the processing module 110. The processing module 110 performs the aforementioned audio enhancements, which are then applied during or immediately following capture of the UGC. Thereby, the processing module 110 generates the enhanced audio data 115 (enhanced audio) that could be rendered without specific support at a playback device.

[0026] Specifically, the processing module 110 may be configured to apply, to the audio data 105, at least one of noise management, loudness management, peak limiting, and timbre management. Accordingly, the processing module 110 in the example apparatus 100 of Fig. 1 comprises a noise management module 130, a loudness management module 140, and a peak limiting module 150. An optional timbre management module is not shown in the figure. It is noted that not all of the aforementioned modules for audio enhancement may be present, depending on the specific application.

[0027] The audio enhancements performed by the processing module 110 are based on respective processing parameters. For example, there may be distinct (sets of) processing parameters for each of noise management, loudness management, peak limiting, and timbre management, if present. As described in more detail below, the processing parameters include band gains and / or full-band gains that are applied during the frame-wise audio enhancement. The band gains or full-band gains may comprise respective gains for each frame of the audio data. Further, the band gains or full-band gains may comprise respective gains for each type of enhancement processing that is applied.

[0028] The noise management module 130 may be adapted for applying noise management, involving suppressing disturbing noises that are oftentimes present in non-professional recording environments. As such, noise management may relate to de-noising, for example. The noise management module 130 may be implemented, for example, by machine learning algorithms or neural networks including recurrent neural networks (RNNs) or convolutional neural networks (CNNs), the implementation details of which are understood to be readily apparent to experts in the field. Further, noise management may involve pitch filtering.

[0029] Processing parameters for noise management may include band gains (e.g., a plurality of band gains) for noise management. These band gains may relate to gains in respective ones among a plurality of frequency bands (e.g., frequency subbands). Further, there may be one such band gain per frame and frequency band. In case of pitch filtering, the processing parameters for noise management may include filter parameters for pitch filtering, such as a center frequency of the filter, for example.

[0030] The loudness management module 140 may be adapted for applying loudness management, involving leveling of the input audio stream (i.e., the audio data 105) to a certain loudness range. Loudness management may relate to level adjustment and / or range control. For example, the input audio stream may be leveled to a loudness range more suitable for later playback by a playback device. As such, the loudness management may adjust the loudness of the audio stream to an appropriate range for better listening experience.

[0031] It may be implemented by automatic gain control (AGC), dynamic range control (DRC), or a combination of the two, the implementation details of which are understood to be readily apparent to experts in the field.

[0032] Processing parameters for loudness management may include gains for loudness management. These gains may relate to full-band gains that uniformly apply to the full frequency range, i.e. apply uniformly to the plurality of frequency bands (e.g., frequency subbands). There may be one such gain per frame.

[0033] The peak limiting module 150 may be adapted for applying peak limiting, involving ensuring that the amplitude of the input audio after enhancements will not exceed a legitimate range allowed by audio storage, distribution, and / or playback. Implementation details again are understood to be readily apparent to experts in the field.

[0034] Processing parameters for peak limiting may include gains for peak limiting. These gains may relate to full-band gains that uniformly apply to the plurality of frequency bands (e.g., frequency subbands). There may be one such gain per frame.

[0035] The timbre management module (not shown) may be adapted for applying timbre management, involving adjusting timbre of the audio data 105.

[0036] Processing parameters for timbre management may include band gains (e.g., a plurality of band gains) for timbre management. These band gains may relate to gains in respective ones among a plurality of frequency bands (e.g., frequency subbands). Further, there may be one such band gain per frame and frequency band.

[0037] The processing module 110 provides one or more (e.g., a plurality of) processing parameters of the frame-wise audio enhancement to the analysis module 120. The processing parameters may be provided in a frame-wise manner. For example, updated values of the processing parameters may be provided for each frame or for each predefined multiple of frames (e.g., for every other frame, once for every N frames, etc.). The processing parameters may include any, some, or all of processing parameters 135 of noise management, processing parameters 145 of loudness management, processing parameters 155 of peak limiting, and processing parameters of timbre management (not shown).

[0038] As further input, the analysis module 120 receives (a version of) the audio data 105.

[0039] The analysis module 120 is adapted to generate metadata 125 (context metadata) for the enhanced audio data 115. Generating the metadata 125 is based on the one or more processing parameters of the frame-wise audio enhancement. For example, the metadata 125 may include the processing parameters (e.g., band gains and / or full-band gains) or an indication thereof.

[0040] The analysis module 120 is further adapted to output the metadata 125. In other words, the analysis module 120 analyzes the audio data 105 and / or the aforementioned audio enhancements performed by the processing module 110 to generate the context metadata 125 for audio enhancements that rely on additional information beyond the information that is available in real time, for further improved audio quality. The generated context metadata 125 can be utilized by specific playback devices or editing tools for better audio quality and user experience.

[0041] Based on the one or more processing parameters of the frame-wise audio enhancement, the analysis module 120 generates first metadata 165 (e.g., enhancement metadata) as part of the context metadata 125. The first metadata 165 are based on the processing parameters, as noted above.

[0042] In addition to the one or more processing parameters of the audio enhancement, the analysis module 120 generates the context metadata 125 further based on a result of analyzing multiple frames of the audio data 105. Such analysis of multiple frames of the audio data 105 (i.e., an analysis of the audio data 105 over time) yields long-term statistics (e.g., file-based statistics) of the audio data 105 (second metadata). Additionally, the analysis of multiple frames of the audio data 105 may yield one or more audio features of the audio data 105. Examples of audio features that may be determined in this manner include a content type of the audio data 105 (e.g., music, speech, movie, effects, etc.), an indication of a capturing environment of the audio data 105 (e.g., a quiet / noisy environment, an environment with / without echo or reverb, etc.), a signal-to-noise ratio, SNR, of the audio data 105, an overall loudness (e.g., file loudness) of the audio data 105, and a spectral shape (e.g., spectral envelope) of the audio data 105.

[0043] Based on the result of analyzing multiple frames of the audio data the analysis module 120 generates second metadata 175 (e.g., long-term metadata) as part of the context metadata 125. The second metadata 175 comprise the long-term statistics and may further comprise the audio features, or indications thereof, for example. The first and second metadata 165, 175 are compiled to obtain compiled metadata as the context metadata 125 for output. It is understood that the context metadata 125 include both of the first metadata 165, based on the one or more processing parameters, and the second metadata 175, based on the analysis of multiple frames of the audio data 105.

[0044] In the example of Fig. 1, the analysis module 120 comprises a processing statistics module 160, a long term statistics module 170, and a metadata compiler module 180 (metadata compiler).

[0045] The processing statistics module 160 implements the generation of the first metadata 165 based on the one or more processing parameters. It tracks the key parameters of the processing applied in the processing module 110, such that at a later time, for example during playback, the rendering system could have a better estimation of the raw audio (prior to capture side audio enhancement), based on an enhanced audio stream comprising the enhanced audio data 115 (enhanced audio data) and the metadata 125 (context metadata). As such, analysis of the one or more processing parameters of the audio enhancement by the processing statistics module may yield processing statistics of the audio enhancement performed by the processing module 110.

[0046] The long term statistics module 170 implements the generation of the second metadata 175 based on the analysis of multiple frames of the audio data 105 (i.e., long-term analysis of the audio data). It analyzes context information of the audio data 105 over a longer time span than allowed in real-time processing, for example over several frames or seconds, or over a whole file. In general, the statistics derived in this manner would be more accurate and stable than real-time statistics.

[0047] The metadata compiler module 180 finally gathers information from both the processing statistics module 160 and the long term statistics module 170 (e.g., the first and second metadata 165, 175) and compiles it into a specific format, so that the information can be retrieved at a later time with a metadata parser. In other words, the metadata compiler module 180 compiles the first and second metadata 165, 175 to obtain compiled metadata as the metadata 125 (context metadata) for output.

[0048] As a consequence of the above processing, the apparatus 100 outputs the enhanced audio data 115 together with the context metadata 125. The enhanced audio data 115 and the context metadata 125 may be output in a suitable format as an enhanced audio stream, for example. The enhanced audio stream may be used for adaptive rendering on playback devices, depending on the devices' capabilities, as described further below.

[0049] While an example apparatus 100 for UGC processing has been described above, the present disclosure likewise relates to corresponding methods of UGC processing. It is understood that any statements made above with regard to the apparatus 100 likewise apply to corresponding methods, and vice versa. An example of such method 200 of UGC processing (e.g., processing of audio data relating to UGC) is illustrated in the flowchart of Fig. 2. Method 200 comprises steps S210 through S240 and may be performed during / subsequent to capture of the UGC. It may be performed by a mobile device, for example.

[0050] At step S210, the audio data is obtained. Obtaining the audio data may include or amount to capturing the audio data by a suitable capturing device. The capturing device may be a binaural capturing device, for example, that can record at least two channel recordings.

[0051] At step S220, frame-wise audio enhancement is applied to the audio data to obtain enhanced audio data. This step corresponds to the processing of the processing module 110 described above. In general, applying the frame-wise audio enhancement to the audio data may include applying at least one of noise management (e.g., as performed by the noise management module 130), loudness management (e.g., as performed by the loudness management module 140), peak limiting (e.g., as performed by the peak limiting module 150), and timbre management (e.g., as performed by the timbre management module). Further, the frame-wise audio enhancement is applied during or immediately after capture of the audio data and may thus be referred to as real-time frame-wise audio enhancement.

[0052] At step S230, metadata (context metadata) is generated for the enhanced audio data, based on one or more processing parameters of the frame-wise audio enhancement. This step corresponds to the processing of the analysis module 120 described above. Accordingly, the one or more processing parameters may include band gains and / or full-band gains applied during the frame-wise audio enhancement. Specifically, the one or more processing parameters may include at least one of band gains for noise management, full-band gains for loudness management, full-band gains for peak limiting, and band gains for timbre management.

[0053] In addition to the one or more processing parameters, the metadata are generated further based on a result of analyzing multiple frames of (e.g., all of) the audio data (e.g., as performed by the long term statistics module 170). Therein, the analysis of multiple frames of the audio data yields long-term statistics (e.g., file-based statistics) of the audio data and may additionally yield one or more audio features of the audio data (e.g., a content type of the audio data, an indication of a capturing environment of the audio data, a signal-to-noise ratio of the audio data, an overall loudness of the audio data, and / or a spectral shape of the audio data, etc.).

[0054] Accordingly, the metadata comprise first metadata (e.g., enhancement metadata) generated based on the one or more processing parameters of the frame-wise audio enhancement (e.g., as generated by the processing statistics module 160) and second metadata (e.g., long-term metadata) generated based on the result of analyzing multiple frames of the audio data (including metadata generated by the long term statistics module 170). In such case, the first and second metadata may be compiled to obtain compiled metadata as the metadata for output (e.g., as done by the metadata compiler module 180).

[0055] At step S240, the enhanced audio data is output together with the generated metadata.

[0056] Next, possible implementations of processing the UGC at a replay or editing device will be described with reference to Fig. 3 to Fig. 5.

[0057] Fig. 3 illustrates a conceptual diagram of an example apparatus (e.g., device, system) 300 for UGC processing for rendering, such as a general audio rendering system for UGC.

[0058] The apparatus 300 comprises a rendering module 310 with a noise management module 320, a loudness management module 330, a timbre management module 340, and a peak limiting module 350. The apparatus 300 only takes the aforementioned enhanced audio data 305 as input and applies blind processing, without the help of any information other than the audio itself. The apparatus 300 finally outputs rendering output 315 for replay. Alternatively, the apparatus 300 may receive but disregard any context metadata that is provided along with the enhanced audio data 305.

[0059] Fig. 4 schematically illustrates an apparatus (e.g., device, system) 400 for processing enhanced audio data 405 relating to UGC (e.g., a rendering apparatus for UGC). Apparatus 400 relates to a replay side for UGC, and as such may correspond to or be included in a mobile device (e.g., mobile phone, tablet computer, PDA, laptop computer, etc.) or any other computing device. Contrary to the blind processing by apparatus 300, apparatus 400 is configured for context-aware processing of UGC, based on received context metadata.

[0060] Thus, in addition to enhanced audio 405 the apparatus 400 also takes the aforementioned context metadata 435 as input, which can be used to steer the rendering processing properly to generate a further enhanced rendering output 425. To this end, the apparatus 400 comprises a metadata parser 430 (e.g., as part of an input module) and several processing components. The processing components in this example may fall into two groups relating to "restore" and "rendering".

[0061] In general, the apparatus 400 comprises the input module (not shown) for receiving the (enhanced) audio data 405 and the (context) metadata 435 for the audio data, a processing module 410 for applying restore processing the audio data 405, and at least one of a rendering module 420 and an editing module (not shown). For example, the audio data 405 and the metadata 435 may be received in the form of a bitstream comprising the audio data 405 and the metadata 435, including retrieving the audio data 405 and the metadata 435 from a storage medium.

[0062] In the example of Fig. 4, the apparatus 400 comprises the metadata parser 430 (e.g., as part of the input module). The metadata parser 430 takes the context metadata 435 (e.g., generated by the aforementioned metadata compiler 180 of apparatus 100) as input.

[0063] In line with the above, the metadata 435 comprises first metadata 440 indicative of one or more processing parameters of a previous (earlier, e.g., capture side) frame-wise audio enhancement of the audio data. Additionally, the metadata 435 comprise second metadata 445 indicative of long-term statistics of the audio data and may comprise further metadata indicative of one or more audio features of the audio data (e.g., a content type of the audio data, an indication of a capturing environment of the audio data, a signal-to-noise ratio of the audio data prior to the previous frame-wise audio enhancement, an overall loudness of the audio data prior to the previous frame-wise audio enhancement, and / or a spectral shape of the audio data prior to the previous frame-wise audio enhancement, etc.). Therein, the statistics of the audio data and / or the audio features of the audio data could be based on the audio prior to or after the previous frame-wise audio enhancement, or even to audio data between two successive previous frame-wise audio enhancements, if applicable.

[0064] The metadata parser 430 retrieves information including processing statistics (e.g., the first metadata 440) and long-term statistics (e.g., the second metadata 445), which in turn are used to steer the processing components, such as the restore module 410, the rendering module 420, and / or the editing module.

[0065] The "restore" group of processing components generates (restored) raw audio from the enhanced audio with the help of the context metadata 435 ( the first metadata 440). Accordingly, the processing module 410 is configured for applying restore processing to the audio data 405, using the context metadata 435. Specifically, the processing module 410 uses the one or more processing parameters (e.g., as indicated by the first metadata 440), to at least partially reverse the previous frame-wise audio enhancement (as performed on the capture side). Thereby, the processing module 410 obtains (restored) raw audio data 415, which may correspond to or be an approximation of the audio data prior to audio enhancement at the UGC capture side.

[0066] Specifically, the processing module 410 may be configured to apply, to the audio data 405, at least one of ambiance restoring, loudness restoring, peak restoring, and timbre restoring.

[0067] To this end, the processing module 410 may comprise corresponding ones of a peak restore module (for peak restore), a loudness restore module 414 (for loudness restore), a noise management restore module 416 (for ambience restore), and a timbre management restore module (not shown; for timbre restore). Therein, the individual restore processes may "mirror" the audio enhancement applied at the UGC capture side. They may be applied in the reverse order compared to the processing at the UGC capture side (e.g., as performed by the apparatus 100 shown in Fig. 1). For example, the kind and / or order of the enhancement processing performed on the UGC capture side may be communicated with the metadata 435, with separate metadata, or may have been previously agreed on (e.g., in the context of standardization, etc.).

[0068] The peak restore aims to recover the over-suppressed peaks in the enhanced audio 405. The loudness restore seeks to bring the audio level back to the original level, and to remove distortions introduced by the loudness management. The noise management restore (ambience restore) brings back the sound events treated as noise (e.g., engine noise) and leave the decision of suppressing or keeping those events to later processing, or to a content creator using an editing tool. Therein, it is understood that noise management / noise suppression at the UGC capture side may suppress ambiance sound as noise, depending on the definition of "noise" and "ambiance". Restoring ambiance sound may be desirable especially in those cases in which the suppressed sound relates to a soundscape or the like.

[0069] As noted above, the restore processing is based on the one or more processing parameters indicated by the metadata 435 ( the first metadata 440). As further noted above, the one or more processing parameters may include band gains (e.g., band gains of previous noise management and / or band gains of previous timbre management) and / or full-band gains (e.g., full-band gains of previous loudness management and / or full-band gains of previous peak limiting) applied during the previous frame-wise audio enhancement. Having knowledge of these gains allows to reverse any enhancement processing that has been performed earlier based on these gains.

[0070] The rendering module 420 is configured for applying frame-wise audio enhancement to the (restored) raw audio data 415 to obtain enhanced audio data as the rendering output 425. The "rendering" group of processing components may be the same as those in the example apparatus 100 in Fig. 1 or the example apparatus 300 (example rendering system) in Fig. 3, including noise management, loudness management, timbre management, and peak limiting. Thus, the rendering module 420 may be configured to apply, to the (restored) raw audio data, at least one of noise management (e.g., by a noise management module 422), loudness management (e.g., by a loudness management module 424), timbre management (e.g., by a timbre management module 426), and peak limiting (e.g., by a peak limiting module 428).

[0071] The above processing is steered by the additional information available in the long-term statistics of the context metadata 435. In other words, the rendering module 420 is configured to apply the frame-wise audio enhancement to the raw audio data 415 based on the second metadata 445.

[0072] For example, the noise management may adjust noise suppression applied earlier to the enhanced audio 405, for example to avoid certain over-suppression, keep sound events, or further suppress certain types of noise in the enhanced audio, given the additional information available in the long-term statistics (e.g., indicated by the second metadata 445) of the context metadata 435. The loudness management may level the enhanced audio 405 (or rather, the raw audio 415) to a more appropriate range, given the additional information available in the long-term statistics of the context metadata 435. The timbre management may rebalance the timbre of the audio based on a content analysis, i.e., based on the long-term statistics of the context metadata.

[0073] The peak limiting may ensure that the amplitude of the audio after the aforementioned enhancements will not exceed the legitimate range allowed by audio playback.

[0074] Alternatively, the restored raw audio 415 obtained by the "restore" group of processing is exported to an editing tool, where some or all of the processing in the "rendering" group could be applied with controls by a content creator, for example via an editing tool UI, and where additional processing could be applied that is not part of "rendering" group. Accordingly, the editing module is a module for applying editing processing to the raw audio data to obtain edited audio data. Also the editing is based on the second metadata 445.

[0075] While an example apparatus 400 for UGC processing for rendering / editing has been described above, the present disclosure likewise relates to corresponding methods of UGC processing for rendering / editing. It is understood that any statements made above with regard to the apparatus 400 likewise apply to corresponding methods, and vice versa. An example of such method 500 of UGC processing (i.e., processing of audio data relating to UGC) is illustrated in the flowchart of Fig. 5. Method 500 comprises steps S510 through S540 and may be performed at a playback device (e.g., a mobile device or generic computing device) or editing device.

[0076] At step S510, the audio data is obtained. This may comprise or amount to receiving a bitstream comprising the audio data, including retrieving the audio data from a storage medium, for example.

[0077] At step S520, metadata for the audio data is obtained. The metadata comprise first metadata indicative of one or more processing parameters of a previous frame-wise audio enhancement of the audio data and second metadata generated based on the result of analyzing multiple frames of audio data, yielding long-term statistics of the audio data. Obtaining the metadata may comprise or amount to receiving a bitstream comprising the metadata (e.g., together with the audio data), including retrieving the metadata (e.g., together with the audio data) from a storage medium, for example.

[0078] At step S530, restore processing is applied to the audio data, using the one or more processing parameters, to at least partially reverse the previous frame-wise audio enhancement, thereby obtaining raw audio data. For example, applying the restore processing to the audio data may include applying at least one of ambiance restoring, loudness restoring, peak restoring, and timbre restoring.

[0079] Accordingly, the one or more processing parameters may include band gains (e.g., band gains of previous noise management and / or band gains of previous timbre management) and / or full-band gains (e.g., full-band gains of previous loudness management and / or full-band gains of previous peak limiting) applied during the previous frame-wise audio enhancement.

[0080] This step may proceed in accordance with the processing of the restore module 410 (and its sub-modules) described above.

[0081] At step S540, frame-wise audio enhancement is applied to the raw audio data to obtain enhanced audio data, and / or editing processing is applied to the raw audio data to obtain edited audio data.

[0082] Here, applying the frame-wise audio enhancement to the raw audio data is based on second metadata included in the metadata. As described above, the second metadata are indicative of long-term statistics of the audio data and may further comprise metadata indicative of one or more audio features of the audio data (e.g., a content type of the audio data, an indication of a capturing environment of the audio data, a signal-to-noise ratio of the audio data prior to the previous frame-wise audio enhancement, an overall loudness of the audio data prior to the previous frame-wise audio enhancement, and / or a spectral shape of the audio data prior to the previous frame-wise audio enhancement, etc.).

[0083] In analogy to the processing applied by step S220 of method 200 shown in Fig. 2, applying the frame-wise audio enhancement to the raw audio data may include applying at least one of noise management, loudness management, peak limiting, and timbre management.

[0084] Step S540 proceeds in accordance with the processing of the rendering module 420 (and its sub-modules) or the editing module described above.

[0085] Examples of methods and apparatus for UGC processing according to embodiments of the disclosure have been described above. It is understood that these methods and apparatus may be implemented by appropriate configuration of computing apparatus (e.g., devices, systems). A block diagram of an example of such computing device 600 is schematically illustrated in Fig. 6. The computing device 600 comprises a processor 610 and a memory 620 coupled to the processor 610. The memory 620 stores instructions for the processor 610. The processor 610 is configured to perform the steps of the methods and / or implement the modules of the apparatus described herein.

[0086] The present disclosure further relates to computer programs comprising instructions that, when executed by a computing device, cause the computing device (e.g., generic computing device 600) to perform the steps of the methods and / or implement the modules of the apparatus described herein.

[0087] The present disclosure further relates to computer-readable storage media storing such computer programs.Interpretation

[0088] Aspects of the systems described herein may be implemented in an appropriate computer-based sound processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0089] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), nonvolatile storage media in various forms, such as optical, magnetic or semiconductor storage media.

[0090] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits ("ASICs"). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, "content activity detectors" described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.

[0091] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.

[0092] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," or "having" and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms "mounted," "connected," "supported," and "coupled" and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.

Claims

1. A method of processing audio data relating to user generated content, the audio data captured by a capture device, the method comprising: obtaining the audio data; applying frame-wise audio enhancement to the audio data to obtain enhanced audio data; generating metadata for the enhanced audio data, based on one or more processing parameters of the frame-wise audio enhancement; and outputting the enhanced audio data together with the generated metadata for rendering at a playback device; wherein the metadata comprises first metadata generated based on the one or more processing parameters of the frame-wise audio enhancement and second metadata generated based on the result of analyzing multiple frames of the audio data; and wherein generating the metadata comprises compiling the first and second metadata to obtain compiled metadata as the metadata for output; wherein the frame-wise audio enhancement is applied during or immediately following capture of the audio data; and wherein the analysis of the multiple frames of the audio data yields long-term statistics of the audio data.

2. The method according to claim 1, wherein applying the frame-wise audio enhancement to the audio data includes applying at least one of: noise management; loudness management; peak limiting; and timbre management.

3. The method according to claim 1 or 2, wherein the one or more processing parameters include band gains and / or full-band gains applied during the frame-wise audio enhancement.

4. The method according to claim 3, wherein the one or more processing parameters include at least one of: band gains for noise management; full-band gains for loudness management; full-band gains for peak limiting; and band gains for timbre management.

5. The method according to any preceding claim, wherein the analysis of multiple frames of the audio data yields one or more audio features of the audio data, wherein the audio features of the audio data optionally relate to at least one of: a content type of the audio data; an indication of a capturing environment of the audio data; a signal-to-noise ratio of the audio data; an overall loudness of the audio data; and a spectral shape of the audio data.

6. A method of processing audio data relating to user generated content, the method comprising: obtaining the audio data; obtaining metadata for the audio data, wherein the metadata comprises first metadata indicative of one or more processing parameters of a previous frame-wise audio enhancement of the audio data, the frame-wise audio enhancement applied during or immediately following the capture of the audio data by a capture device, and second metadata indicative of long-term statistics of the audio data; applying restore processing to the audio data, using the one or more processing parameters, to at least partially reverse the previous frame-wise audio enhancement, thereby obtaining raw audio data; and applying frame-wise audio enhancement to the raw audio data to obtain enhanced audio data, or applying editing processing to the raw audio data to obtain edited audio data; wherein applying the frame-wise audio enhancement to the raw audio data is based on the second metadata, and wherein applying the editing processing is based on the second metadata.

7. The method according to claim 6, wherein applying the restore processing to the audio data includes applying at least one of: ambiance restoring; loudness restoring; peak restoring; and timbre restoring.

8. The method according to claim 6 or 7, wherein the one or more processing parameters include band gains and / or full-band gains applied during the previous frame-wise audio enhancement, wherein the one or more processing parameters optionally include at least one of: band gains of previous noise management; full-band gains of previous loudness management; full-band gains of previous peak limiting; and band gains of previous timbre management.

9. The method according to any one of claims 6 to 8, wherein the second metadata is indicative of one or more audio features of the audio data, wherein the audio features of the audio data optionally relate to at least one of: a content type of the audio data; an indication of a capturing environment of the audio data; a signal-to-noise ratio of the audio data prior to the previous frame-wise audio enhancement; an overall loudness of the audio data prior to the previous frame-wise audio enhancement; and a spectral shape of the audio data prior to the previous frame-wise audio enhancement.

10. The method according to any one of claims 6 to 9, wherein applying the frame-wise audio enhancement to the raw audio data includes applying at least one of: noise management; loudness management; peak limiting; and timbre management.

11. An apparatus for processing audio data relating to user generated content, the audio data captured by a capture device, the apparatus comprising: a processing module for applying frame-wise audio enhancement to audio data to obtain enhanced audio data, and for outputting the enhanced audio data, wherein the processing module is configured to apply the frame-wise audio enhancement during or immediately following capture of the audio data; and an analysis module for generating metadata for the enhanced audio data, based on one or more processing parameters of the frame-wise audio enhancement, and for outputting the metadata; wherein the analysis module is configured to generate the metadata further based on a result of analyzing multiple frames of the audio data, wherein the analysis of multiple frames of the audio data yields long-term statistics of the audio data; and wherein the analysis module is configured to generate first metadata based on the one or more processing parameters of the frame-wise audio enhancement and to generate second metadata based on the result of analyzing multiple frames of the audio data and to compile the first and second metadata, to thereby obtain compiled metadata as the metadata for output.

12. An apparatus for processing audio data relating to user generated content, the apparatus comprising: an input module for receiving audio data and metadata for the audio data, wherein the metadata comprises first metadata indicative of one or more processing parameters of a previous frame-wise audio enhancement of the audio data, the previous frame-wise audio enhancement applied during or immediately following the capture of the audio data by a capture device; the metadata further comprising second metadata indicative of long-term statistics of the audio data; a processing module for applying restore processing to the audio data, using the one or more processing parameters, to at least partially reverse the previous frame-wise audio enhancement, thereby obtaining raw audio data; and at least one of a rendering module and an editing module, wherein the rendering module is a module for applying frame-wise audio enhancement to the raw audio data to obtain enhanced audio data, and the editing module is a module for applying editing processing to the raw audio data to obtain edited audio data; wherein the rendering module is configured to apply the frame-wise audio enhancement and the editing processing to the raw audio data based on the second metadata.

13. An apparatus for processing audio data relating to user generated content, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the method according to any one of claims 1 to 10.

14. A computer program comprising instructions that, when executed by a computing device, cause the computing device to perform all steps of the method according to any one of claims 1 to 10.

15. A computer-readable storage medium storing the computer program according to claim 14.