Automatic Mixing of Audio Descriptions
An automated audio mixing method addresses the inefficiency of manual mixing by calculating loudness and generating mixing parameters, reducing the time needed to integrate audio descriptions with main audio, achieving faster production times.
Patent Information
- Application Number
- JP2022562366
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-13
- Filing Date
- 2021-04-12
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-04-12
AI Technical Summary
The time-consuming process of manually mixing audio descriptions with main audio files, requiring multiple man-hours per hour of content, especially when multiple formats and languages are involved, necessitates a more efficient automated solution.
A computer-implemented method for audio processing that includes receiving audio object and description data, calculating loudness levels, generating mixing parameters, and automatically mixing the audio data based on these parameters, with visualization for engineer adjustment.
Significantly reduces the time required for mixing, enabling an audio mix for a 90-minute movie in approximately 30 minutes compared to traditional methods.
Smart Images

Figure 0007714572000002 
Figure 0007714572000003 
Figure 0007714572000004
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 009,327, filed on Apr. 13, 2020, entitled “Automated Mixing of Audio Description Into Immersive Media”, which is hereby incorporated by reference in its entirety.
[0002] [Field] The present disclosure relates to audio processing, and more particularly, to audio mixing.
Background Art
[0003] Unless otherwise indicated herein, the techniques described in this section are not prior art to the claims of this application and are not admitted to be prior art by inclusion in this section.
[0004] Audio description generally refers to a verbal description of the visual components of audio - visual media such as movies. Audio description helps visually impaired people perceive audio - visual media. For example, audio description can verbally explain the visual aspects of a movie, such as the movement of characters and objects, facial expressions, etc. Audio description is distinguished from the main audio (also called default audio) that refers to the audio aspects of the audio - visual content itself (e.g., dialog, sound effects, background music, etc.).
[0005] Generally, an audio description is generated as a separate file, and an audio engineer mixes it with the main audio file to create a version of the audio that includes the audio description. The audio engineer performs the mixing to produce a consistent listening experience, for example, making the audio description audible in noisy scenes and preventing it from becoming too prominent in quiet scenes. Applying a gain that reduces the loudness level (e.g., a gain less than 1.0) can be referred to as ducking.
[0006] Next, a content provider (e.g., Netflix® service, Amazon Prime Video® service, Hulu® service, Apple TV+® service, etc.) can make available various audio file versions that a consumer can select. These versions can include the main audio file in various formats (such as stereo, 5.1-channel surround sound, etc.), various languages (e.g., English, Spanish, French, Japanese, Korean, etc.), versions with audio descriptions, etc. The content provider stores the audio file versions and provides the selected audio file to the consumer as, for example, the audio component of an audiovisual data stream (e.g., via the Hypertext Transfer Protocol (HTTP) Live Streaming (HLS) protocol).
[0007] As described above, the audio file version may have several formats including monaural, stereo, 5.1 channel surround sound, 7.1 channel surround sound, and the like. Other more recently developed audio formats include the ambisonics format (also called B format) for surround sound, the Dolby Atmos® format, and the like. Generally, the ambisonics format corresponds to a three-dimensional representation of sound pressure and sound pressure gradients in various dimensions. Generally, the Dolby Atmos® format corresponds to a collection of audio objects each including an audio track and metadata defining where the audio track should be output. Summary of the Invention
[0008] One problem with existing systems is the time required to perform mixing. Mixing generally requires an audio engineer to spend multiple man-hours per hour of content. For example, for a 90-minute movie, generating an audio mix including an audio description can involve 16 to 24 man-hours. Further, there may be multiple basic audio formats (e.g., stereo, 5.1 channel surround sound) and languages, and the time required doubles when generating an audio description mix for each combination of format and language. Embodiments are directed to automatically generating a mix including an audio description to reduce the time required by an audio engineer.
[0009] According to one embodiment, a computer-implemented method for audio processing includes receiving audio object data and audio description data, where the audio object data includes a first set of audio objects. The method further includes calculating a long-term loudness of the audio object data and a long-term loudness of the audio description data. The method further includes calculating a short-term loudness of the audio object data and a short-term loudness of the audio description data. The method further includes reading a first set of mixing parameters corresponding to the audio object data. The method further includes generating a second set of mixing parameters based on the first set of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the short-term loudness of the audio object data, and the short-term loudness of the audio description data. The method further includes generating a gain adjustment visualization corresponding to the second set of mixing parameters, the audio object data, and the audio description data. The method further includes generating mixed audio object data by mixing the audio object data and the audio description data according to the second set of mixing parameters. The mixed audio object data includes a second set of audio objects, and the second set of audio objects corresponds to the first set of audio objects mixed with the audio description data according to the second set of mixing parameters.
[0010] According to another embodiment, an apparatus includes a processor and a display. The processor is configured to control the apparatus to implement one or more of the methods described herein. The display is configured to display the gain adjustment visualization. The apparatus may additionally include details similar to one or more of the details of the methods described herein.
[0011] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls an apparatus to perform a process including one or more of the methods described herein.
[0012] The following detailed description and the accompanying drawings provide a further understanding of the nature and advantages of various implementations.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0014] Techniques related to audio processing will be described in this specification. In the following description, for the sake of explanation, numerous examples and specific details are set forth in order to provide a complete understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure as defined by the claims may include some or all of the features in these examples, either alone or in combination with other features described hereinafter, and may further include modifications and equivalents of the features and concepts described herein.
[0015] In the following description, various methods, processes, and procedures are detailed. Although specific steps may be described in a particular order, such order is mainly for convenience and clarity. Specific steps may be repeated more than once, may be performed before or after other steps (even if those steps are otherwise described in a different order), or may be performed in parallel with other steps. A second step is required to follow a first step only if the first step must be completed before the second step is initiated. Such situations are specifically pointed out when not apparent from the context.
[0016] In this specification, the terms "and", "or", and "and / or" are used. Such terms should be read as having an inclusive meaning. For example, "A and B" may mean at least the following: "both A and B", "at least both A and B". As another example, "A or B" may mean at least the following: "at least A", "at least B", "both A and B", "at least both A and B". As another example, "A and / or B" may mean at least the following: "A and B", "A or B". If an exclusive disjunction is intended, it is specifically indicated (e.g., "either A or B", "at most one of A and B").
[0017] This document describes various processing functions related to structures such as blocks, elements, components, circuits, etc. Generally, these structures can be implemented by a processor controlled by one or more computer programs.
[0018] Figure 1 is a block diagram of an audio mixing system 100. The audio mixing system 100 generally receives audio object data 102 and audio description data 104, performs mixing, and generates mixed audio object data 106. The audio mixing system 100 includes an audio object reader 110, an audio description (AD) reader 112, a loudness measurement component 114, a mixing component 116, a metadata reader 118, a visualization component 120, a metadata writer 122, an audio object writer 124, and an object metadata writer 126. The components of the audio mixing system 100 can be implemented by one or more computer programs executed by the processor of the audio mixing system 100. The audio mixing system 100 may include other components such as a display for displaying gain adjustment visualization and speakers for outputting audio that an audio engineer may use when mixing audio using the audio mixing system 100, but these components will not be described in detail.
[0019] The audio object reader 110 reads the audio file 130 and generates audio object data 132. Generally, the audio file 130 is one of several audio files and corresponds to the master version of the main audio of an audiovisual content file. A given audiovisual content file may have several audio files each corresponding to a combination of an audio format (e.g., monaural, stereo, 5.1 channel surround sound, 7.1 channel surround sound, audio object file, etc.) and an interactive language (e.g., English, Spanish, French, Japanese, Korean, etc.). Since the audio mixing system 100 is suitable for mixing audio object files, the audio file 130 is an audio object file for a given interactive language. For example, an audio engineer may select an English dialogue master object audio file as the audio file 130 for a given movie.
[0020] The audio object data 132 corresponds to the audio objects in the audio file 130. Generally, the audio object data 132 includes audio objects. The audio objects generally correspond to the audio file and location metadata. The location metadata instructs the rendering system on how to render the audio file at a given location and can include, for example, the size of the audio (pinpoint vs. diffuse), panning, etc. The rendering system then uses the metadata to perform the rendering and generate an appropriate output taking into account the specific loudspeaker arrangement in the rendering environment. The maximum number of audio objects within the audio object data 132 can vary depending on the specific implementation. For example, audio object data in the Dolby Atmos (registered trademark) format can have up to 128 objects.
[0021] The audio object data 132 may also include an audio bed, for example, as a subtype of the audio object or as a separate bed object. The audio bed generally corresponds to an audio file that is to be rendered at a defined bed position. Generally, each bed position corresponds to a channel output by an array of loudspeakers and is useful for dialog, ambient sound, etc. Typical bed positions include the center channel, the low-frequency effect channel, etc. The bed positions may correspond to surround sound channels such as 5.1-channel surround positions, 7.1-channel surround positions, 7.1-channel surround positions, etc.
[0022] The audio description reader 112 reads the audio description file 134 and generates audio description data 136. Generally, the audio description file 134 is one of several audio description files stored by the audio mixing system 100, and the audio engineer selects the audio description file that is desired to be mixed with the audio-visual content corresponding to the audio object file 130. The audio description of a given audio-visual content can be in multiple formats, for example, monoral, stereo, 5.1-channel surround sound, 7.1-channel surround sound, etc. As a result, there are many combinations of the audio object file 130 and the audio description file 134 that can be selected for a given audio-visual content. The audio description file 134 can be in various file formats such as the ".wav" file format. The audio description data 136 can have one of various coding formats including a pulse code modulation (PCM) signal, a linear PCM (LPCM) signal, an A-law PCM signal, etc.
[0023] The loudness measurement component 114 receives the audio object data 132 and the audio description data 136, calculates various loudness levels, and generates loudness data 138. Specifically, the loudness measurement component calculates the long-term loudness of the audio object data 132, the long-term loudness of the audio description data 136, several short-term loudness levels of the audio object data 132, and several short-term loudness levels of the audio description data 136. Generally, the period of the long-term loudness is a multiple of the period used for the short-term loudness. For example, audio data can be formatted as audio samples (e.g., at a sample rate such as 48 kHz, 96 kHz, 192 kHz, etc.), the short-term loudness can be calculated based on each sample, and the long-term loudness can be calculated over multiple samples. The multiple samples can be grouped into frames (e.g., with frame sizes such as 0.5 ms, 0.833 ms, 1.0 ms, etc.), and the short-term loudness can be calculated based on each frame. The long-term loudness can also be calculated over the entire audio data. Further details of the loudness measurement component 114 are provided with reference to FIG. 2.
[0024] The mixing component 116 receives the audio object data 132, the audio description data 136, and the loudness data 138, applies gains, and generally performs a mixing process, as further described herein. The mixing component 116 also receives metadata 140, generates visualization data 142, generates metadata 144, and generates mixed audio object data 146. The metadata reader 118, the visualization component 120, the metadata writer 122, and the audio object writer 124 can be considered functional components of the mixing component 116 that operate on the inputs to generate outputs.
[0025] The metadata reader 118 receives the metadata 140. Generally, the metadata 140 corresponds to an initial set of mixing parameters (also referred to as default mixing parameters). As further described herein, the initial mixing parameters provide a set of gain adjustments that an audio engineer can adjust as needed, and the adjusted mixing parameters may be referred to as adjusted mixing parameters. The metadata 140 can be in various formats, such as an Extensible Markup Language (XML) format, a JavaScript (registered trademark) Object Notation (JSON) format, and the like.
[0026] The visualization component 120 generates a gain adjustment visualization based on the mixing parameters and the loudness data 138. Generally, the gain adjustment visualization shows the loudness of the audio object data 132, the loudness of the audio description data 136, and the gain applied to the mixing according to the mixing parameters and the loudness data 138. The audio engineer can then use the gain adjustment visualization to evaluate the gain of the proposed mixing and, if necessary, adjust the gain to obtain adjusted mixing parameters. Further details of the gain adjustment visualization are provided with reference to FIG. 4.
[0027] The metadata writer 122 generates the metadata 144. The metadata 144 corresponds to the mixing parameters and the loudness data 138. If the default mixing parameters produce an acceptable audio mix, the parameters in the metadata 140 can be used as the parameters in the metadata 144 without any adjustment. However, generally, the audio engineer will adjust the mixing parameters from the default parameters to generate the metadata 144. The mixing parameters and the loudness data 138 represented by the metadata 144 are referred to as adjusted mixing parameters during the adjustment process and may be referred to as final mixing parameters when the audio engineer completes the gain adjustment.
[0028] The audio object writer 124 and the audio object metadata writer 148 cooperate to generate a mixed audio output. The audio object writer 124 mixes the audio object data 132 and the audio description data 136 according to the mixing parameters and the loudness data 138 to generate the mixed audio object data 146. The mixing parameters can be initial mixing parameters or adjusted mixing parameters. Next, the mixed audio object data 146 includes gain-adjusted audio object data and gain-adjusted audio description data. The mixed audio object data 146 can include audio objects, audio bed channels, and the like. The audio object data 132 and the audio description data 136 can be mixed according to two options (including their gain adjustment).
[0029] One option is to mix the audio description into one or more appropriate bed channels according to the format of the audio description. For example, a monoral audio description can be mixed into the central channel bed, a stereo audio description can be mixed into the left and right channel beds, a 5.1-channel audio description can be mixed into the 5.1-channel audio bed, and so on. This option is useful when the total number of available audio objects is limited.
[0030] Another option is to create one or more new audio objects corresponding to one or more appropriate positions corresponding to the format of the audio description. For example, in the case of a monoral audio description, an audio object can be generated at the central position, in the case of a stereo audio description, two audio objects can be generated at the left and right positions respectively, and in the case of a 5.1 channel audio description, five audio objects can be generated at the 5.1 channel surround positions, and so on.
[0031] The audio object metadata writer 126 generates audio object metadata 148 related to the mixed audio object data 146. For example, the audio object metadata 148 can include position information of each audio object, size information of each audio object, and the like.
[0032] The outline of the operation of the audio mixing system 100 is as follows. The audio engineer selects an audio object file and an audio description file, the audio object reader 110 generates the corresponding audio object data 132, and the audio description reader generates the corresponding audio description data 136. The loudness measurement component 114 generates loudness data 138. The mixing component 116 reads the metadata 140 and applies the mixing parameters to the loudness data 138 to generate gain visualization data 142. The audio mixing system 100 displays the gain visualization data 142, and the audio engineer evaluates the gain visualization.
[0033] Based on the gain visualization, an audio engineer can adjust the gain. The mixing component 116 adjusts the mixing parameters to correspond to the adjusted gain and displays the adjusted gain visualization corresponding to the adjusted mixing parameters. The display, evaluation, and adjustment processes may be executed multiple times, iteratively, etc.
[0034] When the evaluation of the mixing engineer is completed (based on either the initial mixing parameters or the adjusted mixing parameters), the mixing component 116 generates metadata 144 corresponding to the final mixing parameters and generates audio object data 146 mixed based on the final mixing parameters.
[0035] The mixing process executed using the audio mixing system 100 may result in the generation of mixed audio more quickly than existing mixing systems. For example, an audio mix using the initial mixing parameters for a 90 - minute movie can be generated in 30 minutes.
[0036] Further details of the audio mixing system 100 are as follows.
[0037] FIG. 2 is a block diagram of the loudness measurement component 200. The loudness measurement component 200 can be used as the loudness measurement component 114 (see FIG. 1). The loudness measurement component 200 generally receives audio object data 132 and audio description data 136, performs a loudness measurement, and generates loudness data 138 (see FIG. 1). The loudness measurement component 200 includes a spatial coding component 202, a renderer 204, and a loudness measurer 206.
[0038] The spatial coding component 202 receives the audio object data 132, performs spatial coding, and generates the clustered data 210. The audio object data 132 generally includes audio objects, and each audio object includes audio data and location metadata indicating where the audio data should be output. The audio object data 132 may also include an audio bed. The spatial coding component 202 performs spatial coding to reduce the number of objects and beds to a smaller number of clusters. For example, the audio object data 132 may include up to 128 objects and beds, and the spatial coding component groups these into elements (also called clusters). The clustered data 210 may include several clusters, such as 12 or 16 clusters. The clusters may be in a surround sound channel format, for example, an 11.1 channel format for 12 clusters or a 15.1 channel format for 16 clusters. Generally, the spatial coding component 202 performs spatial coding by dynamically grouping audio objects into dynamic clusters, where the audio objects can move from one cluster to another as the location information of the audio objects changes, and the clusters can move as well.
[0039] The renderer 204 receives the clustered data 210, performs rendering, and generates the rendered data 212. Generally, the renderer 204 performs rendering by associating the clusters within the clustered data 210 with channels within the rendered data 212. The rendered data 212 can be in one of various channel formats, including a monaural format (1 channel), a stereo format (2 channels), a 5.1 channel surround format (6 channels), a 7.1 channel surround format (8 channels), and the like. The rendered data 212 can have one of various coding formats, including a pulse code modulation (PCM) signal, a linear PCM (LPCM) signal, an A-law PCM signal, and the like. According to a particular exemplary embodiment, the rendered data 212 is a 5.1 channel LPCM signal.
[0040] The loudness measurer 206 receives the rendered data 212 and the audio description data 136, performs a loudness measurement, generates long-term loudness data 220 and short-term loudness data 222 for the rendered data 212, and generates long-term loudness data 224 and short-term loudness data 226 for the audio description data 136. Collectively, the loudness data 220, 222, 224, and 226 correspond to the loudness data 138 (see FIG. 1).
[0041] The loudness measurer 206 can implement one of several loudness measurement processes. Exemplary loudness measurement processes include a Leq (loudness equivalent continuous sound pressure level) process, an LKFS (loudness, K-weighted, relative to full scale) process, an LUFS (loudness units relative to full scale) process, and the like.
[0042] The loudness measurer 206 can be implemented by a computer program such as the Dolby (registered trademark) Professional Loudness Measurement (DPIM) development kit. The loudness measurer 206 calculates the long-term loudness data 220 and 224 to determine the overall level of this input. This value is used to perform (and obtain) the normalization of the input (referred to as dialog normalization or "dialnorm") so that there is no disturbing loudness difference between the rendered data 212 and the audio description data 136. An exemplary target value for dialog normalization is -31 dB.
[0043] Generally, the short-term loudness data 222 and 226 correspond to the time-ordered data, where each loudness measurement corresponds to the loudness of a specific part of the input (e.g., sample, frame, etc.). Generally, the long-term loudness data 220 and 224 correspond to the overall loudness of each respective input, but if they are calculated over multiple parts of the input, they can likewise be time-ordered data. The loudness data 138 can be formatted in a hierarchical format, for example, as Extensible Markup Language (XML) data.
[0044] FIG. 3 is a block diagram showing additional components of the mixing component 116 (see FIG. 1). These additional components are generally used when processing the loudness data 138 according to the mixing parameters. The additional components include a look-ahead component 302, a ramp component 304, and a maximum delta component 306. The mixing component 116 can include components for processing other mixing parameters as needed.
[0045] Mixing parameters are first provided in metadata 140 (see FIG. 1). For example, there may be several initial sets of mixing parameters available for selection by a mixing engineer corresponding to various genres of audio-visual content. The mixing engineer then selects a set of mixing parameters corresponding to the genre of the audio object file 134, and these initial parameters are provided to the mixing component 116 as metadata 140. Exemplary genres include action genre, horror genre, suspense genre, news genre, conversation genre, sports genre, and talk show genre.
[0046] An example of an initial set of mixing parameters included in metadata 140 for the action genre is shown in Table 1;
Table 1
[0047] The look-ahead component 302 processes look-ahead parameters. The look-ahead length of the main audio corresponds to the forward-looking period used by the mixing component 116 when processing loudness data 138 to duck the main audio when an audio description exists. If the audio description stops and starts again before the value of this parameter (e.g., 1.0 second), the lamp is not released during this stop period (an example is shown in FIG. 4 and explained in more detail there). This parameter prevents large fluctuations in the loudness of the main audio that would occur during short pauses in the audio description. This parameter may be different in other genres, for example, it may be increased in the news genre (e.g., 2.0 seconds).
[0048] The look-ahead length of the audio description corresponds to the forward-looking period used by the mixing component 116 when processing the loudness data 138 to adjust the value of the gain for the audio description. For example, the look-ahead component 302 may process the short-term loudness data 226 of the audio description (see FIG. 2) over the next period (e.g., 2.0 seconds) corresponding to the value of this parameter, and based on this processing, may increase or decrease the gain applied to the audio description. As another example, the look-ahead component 302 may process both the short-term loudness data 226 of the audio description and the short-term loudness data 222 of the main audio over the next time period (e.g., 2.0 seconds) corresponding to the value of this parameter, and based on this processing, may increase or decrease the gain applied to both the audio description and the main audio.
[0049] The ramp component 304 processes the ramp parameters. The ramp start offset corresponds to the length of time over which the gain for the main audio is gradually applied when the audio description starts. This gain is gradually applied and not instantaneously to reduce the possibility that the reduction of the main audio will interfere with the listener experience. For example, when the gain applied to the main audio when mixing the audio description is 0.3, instead of instantaneously changing the gain from 1.0 to 0.3, the gain is gradually changed over the ramp start offset period. In the action genre, a period of 0.192 seconds functions well. For other genres, this period may be adjusted. For example, in the drama genre, a longer period (e.g., 0.384 seconds) functions well.
[0050] The lamp-off offset corresponds to the length of time during which the gain to the main audio is gradually released when the audio description ends. For example, when a gain of 0.3 is applied during the audio description, the gain gradually returns to 1.0 over the lamp-off offset period (e.g., 0.192 seconds). The lamp-off offset period may be different from or the same as the lamp-on offset period. In the action genre, a period of 0.192 seconds functions well. For other genres, this period may be adjusted. For example, in the drama genre, a longer period (e.g., 0.384 seconds) functions well.
[0051] The maximum delta component 306 processes the target maximum delta and the minimum gain parameter. The target maximum delta corresponds to the difference in loudness level between the main audio and the audio description to which the gain is applied to the main audio. If the loudness difference is less than this level, no gain will be applied to the main audio even if an audio description exists. This feature is useful when there is background music in a quiet scene and an audio description exists. When the main audio is ducked, the background music cannot be heard in the audio description, and there is a risk that the director's intention for the audio scene will not be conveyed.
[0052] The minimum gain corresponds to the minimum gain applied to the audio description. This value prevents the audio description from becoming too loud compared to the main audio. For example, in a quiet scene, the audio description may become too loud and prevent the listener from immersing in the audio scene. In these extreme cases, by ducking the audio description, the listener can immerse in the audio scene.
[0053] As described above, the parameters in Table 1 are an initial set of parameters provided to the mixing component 116 via the metadata 140. These correspond to genres, and the audio mixing system 100 may store several sets of mixing parameters, each corresponding to one of several genres. Additionally, the values of the parameters used as the initial parameters may be adjusted as well. For example, in the case of the action genre (see Table 1), the lamp start offset value may be changed from -0.192 to -0.182. Then, the value of -0.182 is used as one of the initial parameters provided via the metadata 140. This allows the mixing engineer to adjust the default mixing parameters before the default mixing parameters are input to the mixing component 116. Further, there may be multiple sets of parameters for a given genre. For example, in the case of the action genre, one set of parameters may have a lamp start offset value of -0.190, and another set of parameters may have a lamp start offset value of -0.195.
[0054] The audio mixing system 100 may process mixing parameters other than those detailed in Table 1.1. For example, the default main audio ducking parameter may set the default gain value applied when ducking the main audio. This parameter may be defined as a gain level (e.g., a gain of 0.3), a decibel level (e.g., -16 dB), etc. As another example, enabling a minimum gain parameter for ducking the audio description (as described above) is an artistic choice that can be toggled on or off according to the parameter.
[0055] FIG. 4 is a graph 400 showing a visualization 402 of visualization data 142 (see FIG. 1). In graph 400, the x-axis is the sample index of main audio data (e.g., audio object data 132) and audio description data (e.g., audio description data 136). The x-axis can be regarded as a time index, with the start of the content being 0 on the left side and the end of the content being on the right side. The left y-axis indicates the gain applied to the main audio and the audio description, and the right side indicates the loudness level (in dB) of the main audio and the audio description.
[0056] The visualization 402 is an example of a selected audio object file (e.g., 130) and a selected audio description file (e.g., 134) showing gain and loudness. The gain is indicated by a dashed line, where line 410 indicates the gain applied to the main audio and line 412 indicates the gain applied to the audio description. As described above, these gains correspond to the mixing parameters applied to the loudness data 138 (see FIG. 1). The loudness levels are lines 414 and 416, where line 414 indicates the loudness of the main audio and line 416 indicates the loudness of the audio description. Note that line 416 is discontinuous. Where line 416 does not exist, the audio description does not exist.
[0057] The visualization 402 shows several features. Note that the gain (line 412) applied to the audio description is constant at 1.0. This indicates that, as a result of considering the mixing parameters and the loudness data 138, the mixing component 116 has determined that there is no need to apply gain adjustment to the audio description. For example, the overall loudness comparison between the main audio and the audio description may be within the range of values defined by the mixing parameters.
[0058] Note that the gain applied to the main audio (line 410) is mainly in the range between 1.0 and 0.3, except near point 420. The ramp-down from 1.0 to 0.3 and the ramp-up from 0.3 to 1.0 are not easily visible in relation to the scale of the x-axis, but they exist according to the mixing parameters of the ramp start offset and the ramp end offset (see Table 1). Further, note that the gain of 0.3 can be configured using the mixing parameters, for example, using the default main audio ducking parameters. Near point 420, the applied gain is approximately 0.32. This is the result of the interaction between the mixing parameters applied to the short-term loudness. For example, the default parameters may result in this gain, or the mixing engineer may adjust the mixing parameters (e.g., in response to listening to the mixed audio to produce a more acceptable mix) so that this gain is achieved.
[0059] Note that the gain applied to the main audio (line 410) generally exists when there is an audio description (line 416). However, line 410 also exists at some indices where there is no audio description, such as near point 422. This indicates a short break in the audio description that is smaller than the look-ahead length value defined in the mixing parameters (see Table 1).
[0060] The mixing engineer may use visualization 402 to evaluate the proposed gain applied in mixing. For example, the default parameters result in a first visualization, which the mixing engineer evaluates. If the first visualization appears to show that an acceptable mix is obtained, the mixing engineer may instruct the audio mixing system 100 to generate an audio mix without any adjustments. However, if the first visualization shows some discontinuities or other visual features indicating that an unacceptable mix will be obtained, the mixing engineer may adjust the mixing parameters, and the audio mixing system 100 may generate a second visualization based on the adjusted parameters (e.g., the mixing parameters may be adjusted so that line 410 has a slightly different appearance near point 420). The process of displaying the modified visualization, evaluating the modified visualization, and adjusting the mixing parameters may be repeatedly (or otherwise multiple times) executed until the visual features of the modified visualization indicate that an acceptable mix will be obtained.
[0061] In addition, it should be recalled that the metadata 144 corresponding to the final mixing parameters is generated before and after the time when the mixed audio object data 146 is generated. Thereby, the mixing engineer can evaluate the mixed audio. If unacceptable, the mixing engineer may instruct the audio mixing system 100 to use the metadata 144 as an input to the mixing component 116 (e.g., as the mixing parameter 140), and then perform evaluation and adjustment based on the adjusted parameters instead of the default parameters.
[0062] FIG. 5 is a block diagram of an audio mixing system 500. Compared with the audio mixing system 100 (see FIG. 1), which is described as processing object audio, the audio mixing system 500 can be used to process other types of audio. The audio mixing system 500 includes a converter 502, an audio mixing system 100, and a converter 504.
[0063] Converter 502 receives audio data 510, converts the audio data 510, and generates audio object data (e.g., audio object file 130, audio object data 132, etc.). The audio data 510 generally corresponds to audio data that does not include audio objects, and the converter 502 performs a conversion that converts the audio data 510 into object audio data. For example, the audio data 510 may be in an ambisonics format, the audio object file 130 may be in a Dolby Atmos (registered trademark) format, and the converter 502 may implement a conversion from ambisonics to Dolby Atmos (registered trademark). The audio data 510 may generally correspond to the main audio of audio-visual content (e.g., a movie soundtrack).
[0064] As described above, the audio mixing system 100 processes the audio object file 130 (resulting from the conversion) and generates mixed audio object data 146 and audio object metadata 148.
[0065] Converter 504 receives the mixed audio object data 146, converts the mixed audio object data 146, and generates mixed audio data 512. Converter 504 may also receive mixed audio object data 148. The mixed audio data 512 corresponds to the audio data 510 and the mixed audio description. Generally, converter 504 performs the reverse of the conversion performed by converter 502. For example, when converter 502 implements a conversion from ambisonics to Dolby Atmos®, converter 504 implements a conversion from Dolby Atmos® to ambisonics.
[0066] In this way, the audio mixing system 500 enables the audio mixing system 100 to be used with other types of audio.
[0067] Figure 6 is a device architecture 600 for implementing the features and processes described herein, according to one embodiment. Architecture 600 can be implemented in any electronic device, including but not limited to desktop computers, home audio / video (AV) devices, radio broadcast devices, mobile devices (e.g., smartphones, tablet computers, laptop computers, wearable devices), etc. In the exemplary embodiment shown, architecture 600 is for a laptop computer and includes one or more processors 601, a peripheral interface 602, an audio subsystem 603, loudspeakers 604, a microphone 605, sensors 606 (e.g., accelerometers, gyroscopes, barometers, magnetometers, cameras), a location processor 607 (e.g., GNSS receiver), a wireless communication subsystem 608 (e.g., Wi-Fi, Bluetooth®, cellular), and an I / O subsystem(s) 609 that includes a touch controller 610 and other input controllers 611, a touch surface 612, and other input / control devices 613. Other architectures with more or fewer components can also be used to implement the disclosed embodiments.
[0068] A memory interface 614 is coupled to the processor 601, the peripheral interface 602, and a memory 615 (e.g., flash, RAM, ROM). The memory 615 stores computer program instructions and data, including but not limited to operating system instructions 616, communication instructions 617, GUI instructions 618, sensor processing instructions 619, telephone instructions 620, electronic messaging instructions 621, web browsing instructions 622, audio processing instructions 623, GNSS / navigation instructions 624, and applications / data 625. The audio processing instructions 623 include instructions for performing the audio processing described herein.
[0069] As a specific example, the device architecture 600 may implement an audio mixing system 100 (see FIG. 1) by, for example, executing one or more computer programs. The device architecture may access an audio file 130 via a peripheral device interface 602 (connected to non-volatile storage such as a solid state drive, for example), may calculate loudness data 138 using a processor 601, may display visualization data 142 via a peripheral device interface 602 (connected to a display device, for example), and may generate mixed audio object data 146 using a processor 601.
[0070] FIG. 7 is a flowchart of a method 700 for audio processing. Method 700 may be executed by a device (such as a laptop computer, a desktop computer, etc.) having the components of the architecture of FIG. 6 to implement functions such as an audio mixing system 100 (see FIG. 1) by, for example, executing one or more computer programs.
[0071] At 702, audio object data and audio description data are received. The audio object data includes a first set of audio objects. For example, a loudness measurement component 114 and a mixing component 116 (see FIG. 1) may receive audio object data 132 and audio description data 136.
[0072] At 704, the long-term loudness of the audio object data and the long-term loudness of the audio description data are calculated. For example, a loudness measurement component 114 (see FIG. 1) may calculate the long-term loudness as part of the loudness data 138. The long-term loudness may be calculated over the entire data.
[0073] At 706, some short-term loudness of the audio object data and some short-term loudness of the audio description data are calculated. For example, the loudness measurement component 114 (see FIG. 1) may calculate the short-term loudness as part of the loudness data 138. The short-term loudness may be calculated on a continuous basis, such as for each sample, for each frame, etc.
[0074] At 708, a first set of mixing parameters corresponding to the audio object data is read. For example, the metadata reader 118 (see FIG. 1) may read the metadata 140 including the initial mixing parameters.
[0075] At 710, a second set of mixing parameters is generated based on the first set of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the short-term loudness of the audio object data, and the short-term loudness of the audio description data. For example, the mixing component 116 (see FIG. 1) may process the loudness data 138 according to the initial mixing parameters to generate a set of proposed gains for the main audio and the audio description (corresponding to lines 410 and 412 in FIG. 4, for example).
[0076] At 712, a gain adjustment visualization is generated. The gain adjustment visualization corresponds to the second set of mixing parameters, the audio object data, and the audio description data. For example, the gain adjustment visualization may correspond to the visualization 402 (see FIG. 4) showing the loudness of the main audio, the loudness of the audio description, and the proposed gain.
[0077] At 714, the gain adjustment visualization is evaluated to determine whether applying the mixing parameters produces an acceptable result. For example, a mixing engineer may evaluate the visualization 402 (see FIG. 4). If the result is not acceptable, the flow proceeds to 716; if the result is acceptable, the flow proceeds to 718.
[0078] At 716, a second set of mixing parameters is adjusted. For example, a mixing engineer may adjust the proposed gain, and the mixing component 116 may accordingly adjust the mixing parameters to correspond to the adjusted gain.
[0079] At 718, the mixed audio object data is generated by mixing the audio object data and the audio description data according to the second set of mixing parameters. The mixed audio object data includes a second set of audio objects, where the second set of audio objects corresponds to the first set of audio objects mixed with the audio description data according to the second set of mixing parameters. For example, the audio object writer 124 (see FIG. 1) may generate the mixed audio object data 146. Audio object metadata related to the mixed audio object data may also be generated. For example, the audio object metadata writer 126 may generate the audio object metadata 148.
[0080] Method 700 may include additional steps to accommodate other functions such as the audio mixing system 100 as described herein. For example, default mixing parameters may be selected based on the genre of the audio visual content being mixed. The mixing parameters may include look-ahead parameters, ramp parameters, maximum delta parameters, and the like. Method 700 may include a conversion step of converting non-object audio to object audio for processing by the audio mixing system 100, and a conversion step of converting the mixed object audio to mixed non-object audio.
[0081] Additional Details
[0082] The description has focused on the mixing of audio descriptions, but embodiments may also be used to mix other types of audio content to achieve similar improvements in time, effort, and efficiency. For example, the audio mixing system 100 (see FIG. 1) may be used to mix a director's commentary.
[0083] Details of Implementations
[0084] Embodiments may be implemented in hardware, in executable modules stored on a computer-readable medium, or in a combination of both (such as a programmable logic array). Unless otherwise specified, the steps performed by the embodiments need not be inherently related to any particular computer or other device, although they may be in certain embodiments. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized device (e.g., an integrated circuit) to perform the required method steps. Accordingly, embodiments may be implemented in one or more computer programs executed on one or more programmable computer systems each comprising at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0085] Each such computer program is preferably stored on or downloaded to a storage medium or device readable by a general-purpose or special-purpose programmable computer (e.g., solid-state memory or media, or magnetic or optical media) for configuring and operating a computer when the storage medium or device is read by the computer system to perform the procedures described herein. The system of the present invention may also be considered to be implemented as a computer-readable storage medium configured with a computer program, where the storage medium so configured causes a computer system to operate in a particular predefined manner to perform the functions described herein (software itself and intangible or transient signals are excluded insofar as they are non-patentable subject matter).
[0086] Aspects of the systems described herein can be implemented in a suitable computer-based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system can include one or more networks comprising any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on various different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0087] One or more of the components, blocks, processes, or other functional components can be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein can be described in terms of any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media, with respect to their behavior, register transfers, logic components, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions can be embodied include various forms of physical (non-transitory) non-volatile storage media including, but not limited to, optical storage media, magnetic storage media, or semiconductor storage media.
[0088] The above description shows various embodiments of the present disclosure, along with examples of how aspects of the present disclosure can be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of the present disclosure as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and can be used without departing from the spirit and scope of the present disclosure as defined by the claims.
Claims
1. A computer-implemented method for audio processing, comprising: Receiving audio object data and audio description data, wherein the audio object data includes a first plurality of audio objects; Calculating a long-term loudness of the audio object data and a long-term loudness of the audio description data; Calculating a plurality of short-term loudnesses of the audio object data and a plurality of short-term loudnesses of the audio description data; Reading a first plurality of mixing parameters corresponding to the audio object data; Generating a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data; Generating a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data; Generating mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data includes a second plurality of audio objects, and the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters. A method.
2. The long-term loudness of the audio object data is calculated over a plurality of samples of the audio object data, and the long-term loudness of the audio description data is calculated over a plurality of samples of the audio description data. Each of the plurality of short-term loudnesses of the audio object data is calculated over a single sample of the audio object data, and each of the plurality of short-term loudnesses of the audio description data is calculated over a single sample of the audio description data. The method according to claim 1.
3. The method according to claim 1 or 2, wherein the first plurality of mixing parameters are associated with one of a plurality of genres, and each of the plurality of genres is associated with a corresponding set of mixing parameters.
4. The method according to claim 3, wherein the plurality of genres include an action genre, a horror genre, a suspense genre, a news genre, a conversation genre, a sports genre, and a talk show genre.
5. The method according to any one of claims 1 to 4, wherein the first plurality of mixing parameters include a look-ahead parameter, a ramp parameter, and a maximum delta parameter.
6. The method according to claim 5, wherein the look-ahead parameter corresponds to maintaining a uniform gain adjustment during an audio pause of the audio description data.
7. The method according to claim 5 or 6, wherein the ramp parameter corresponds to a period during which a gain adjustment is gradually applied.
8. The method according to any one of claims 5 to 7, wherein the maximum delta parameter corresponds to a maximum loudness difference between a frame of the audio object data and a corresponding frame of the audio description data.
9. Receiving user input for adjusting the second plurality of mixing parameters before generating the mixed audio object data; Generating a modified gain adjustment visualization corresponding to the second plurality of mixing parameters adjusted according to the user input; further comprising The mixed audio object data is generated based on the adjusted second plurality of mixing parameters. The method according to any one of claims 1 to 8.
10. Before receiving the audio object data, Receiving audio data, where the audio data does not include audio objects. converting the audio data into the audio object data; after generating the mixed audio object data, converting the mixed audio object data into mixed audio data, further comprising, wherein the mixed audio data corresponds to the audio data mixed with the audio description data, The method according to any one of claims 1 to 9.
11. A non-transitory computer-readable medium storing a computer program that controls an apparatus to execute a process including the method according to any one of claims 1 to 10 when executed by a processor.
12. An apparatus for audio processing, comprising: a processor, wherein the processor is configured to control the apparatus to receive audio object data and audio description data, wherein the audio object data includes a first plurality of audio objects; wherein the processor is configured to control the apparatus to calculate a long-term loudness of the audio object data and a long-term loudness of the audio description data; wherein the processor is configured to control the apparatus to calculate a plurality of short-term loudnesses of the audio object data and a plurality of short-term loudnesses of the audio description data; wherein the processor is configured to control the apparatus to read a first plurality of mixing parameters corresponding to the audio object data; wherein the processor is configured to control the apparatus to generate a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data. The processor is configured to control the apparatus to generate gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data. The processor is configured to control the apparatus to generate mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data includes a second plurality of audio objects, and the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters. Apparatus. **Claim 13** A display configured to display the gain adjustment visualization The apparatus according to claim 12, further comprising. **Claim 14** The long-term loudness of the audio object data is calculated over a plurality of samples of the audio object data, and the long-term loudness of the audio description data is calculated over a plurality of samples of the audio description data. Each of the plurality of short-term loudnesses of the audio object data is calculated over a single sample of the audio object data, and each of the plurality of short-term loudnesses of the audio description data is calculated over a single sample of the audio description data. The apparatus according to claim 12 or 13. **Claim 15** The apparatus according to any one of claims 12 to 14, wherein the first plurality of mixing parameters are associated with one of a plurality of genres, and each of the plurality of genres is associated with a corresponding set of mixing parameters. **Claim 16** The apparatus according to any one of claims 12 to 15, wherein the first plurality of mixing parameters include a look-ahead parameter, a ramp parameter, and a maximum delta parameter. **Claim 17** The apparatus according to claim 16, wherein the look-ahead parameter corresponds to maintaining a uniform gain adjustment during an audio pause of the audio description data. **Claim 18** The apparatus according to claim 16 or 17, wherein the lamp parameter corresponds to a period during which the gain adjustment is gradually applied. **Claim 19** The apparatus according to any one of claims 16 to 18, wherein the maximum delta parameter corresponds to a maximum loudness difference between a frame of the audio object data and a corresponding frame of the audio description data. **Claim 20** The processor is configured to control the apparatus to receive a user input for adjusting the second plurality of mixing parameters before generating the mixed audio object data, The processor is configured to control the apparatus to generate a modified gain adjustment visualization corresponding to the second plurality of mixing parameters adjusted according to the user input, The mixed audio object data is generated based on the adjusted second plurality of mixing parameters. The apparatus according to any one of claims 12 to 19.
Citation Information
Patent Citations
Object-based audio signal balancing method
JP2019501563A
Dynamic audio ducking
US20100211199A1