Automatic mixing of audio descriptions

By using a computer-automated audio mixing method to calculate loudness and generate mixing parameters, the problem of long audio mixing time in existing technologies is solved, and efficient generation of audio mixes containing audio descriptions is achieved.

CN115668765BActive Publication Date: 2026-05-12DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2021-04-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing audio mixing systems require a significant amount of manual time to generate audio mixes that include audio descriptions, especially for combinations of different formats and languages, resulting in inefficiency.

Method used

The computer-implemented audio processing method receives audio object data and audio description data, calculates loudness and generates mixing parameters, automatically mixes the audio object and audio description data, provides gain adjustment visualization to assist engineers in making adjustments, and finally generates the mixed audio object data.

Benefits of technology

It significantly reduces the working time of audio engineers, enabling the generation of high-quality audio mixes with audio descriptions in a short time, thus improving mixing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668765B_ABST
    Figure CN115668765B_ABST
Patent Text Reader

Abstract

A computer-implemented audio processing method, the method comprising: receiving audio object data and audio description data, wherein the audio object data comprises a first plurality of audio objects; calculating a long-term loudness of the audio object data and a long-term loudness of the audio description data; calculating a plurality of short-term loudnesses of the audio object data and a plurality of short-term loudnesses of the audio description data; reading a first plurality of mixing parameters corresponding to the audio object data; generating a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data; generating a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data; and generating mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data comprises a second plurality of audio objects, wherein the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is related to U.S. Provisional Application No. 63 / 009,327, filed April 13, 2020, entitled “Automatic mixing of audio descriptions into immersive media,” which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to audio processing, and more particularly to audio mixing. Background Technology

[0004] Unless otherwise indicated herein, the methods described in this section are not prior art to the claims of this application and are not acknowledged as prior art by virtue of their inclusion in this section.

[0005] Audio descriptions typically refer to verbal descriptions of the visual elements of audiovisual media, such as films. Audio descriptions help visually impaired consumers perceive audiovisual media. For example, an audio description can verbally describe the visual aspects of a film, such as the movement of characters and objects, facial expressions, etc. Audio descriptions differ from content known as the main audio (also called the default audio), which refers to the audio aspects of the audiovisual content itself (e.g., dialogue, sound effects, background music, etc.).

[0006] Typically, the audio description is generated as a separate file. Audio engineers mix this separate file with the main audio file to create an audio version that now includes the audio description. Audio engineers perform this mixing to create a harmonious listening experience, such as making the audio description audible in noisy scenes and not too loud in quiet scenes. Applying a gain that reduces the loudness level (e.g., a gain less than 1.0) can be referred to as ducking.

[0007] Content providers (e.g., Netflix) TM Services, Amazon Prime Video TM Services, Hulu TM Services, Apple TV+ TM The content provider (such as the service provider) can then offer consumers a variety of audio file versions to choose from. These versions may include main audio files in various formats (stereo, 5.1 channel surround sound, etc.), in various languages ​​(e.g., English, Spanish, French, Japanese, Korean, etc.), versions with audio descriptions, etc. The content provider stores the audio file versions and provides selected audio files to consumers, for example, as audio components of an audiovisual data stream (e.g., via the Hypertext Transfer Protocol (HTTP) Real-Time Streaming (HLS) protocol).

[0008] As mentioned above, audio file versions can come in various formats, including mono, stereo, 5.1 channel surround sound, and 7.1 channel surround sound. Other more recently developed audio formats include Ambisonics (also known as B format) high-fidelity surround sound and Dolby Atmos. TM Formats, etc. Typically, high-fidelity stereo formats correspond to a three-dimensional representation of sound pressure levels and gradients in various dimensions. Dolby Atmos is a common example. TM The format corresponds to a collection of audio objects, each of which includes an audio track and metadata defining where the audio track should be output. Summary of the Invention

[0009] One problem with existing systems is the time required to perform mixing. Mixing typically requires audio engineers to spend multiple hours on a single hour of content. For example, a 90-minute movie might involve 16 to 24 hours to generate an audio mix containing audio descriptions. Furthermore, multiple basic audio formats (e.g., stereo, 5.1 channel surround sound) and multiple languages ​​can exist; generating audio description mixes for each combination of format and language multiplies the required time. Implementations involve automating the generation of mixes containing audio descriptions to reduce the time required by audio engineers.

[0010] According to an embodiment, a computer-implemented audio processing method includes receiving audio object data and audio description data, wherein the audio object data includes a first set of audio objects. The method further includes calculating the long-term loudness of the audio object data and the long-term loudness of the audio description data. The method further includes calculating the short-term loudness of the audio object data and the short-term loudness of the audio description data. The method further includes reading a first set of mixing parameters corresponding to the audio object data. The method further includes generating a second set of mixing parameters based on the first set of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the short-term loudness of the audio object data, and the short-term loudness of the audio description data. The method further includes generating a gain adjustment visualization corresponding to the second set of mixing parameters, the audio object data, and the audio description data. The method further includes generating mixed audio object data by mixing the audio object data and the audio description data according to the second set of mixing parameters. The mixed audio object data includes a second set of audio objects, and the second set of audio objects corresponds to the first set of audio objects mixed with the audio description data according to the second set of mixing parameters.

[0011] According to another embodiment, an apparatus includes a processor and a display. The processor is configured to control the apparatus to implement one or more of the methods described herein. The display is configured to display a gain adjustment visualization. The apparatus may additionally include details similar to those of the methods described herein.

[0012] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls means to perform processing including one or more of the methods described herein.

[0013] The following detailed description and accompanying drawings provide a further understanding of the nature and advantages of the various embodiments. Attached Figure Description

[0014] Figure 1 This is a block diagram of the audio mixing system 100.

[0015] Figure 2 This is a block diagram of the loudness measuring component 200.

[0016] Figure 3 The hybrid component 116 is shown (see Figure 1 A block diagram of the additional components.

[0017] Figure 4 It shows the visualization data 142 (see Figure 1 The visualization of 402 is shown in the curve 400.

[0018] Figure 5 This is a block diagram of the audio mixing system 500.

[0019] Figure 6 This is an apparatus architecture 600 according to embodiments for implementing the features and processes described herein.

[0020] Figure 7 This is a flowchart of audio processing method 700. Detailed Implementation

[0021] This document describes techniques related to audio processing. In the following description, numerous examples and specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure, as defined by the claims, may include some or all of these examples, either alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.

[0022] The following description details various methods, processes, and procedures. While specific steps may be described in a particular order, this order is primarily for convenience and clarity. A particular step may be performed more than once, may occur before or after other steps (even if these steps are described in a different order), and may occur in parallel with other steps. A second step is only necessary if it must be completed before the first step can begin. This will be specifically indicated when it is unclear from the context.

[0023] In this document, the terms “and,” “or,” and “and / or” are used. These terms should be interpreted as inclusive. For example, “A and B” can at least mean: “both A and B,” or “at least both A and B.” As another example, “A or B” can at least mean: “at least A,” “at least B,” “both A and B,” or “at least both A and B.” As yet another example, “A and / or B” can at least mean: “A and B,” or “A or B.” When XOR is intended to be used, it will be specifically indicated (e.g., “either A or B,” or “at most one of A and B”).

[0024] This document describes the various processing functions associated with structures such as blocks, elements, components, and circuits. Typically, these structures can be implemented by a processor controlled by one or more computer programs.

[0025] Figure 1 This is a block diagram of an audio mixing system 100. The audio mixing system 100 typically receives audio object data 102 and audio description data 104, performs mixing, and generates mixed audio object data 106. The audio mixing system 100 includes an audio object reader 110, an audio description (AD) reader 112, a loudness measurement unit 114, a mixing unit 116, a metadata reader 118, a visualization unit 120, a metadata writer 122, an audio object writer 124, and an object metadata writer 126. The components of the audio mixing system 100 can be implemented by one or more computer programs executed by the processor of the audio mixing system 100. The audio mixing system 100 may include other components that audio engineers can use when mixing audio using the audio mixing system 100, such as a display for visualizing gain adjustments, speakers for outputting audio, etc.; these components are not discussed in detail.

[0026] Audio object reader 110 reads audio file 130 and generates audio object data 132. Typically, audio file 130 is one of multiple audio files and corresponds to the main version of the main audio of an audiovisual content file. A given audiovisual content file can have multiple audio files, each corresponding to a combination of audio format (e.g., mono, stereo, 5.1 channel surround sound, 7.1 channel surround sound, audio object file, etc.) and dialogue language (e.g., English, Spanish, French, Japanese, Korean, etc.). Audio mixing system 100 is adapted for mixing audio object files, so audio file 130 is an audio object file for a given dialogue language. For example, an audio engineer could select the main English dialogue object audio file of a given movie as audio file 130.

[0027] Audio object data 132 corresponds to audio objects in audio file 130. Typically, audio object data 132 includes audio objects. Audio objects typically correspond to an audio file and location metadata. Location metadata instructs the rendering system how to render the audio file at a given location; this can include audio size (precise positioning and diffusion), translation, etc. The rendering system then uses the metadata to perform rendering to generate appropriate output given a specific loudspeaker arrangement in the rendering environment. The maximum number of audio objects in audio object data 132 may vary depending on the specific implementation. For example, Dolby Atmos... TM The formatted audio object data can have a maximum of 128 objects.

[0028] Audio object data 132 may also include audio beds, for example, as a subtype of the audio object or as a separate bed object. An audio bed typically corresponds to an audio file to be rendered at a defined bed position. Typically, each bed position corresponds to a channel output by the amplifier array and is useful for dialogue, ambient sounds, etc. Typical bed positions include the center channel, low-frequency effects channel, etc. Bed positions can correspond to surround sound channels, such as 5.1 channel surround positions, 7.1 channel surround positions, 7.1.4 channel surround positions, etc.

[0029] Audio description reader 112 reads audio description file 134 and generates audio description data 136. Typically, audio description file 134 is one of multiple audio description files stored in audio mixing system 100, and the audio engineer selects the audio description file they wish to mix with the audiovisual content corresponding to audio object file 130. The audio description of a given audiovisual content can be in various formats, such as mono, stereo, 5.1 channel surround sound, 7.1 channel surround sound, etc. Therefore, there are many possible combinations of audio object file 130 and audio description file 134 for a given audiovisual content. Audio description file 134 can be in various file formats, such as the ".wav" file format. Audio description data 136 can have one of various encoding formats, including pulse code modulation (PCM) signals, linear PCM (LPCM) signals, A-law PCM signals, etc.

[0030] Loudness measurement unit 114 receives audio object data 132 and audio description data 136, calculates various loudnesses, and generates loudness data 138. Specifically, the loudness measurement unit calculates the long-term loudness of the audio object data 132, the long-term loudness of the audio description data 136, multiple short-term loudnesses of the audio object data 132, and multiple short-term loudnesses of the audio description data 136. Typically, the time period for the long-term loudness is a multiple of the time period used for the short-term loudness. For example, the audio data can be formatted as audio samples (e.g., sampling rates of 48kHz, 96kHz, 192kHz); the short-term loudness can be calculated based on each sample, and the long-term loudness can be calculated using multiple samples. Multiple samples can be organized into frames (e.g., frame sizes of 0.5ms, 0.833ms, 1.0ms, etc.), and the short-term loudness can be calculated based on each frame. The long-term loudness can also be calculated using all the audio data. (See reference) Figure 2 Further details are provided for the loudness measuring component 114.

[0031] The mixing unit 116 receives audio object data 132, audio description data 136, and loudness data 138, applies gain, and typically performs a mixing process as further described herein. The mixing unit 116 also receives metadata 140, generates visualization data 142, generates metadata 144, and generates the mixed audio object data 146. The metadata reader 118, visualization unit 120, metadata writer 122, and audio object writer 124 can be considered as functional components of the mixing unit 116 that manipulate the inputs to generate the output.

[0032] Metadata reader 118 receives metadata 140. Typically, metadata 140 corresponds to a set of initial mixing parameters (also referred to as default mixing parameters). As discussed further in this document, the initial mixing parameters produce a set of gain adjustments that audio engineers can adjust as needed; the adjusted mixing parameters can be referred to as the adjusted mixing parameters. Metadata 140 can be in various formats, such as Extensible Markup Language (XML) format, JavaScript Object Notation (JSON) format, etc.

[0033] The visualization component 120 generates a gain adjustment visualization based on the mixing parameters and loudness data 138. Typically, this gain adjustment visualization shows the loudness of the audio object data 132, the loudness of the audio description data 136, and the gain to be applied for mixing according to the mixing parameters and loudness data 138. Audio engineers can then use the gain adjustment visualization to evaluate the gain of the proposed mix and adjust these gains as needed, thereby producing adjusted mixing parameters. (Reference) Figure 4 Provides further details on gain adjustment visualization.

[0034] Metadata writer 122 generates metadata 144. Metadata 144 corresponds to mixing parameters and loudness data 138. If the default mixing parameters produce an acceptable audio mix, the parameters in metadata 140 can be used as the parameters in metadata 144 without any adjustment. However, audio engineers typically adjust the mixing parameters based on the default parameters to generate metadata 144. The mixing parameters and loudness data 138 represented by metadata 144 can be referred to as adjusted mixing parameters during the adjustment process, and can be referred to as final mixing parameters once the audio engineer has completed the gain adjustment.

[0035] Audio object writer 124 and audio object metadata writer 148 work together to generate a mixed audio output. Audio object writer 124 mixes audio object data 132 and audio description data 136 according to mixing parameters and loudness data 138, generating mixed audio object data 146. The mixing parameters can be initial mixing parameters or adjusted mixing parameters. The mixed audio object data 146 then includes gain-adjusted audio object data and gain-adjusted audio description data. The mixed audio object data 146 can include audio objects, audio bed channels, etc. Audio object data 132 and audio description data 136 can be mixed according to two options (including their gain adjustment).

[0036] One option is to mix the audio description into one or more appropriate bed channels based on its format. For example, a mono audio description can be mixed into a center bed, a stereo audio description can be mixed into left and right bed channels, and a 5.1 channel audio description can be mixed into a 5.1 channel bed, etc. This option is useful when the total number of available audio objects is limited.

[0037] Another option is to create one or more new audio objects corresponding to one or more appropriate positions, which correspond to the format of the audio description. For example, an audio object located at the center position can be generated for a mono audio description, two audio objects located at corresponding left and right positions can be generated for a stereo audio description, and five audio objects located at 5.1 channel surround positions can be generated for a 5.1 channel audio description, and so on.

[0038] The audio object metadata writer 126 generates audio object metadata 148 related to the mixed audio object data 146. For example, the audio object metadata 148 may include the location information of each audio object, the size information of each audio object, etc.

[0039] The following is a brief overview of the operation of the audio mixing system 100. The audio engineer selects an audio object file and an audio description file; the audio object reader 110 generates corresponding audio object data 132, and the audio description reader generates corresponding audio description data 136. The loudness measurement unit 114 generates loudness data 138. The mixing unit 116 reads metadata 140 and applies mixing parameters to the loudness data 138 to generate gain visualization data 142. The audio mixing system 100 displays the gain visualization data 142, and the audio engineer evaluates the gain visualization.

[0040] Based on gain visualization, audio engineers can adjust the gain; the mixing unit 116 adjusts the mixing parameters to correspond to the adjusted gain and displays the adjusted gain visualization corresponding to the adjusted mixing parameters. The display, evaluation, and adjustment process can be performed multiple times, iteratively, etc.

[0041] Once the mixing engineer completes the evaluation (based on the initial mixing parameters or the adjusted mixing parameters), the mixing component 116 generates metadata 144 corresponding to the final mixing parameters and generates mixed audio object data 146 based on the final mixing parameters.

[0042] The mixing process performed using the audio mixing system 100 can generate mixed audio much faster than existing mixing systems. For example, an audio mix of a 90-minute movie can be generated in 30 minutes using the initial mixing parameters.

[0043] The following are further details of the audio mixing system 100.

[0044] Figure 2 This is a block diagram of loudness measuring component 200. Loudness measuring component 200 can be used as loudness measuring component 114 (see [link]). Figure 1 The loudness measurement unit 200 typically receives audio object data 132 and audio description data 136, performs a loudness measurement, and generates loudness data 138 (see [link to relevant documentation]). Figure 1 The loudness measurement component 200 includes a spatial encoding component 202, a renderer 204, and a loudness meter 206.

[0045] Spatial encoding component 202 receives audio object data 132, performs spatial encoding, and generates cluster data 210. Audio object data 132 typically includes audio objects, each containing audio data and location metadata indicating where the audio data should be output. Audio object data 132 may also contain audio beds. Spatial encoding component 202 performs spatial encoding to reduce the number of objects and beds to a smaller number of clusters. For example, audio object data 132 may contain up to 128 objects and beds, which the spatial encoding component groups into elements (also referred to as clusters). Cluster data 210 may contain multiple clusters, such as 12 or 16 clusters. These clusters may be in surround sound channel formats, such as an 11.1-channel format for 12 clusters, a 15.1-channel format for 16 clusters, etc. Typically, spatial encoding component 202 performs spatial encoding by dynamically grouping audio objects into dynamic groups, where audio objects can move from one cluster to another as the location information of the audio objects changes, and clusters can also move.

[0046] Renderer 204 receives cluster data 210, performs rendering, and generates render data 212. Typically, renderer 204 performs rendering by associating clusters in cluster data 210 with channels in render data 212. Render data 212 can be one of various channel formats, including mono (1 channel), stereo (2 channels), 5.1 channel surround (6 channels), 7.1 channel surround (8 channels), etc. Render data 212 can have one of various encoding formats, including pulse code modulation (PCM) signals, linear PCM (LPCM) signals, A-law PCM signals, etc. According to a specific example embodiment, render data 212 is a 5.1 channel LPCM signal.

[0047] Loudness meter 206 receives rendering data 212 and audio description data 136, performs loudness measurement, generates long-term loudness data 220 and short-term loudness data 222 for the rendering data 212, and generates long-term loudness data 224 and short-term loudness data 226 for the audio description data 136. In general, loudness data 220, 222, 224, and 226 correspond to loudness data 138 (see [link to relevant documentation]). Figure 1 ).

[0048] The loudness meter 206 can perform one of several loudness measurement procedures. Example loudness measurement procedures include the Leq (loudness equivalent continuous sound pressure level) procedure, the LKFS (loudness, K-weighted, relative to full scale) procedure, the LUFS (loudness unit relative to full scale) procedure, and so on.

[0049] Loudness meter 206 can be made by, for example, Dolby TM Computer programs such as the Professional Loudness Measurement (DPLM) Development Kit are implemented. Loudness meter 206 calculates long-term loudness data 220 and 224 to determine the overall level of this input; this value can be used to normalize the input (referred to as dialogue normalization or "dialogue normalization") so that there is no interfering loudness difference between the rendered data 212 and the audio description data 136. An example target value for dialogue normalization is -31 dB.

[0050] Typically, short-term loudness data 222 and 226 correspond to values ​​ordered by time, where each loudness measurement corresponds to the loudness of a specific portion of the input (e.g., a sample, frame, etc.). Typically, long-term loudness data 220 and 224 correspond to the overall loudness of each corresponding input, but they can also be ordered by time if these long-term loudness data are calculated from multiple portions of the input. Loudness data 138 can be formatted in a hierarchical format, such as as Extensible Markup Language (XML) data.

[0051] Figure 3 The hybrid component 116 is shown (see Figure 1 A block diagram of the additional components. These additional components are typically used when processing loudness data 138 according to mixing parameters. The additional components include a leading component 302, a ramp component 304, and a maximum difference component 306. The mixing component 116 may include components for processing other mixing parameters as needed.

[0052] Originally in metadata 140 (see Figure 1The mixing parameters are provided in the mixing component 116. Multiple sets of initial mixing parameters may exist for the mixing engineer to choose from, such as multiple sets of initial mixing parameters corresponding to various genres of the audiovisual content. The mixing engineer then selects a set of mixing parameters corresponding to the genre of the audio object file 134 and provides these initial parameters as metadata 140 to the mixing component 116. Example genres include action, horror, suspense, news, dialogue, sports, and talk show genres.

[0053] Table 1 provides examples of the initial blending parameters for the action genre included in metadata 140:

[0054] Parameters (units) value The lead length (s) of the main audio. 1.0 The slope begins to shift (s) -0.192 Ramp end offset (s) -0.192 Target maximum difference (dB) 30 Minimum gain 0.4 The preceding length (s) of the audio description 2.0

[0055] Table 1

[0056] The lead component 302 processes the lead parameter. The lead length of the main audio corresponds to the lead time period used by the mixing component 116 when processing loudness data 138 to skip the main audio in the presence of an audio description. If the audio description has stopped and restarts before the value of this parameter (e.g., 1.0 seconds), the ramp is not released during the stop period. (Example in...) Figure 4 (This is shown in the diagram and discussed in more detail there.) This parameter prevents large fluctuations in the loudness of the main audio, which could otherwise occur during short pauses in the audio description. This parameter may differ for other genres; for example, it may be increased (e.g., 2.0 seconds) for news genres.

[0057] The lead length of the audio description corresponds to the lead time period used by the mixing component 116 to tune the gain value of the audio description when processing the loudness data 138. For example, the lead component 302 can process the short-term loudness data 226 of the audio description within the upcoming time period corresponding to the value of this parameter (e.g., 2.0 seconds). Figure 2 Based on this processing, the gain to be applied to the audio description can be increased or decreased. As another example, the advance component 302 can process both the short-term loudness data 226 of the audio description and the short-term loudness data 222 of the main audio in the upcoming time period corresponding to the value of this parameter (e.g., 2.0 seconds), and based on this processing, the gain to be applied to both the audio description and the main audio can be increased or decreased.

[0058] The ramp component 304 processes ramp parameters. The ramp start offset corresponds to the length of time it takes for the gain to be gradually applied to the main audio as the audio description begins. This gain is applied gradually rather than instantaneously to reduce the likelihood of the reduced main audio quality disrupting the listener's experience. For example, if the gain to be applied to the main audio in the mixed audio description is 0.3, the gain will not change instantaneously from 1.0 to 0.3, but will change gradually over the ramp start offset period. For action genres, a time period of 0.192 seconds works well. This period can be adjusted for other genres. For example, a larger time period (e.g., 0.384 seconds) works well for dramatic genres.

[0059] The ramp-end offset corresponds to the length of time it takes for the gain to be gradually released to the main audio as the audio description ends. For example, if a gain of 0.3 has already been applied during the audio description, the gain gradually increases back to 1.0 over the ramp-end offset period (e.g., 0.192 seconds). The ramp-end offset period may differ from the ramp-start offset period, or they may be the same. For action genres, a 0.192-second period works well. This period can be adjusted for other genres. For example, a larger period (e.g., 0.384 seconds) works well for dramatic genres.

[0060] The maximum difference component 306 processes the target maximum difference and minimum gain parameters. The target maximum difference corresponds to the difference in loudness level between the main audio and the audio description; when this difference is exceeded, gain will be applied to the main audio. If the loudness difference is less than this level, gain will not be applied to the main audio, even if the audio description is present. This feature is useful when there is a quiet scene with background music and an audio description; if the main audio is skipped, the background music may not be audible from the audio description, thus disrupting the director's intent for the audio scene.

[0061] Minimum gain corresponds to the minimum gain applied to the audio description. This value prevents the audio description from being too loud compared to the main audio description; for example, in other quiet scenarios, the audio description might be so loud that it disrupts the listener's immersion in the audio scene. In these extreme cases, skipping the audio description allows the listener to remain immersed in the audio scene.

[0062] As mentioned above, the parameters in Table 1 are a set of initial parameters provided to the mixing unit 116 via metadata 140. Corresponding to genre, and the audio mixing system 100 can store multiple sets of mixing parameters, each set corresponding to one of multiple genres. Additionally, the values ​​of the parameters used as initial parameters can be adjusted. For example, for the action genre (see Table 1), the ramp start offset value can be changed from -0.192 to -0.182. The value of -0.182 is then used as one of the initial parameters provided via metadata 140. This allows mixing engineers to adjust these default mixing parameters before they are input into the mixing unit 116. Furthermore, multiple sets of parameters can exist for a given genre. For example, for the action genre, one set of parameters can have a ramp start offset value of -0.190, and another set of parameters can have a ramp start offset value of -0.195.

[0063] The audio mixing system 100 can handle mixing parameters other than those detailed in Table 1. For example, the default master audio ducking parameter can set the default gain value to be applied when ducking the master audio. This parameter can be defined as a gain level (e.g., a gain of 0.3), a decibel level (e.g., -16dB), etc. As another example, enabling the minimum gain parameter to be used for ducked audio descriptions (as discussed above) is an artful option that can be toggled on and off based on the parameter.

[0064] Figure 4 It shows the visualization data 142 (see Figure 1 The visualization 402 is represented by a graph 400. In graph 400, the x-axis is the sample index of the main audio data (e.g., audio object data 132) and the audio description data (e.g., audio description data 136). The x-axis can be viewed as a time index, where the beginning of the content is zero on the left and the end of the content is on the right. The y-axis on the left shows the gain to be applied to the main audio and the audio description, and the y-axis on the right shows the loudness level (in dB) of the main audio and the audio description.

[0065] Visualization 402 is an example of a selected audio object file (e.g., 130) and a selected audio description file (e.g., 134), showing gain and loudness. Gain is shown as dashed lines; line 410 shows the gain to be applied to the main audio, and line 412 shows the gain to be applied to the audio description. As discussed above, these gains correspond to those applied to the loudness data 138 (see...). Figure 1 The loudness level is represented by lines 414 and 416; line 414 indicates the loudness of the main audio, and line 416 indicates the loudness of the audio description. Note that line 416 is discontinuous; without line 416, there is no audio description.

[0066] Visualization 402 shows several features. Note that the gain (line 412) applied to the audio description is constant at 1.0. This indicates that mixing unit 116 determines, after considering mixing parameters and loudness data 138, that gain adjustment does not need to be applied to the audio description. For example, a comparison of the overall loudness between the main audio and the audio description might be within the values ​​defined by the mixing parameters.

[0067] Note that the gain to be applied to the master audio (line 410) is primarily in the range of 1.0 to 0.3, except around point 420. The downslope from 1.0 to 0.3 and the upslope from 0.3 to 1.0 are not easily visible due to the x-axis scale, but they exist according to the mixing parameters for the start and end offsets of the ramps (see Table 1). Furthermore, note that the gain of 0.3 can be configured using mixing parameters, such as the default master audio ducking parameters. Around point 420, the gain to be applied is approximately 0.32; this is a result of the interaction between mixing parameters applied to short-term loudness. For example, the default parameters can produce this gain, or a mixing engineer can adjust the mixing parameters to produce this gain (e.g., in response to hearing the mixed audio to generate a more acceptable mix).

[0068] Note that when an audio description (line 416) is present, there is usually a gain (line 410) to be applied to the main audio. However, line 410 also exists at some indices where no audio description is present, such as around point 422. This indicates a short break in the audio description that is smaller than the precedence length value defined in the mixing parameters (see Table 1).

[0069] A mixing engineer can use visualization 402 to evaluate proposed gains to be applied in the mix. For example, default parameters produce a first visualization, which the mixing engineer evaluates. If the first visualization appears to indicate that an acceptable mix will be produced, the mixing engineer can instruct the audio mixing system 100 to generate an audio mix without any adjustments. However, if the first visualization reveals some discontinuities or other visual features that indicate an unacceptable mix will be produced, the mixing engineer can adjust the mixing parameters, and the audio mixing system 100 can generate a second visualization based on the adjusted parameters. (For example, the mixing parameters can be adjusted to produce a slightly different appearance of line 410 around point 420.) The process of displaying a revised visualization, evaluating the revised visualization, and adjusting the mixing parameters can be performed iteratively (or otherwise multiple times) until the visual features of the revised visualization indicate that an unacceptable mix will be produced.

[0070] Additionally, to recap, the metadata 144 corresponding to the final mix parameters is generated around the time the mixed audio object data 146 is generated. This allows mixing engineers to evaluate the mixed audio; if the mixed audio is unacceptable, the mixing engineer can instruct the audio mixing system 100 to use the metadata 144 as input to the mixing component 116 (e.g., as mix parameters 140), and then perform evaluation and adjustments based on the adjusted parameters instead of the default parameters.

[0071] Figure 5 This is a block diagram of audio mixing system 500. It is related to audio mixing system 100, which is described as processing object audio (see [link]). Figure 1 In comparison, the audio mixing system 500 can be used to process other types of audio. The audio mixing system 500 includes a converter 502, an audio mixing system 100, and a converter 504.

[0072] Converter 502 receives audio data 510, converts the audio data 510, and generates audio object data (e.g., audio object file 130, audio object data 132, etc.). Audio data 510 typically corresponds to audio data that does not include audio objects, and converter 502 performs a conversion to transform audio data 510 into object audio data. For example, audio data 510 may be in a high-fidelity stereo format, and audio object file 130 may be in Dolby Atmos format. TM Format; Converter 502 can implement high-fidelity stereo to Dolby Atmos TM Conversion. Audio data 510 can typically correspond to the main audio of audiovisual content (e.g., movie soundtrack).

[0073] The audio mixing system 100 processes the audio object file 130 (generated by the conversion) and generates mixed audio object data 146 and audio object metadata 148, as discussed above.

[0074] Converter 504 receives the mixed audio object data 146, converts the mixed audio object data 146, and generates mixed audio data 512. Converter 504 may also receive mixed audio object data 148. The mixed audio data 512 then corresponds to an audio description mixed with audio data 510. Typically, converter 504 performs the inverse operation of the conversion performed by converter 502. For example, when converter 502 implements high-fidelity stereo to Dolby Atmos... TM During conversion, converter 504 implements Dolby Atmos. TM Convert to high-fidelity stereo.

[0075] In this way, the audio mixing system 500 enables the audio mixing system 100 to be used with other types of audio.

[0076] Figure 6 This is a device architecture 600 according to embodiments for implementing the features and processes described herein. Architecture 600 can be implemented in any electronic device, including but not limited to: desktop computers, consumer audio / video (AV) devices, radio broadcasting devices, mobile devices (e.g., smartphones, tablets, laptops, wearable devices), etc. In the example embodiments shown, architecture 600 is for a laptop computer and includes processor(s) 601, peripheral interface 602, audio subsystem 603, loudspeaker 604, microphone 605, sensors 606 (e.g., accelerometer, gyroscope, barometer, magnetometer, camera), location processor 607 (e.g., GNSS receiver), wireless communication subsystem 608 (e.g., Wi-Fi, Bluetooth, cellular), and I / O subsystem(s) 609 (including touch controller 610 and other input controllers 611), touch surface 612, and other input / control devices 613. Other architectures with more or fewer components can also be used to implement the disclosed embodiments.

[0077] Memory interface 614 is coupled to processor 601, peripheral interface 602, and memory 615 (e.g., flash memory, RAM, ROM). Memory 615 stores computer program instructions and data, including but not limited to: operating system instructions 616, communication instructions 617, GUI instructions 618, sensor processing instructions 619, telephone instructions 620, electronic messaging instructions 621, web browsing instructions 622, audio processing instructions 623, GNSS / navigation instructions 624, and application / data 625. Audio processing instructions 623 include instructions for performing the audio processing described herein.

[0078] As a specific example, device architecture 600 can implement audio mixing system 100, for instance, by executing one or more computer programs (see [link to documentation]). Figure 1 The device architecture can access audio file 130 via peripheral device interface 602 (e.g., connected to non-volatile storage such as solid-state drive), can use processor 601 to calculate loudness data 138, can display visualization data 142 via peripheral device interface 602 (e.g., connected to display device), and can use processor 601 to generate mixed audio object data 146.

[0079] Figure 7 This is a flowchart of audio processing method 700. Method 700 can be performed by a... Figure 6The components of the architecture 600 are executed by a device (e.g., a laptop computer, desktop computer, etc.) to implement the audio mixing system 100, for example, by executing one or more computer programs (see...). Figure 1 Functions such as )

[0080] At 702, audio object data and audio description data are received. The audio object data includes a first set of audio objects. For example, loudness measurement component 114 and mixing component 116 (see...). Figure 1 It can receive audio object data 132 and audio description data 136.

[0081] At 704, the long-term loudness of the audio object data and the long-term loudness of the audio description data are calculated. For example, loudness measurement component 114 (see...) Figure 1 Long-term loudness can be calculated as part of the loudness data 138. Long-term loudness can also be calculated using all the data.

[0082] At 706, multiple short-term loudnesses of the audio object data and multiple short-term loudnesses of the audio description data are calculated. For example, loudness measurement component 114 (see...) Figure 1 Short-term loudness can be calculated as part of loudness data 138. Short-term loudness can be calculated on a continuous basis (e.g., per sample, per frame, etc.).

[0083] At 708, the first set of mixing parameters corresponding to the audio object data is read. For example, metadata reader 118 (see...) Figure 1 It can read metadata 140 containing initial mixing parameters.

[0084] At 710, a second set of mixing parameters is generated based on the first set of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the short-term loudness of the audio object data, and the short-term loudness of the audio description data. For example, mixing component 116 (see...) Figure 1 The loudness data 138 can be processed based on the initial mixing parameters to generate a set of proposed gains for the main audio and audio description (e.g., with...). Figure 4 (Lines 410 and 412 correspond to each other).

[0085] At 712, a gain adjustment visualization is generated. This gain adjustment visualization corresponds to the second set of mixing parameters, audio object data, and audio description data. For example, this gain adjustment visualization could correspond to visualization 402, which shows the loudness of the main audio, the loudness of the audio description, and the proposed gain (see...). Figure 4 ).

[0086] At 714, evaluate the gain adjustment visualization to determine if applying the mixing parameters will produce acceptable results. For example, a mixing engineer could evaluate visualization 402 (see...). Figure 4 If the result is unacceptable, the process proceeds to 716; if the result is acceptable, the process proceeds to 718.

[0087] At 716, the second set of mixing parameters is adjusted. For example, a mixing engineer can adjust the proposed gain, and the mixing unit 116 can accordingly adjust the mixing parameters to correspond to the adjusted gain.

[0088] At point 718, mixed audio object data is generated by mixing audio object data and audio description data according to a second set of mixing parameters. The mixed audio object data includes a second set of audio objects, which corresponds to a first set of audio objects mixed with the audio description data according to the second set of mixing parameters. For example, audio object writer 124 (see...) Figure 1 It can generate mixed audio object data 146. It can also generate audio object metadata related to the mixed audio object data. For example, the audio object metadata writer 126 can generate audio object metadata 148.

[0089] Method 700 may include additional steps corresponding to other functions of the audio mixing system 100, as described herein. For example, default mixing parameters may be selected based on the genre of the audiovisual content being mixed. Mixing parameters may include precedence parameters, ramp parameters, maximum difference parameters, etc. Method 700 may include a conversion step that converts non-object audio to object audio for processing by the audio mixing system 100 and converts the mixed object audio back to mixed non-object audio.

[0090] Additional details

[0091] While this description focuses on audio mixing, the embodiments can also be used to mix other types of audio content to achieve similar improvements in time, effort, and efficiency. For example, audio mixing system 100 (see...) Figure 1 It can be used to mix director's comments.

[0092] Implementation details

[0093] Embodiments may be implemented in hardware, as executable modules stored on a computer-readable medium, or a combination of both (e.g., a programmable logic array). Unless otherwise stated, the steps performed by the embodiments do not need to be inherently associated with any particular computer or other device, although they may be relevant in some embodiments. Specifically, various general-purpose machines may be used with programs written in accordance with the teachings herein, or more specialized devices (e.g., integrated circuits) may be more readily constructed to perform the desired method steps. Therefore, embodiments may be implemented in one or more computer programs that execute on one or more programmable computer systems, each including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0094] Each such computer program is preferably stored or downloaded to a storage medium or device (e.g., solid-state memory or medium, or magnetic or optical medium) readable by a general-purpose or special-purpose programmable computer, for configuring and operating the computer to execute the program described herein when the computer system reads the storage medium or device. The system of the present invention can also be considered as an embodiment of a computer-readable storage medium configured with a computer program, wherein such a configuration causes the computer system to operate in a specific and predefined manner to perform the functions described herein. (Software itself and intangible or transient signals are excluded in the sense that they are not patentable subject matter.)

[0095] The aspects of the system described herein can be implemented in a suitable computer-based sound processing network environment to process digital or digitized audio files. Parts of the adaptive audio system may include one or more networks comprising any desired number of independent machines, including one or more routers (not shown) for buffering and routing data transmitted between computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0096] One or more components, blocks, processes, or other functional elements may be implemented by a computer program executed by a processor-based computing device of a control system. It should also be noted that the various functions disclosed herein may be described from the perspective of behavior, register transfer, logic components, and / or other characteristics using any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media. Computer-readable media that may embody such formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media such as optical, magnetic, or semiconductor storage media.

[0097] The above description illustrates various embodiments of the present disclosure and examples of how aspects of the present disclosure may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of the present disclosure defined by the appended claims. Based on the above disclosure and the appended claims, other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be employed without departing from the spirit and scope of the present disclosure defined by the claims.

Claims

1. A computer-implemented audio processing method, the method comprising: Receive audio object data and audio description data, wherein the audio object data includes a first plurality of audio objects; Calculate the long-term loudness of the audio object data and the long-term loudness of the audio description data; Calculate multiple short-term loudnesses of the audio object data and multiple short-term loudnesses of the audio description data; Read a first plurality of mixing parameters corresponding to the audio object data, wherein the first plurality of mixing parameters include an advance parameter, a ramp parameter, and a maximum difference parameter; Based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data, a second plurality of mixing parameters are generated; Generate a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data; and Mixed audio object data is generated by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data includes a second plurality of audio objects, wherein the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters.

2. The method as described in claim 1, wherein, The long-term loudness of the audio object data is calculated from multiple samples of the audio object data, and the long-term loudness of the audio description data is calculated from multiple samples of the audio description data. Each of the plurality of short-term loudnesses of the audio object data is calculated using a single sample of the audio object data, and each of the plurality of short-term loudnesses of the audio description data is calculated using a single sample of the audio description data.

3. The method according to any one of claims 1 to 2, wherein, The first plurality of blending parameters are associated with one of a plurality of genres, wherein each of the plurality of genres is associated with a corresponding set of blending parameters.

4. The method of claim 3, wherein, The various genres include action, horror, suspense, news, dialogue, sports, and talk show.

5. The method of claim 1, wherein, The preceding parameters correspond to maintaining a uniform gain adjustment during audio pauses in the audio description data.

6. The method as described in any one of claims 1 or 5, wherein, The ramp parameter corresponds to the time period during which gain adjustment is gradually applied.

7. The method as described in any one of claims 1 or 5, wherein, The maximum difference parameter corresponds to the maximum loudness difference between the frames of the audio object data and the corresponding frames of the audio description data.

8. The method of any one of claims 1 to 2, further comprising: Before generating the mixed audio object data, user input is received to adjust the second plurality of mixing parameters; as well as Generate a modified gain adjustment visualization corresponding to the second plurality of mixing parameters that have been adjusted according to the user input. The mixed audio object data is generated based on the adjusted second plurality of mixing parameters.

9. The method of any one of claims 1 to 2, further comprising: Before receiving the audio object data: Receive audio data, wherein the audio data does not include an audio object; and Convert the audio data into the audio object data, and After generating the mixed audio object data: The mixed audio object data is converted into mixed audio data, wherein the mixed audio data corresponds to the audio data mixed with the audio description data.

10. A non-transitory computer-readable medium storing a computer program that, when executed by a processor, controls means to perform processing including the method as described in any one of claims 1 to 9.

11. An apparatus for audio processing, the apparatus comprising: processor, The processor is configured to control the device to receive audio object data and audio description data, wherein the audio object data includes a first plurality of audio objects. The processor is configured to control the device to calculate the long-term loudness of the audio object data and the long-term loudness of the audio description data. The processor is configured to control the device to calculate multiple short-term loudnesses of the audio object data and multiple short-term loudnesses of the audio description data. The processor is configured to control the device to read a first plurality of mixing parameters corresponding to the audio object data, wherein the first plurality of mixing parameters includes an anticipation parameter, a ramp parameter, and a maximum difference parameter. The processor is configured to control the device to generate a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, a plurality of short-term loudnesses of the audio object data, and a plurality of short-term loudnesses of the audio description data. The processor is configured to control the device to generate a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data. The processor is configured to control the device to generate mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data includes a second plurality of audio objects, wherein the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters.

12. The apparatus of claim 11, further comprising: A display configured to display the gain adjustment visualization.

13. The apparatus as claimed in any one of claims 11 to 12, wherein, The long-term loudness of the audio object data is calculated from multiple samples of the audio object data, and the long-term loudness of the audio description data is calculated from multiple samples of the audio description data. Each of the plurality of short-term loudnesses of the audio object data is calculated using a single sample of the audio object data, and each of the plurality of short-term loudnesses of the audio description data is calculated using a single sample of the audio description data.

14. The apparatus as claimed in any one of claims 11 to 12, wherein, The first plurality of blending parameters are associated with one of a plurality of genres, wherein each of the plurality of genres is associated with a corresponding set of blending parameters.

15. The apparatus of claim 11, wherein, The preceding parameters correspond to maintaining a uniform gain adjustment during audio pauses in the audio description data.

16. The apparatus as claimed in any one of claims 11 or 15, wherein, The ramp parameter corresponds to the time period during which gain adjustment is gradually applied.

17. The apparatus as claimed in any one of claims 11 or 15, wherein, The maximum difference parameter corresponds to the maximum loudness difference between the frames of the audio object data and the corresponding frames of the audio description data.

18. The apparatus as claimed in any one of claims 11-12, wherein, The processor is configured to control the device to receive user input to adjust the second plurality of mixing parameters before generating the mixed audio object data; The processor is configured to: control the device to generate a corrected gain adjustment visualization corresponding to the second plurality of mixing parameters adjusted according to the user input; and The mixed audio object data is generated based on the adjusted second plurality of mixing parameters.