Audio production model file creation method and electronic equipment

By generating audio production model files, the connection problem between metadata and audio files is solved, efficient management and accurate mapping are achieved, the efficiency and quality of audio production are improved, and dynamic updates and immersion are improved.

CN120448580APending Publication Date: 2025-08-08SINE MICRO (BEIJING) ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399799.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, the metadata file lacks an effective connection method with the original audio file, which causes the metadata to be unable to accurately map to the corresponding parts of the audio file, reducing the efficiency and quality of audio production.

Method used

By obtaining the original audio file, generating preset metadata elements, and establishing reference relationships between these elements, creating channel allocation files for audio tracks, generating audio production model files, supporting dynamic updates and real-time binding, and adjusting reverberation parameters based on room acoustic measurement data.

Benefits of technology

It realizes efficient management and accurate correspondence between metadata and audio files, improves the efficiency and quality of audio production, supports dynamic updates, reduces data redundancy, and improves user immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448580A_ABST
    Figure CN120448580A_ABST
Patent Text Reader

Abstract

The invention relates to an audio production model file creation method and electronic equipment. The method comprises the steps of obtaining an original audio file; analyzing the original audio file through a renderer, and generating corresponding preset metadata elements; wherein the preset metadata elements comprise an audio program element, an audio content element, an audio object element, an audio track unique identification element, an audio packet format element, an audio channel format element, an audio stream format element and an audio track format element; based on the relation between the preset metadata elements, establishing a reference relation of the corresponding elements; generating a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship; creating a channel allocation file for connecting the metadata file and an audio track of the original audio file; and establishing a guide relationship between the metadata file and the channel allocation file, and generating an audio production model file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular to a method for creating an audio production model file and an electronic device. Background Art

[0002] As technology advances, audio becomes increasingly complex. From early mono audio to stereo, the focus has been on correctly processing the left and right channels. However, with the advent of surround sound, processing has become increasingly complex. Surround 5.1 speaker systems prioritize the sequencing of multiple channels, and subsequent 6.1 and 7.1 speaker systems have further complicated audio processing, ensuring the correct signal is delivered to the appropriate speakers to create a coherent effect. Consequently, as sound becomes more immersive and interactive, the complexity of audio processing has increased significantly.

[0003] Audio channels (or channels) refer to independent audio signals collected or played back at different spatial locations during recording or playback. The number of channels refers to the number of sound sources during recording or the number of corresponding speakers during playback. For example, a 5.1 surround speaker system includes six audio signals at different spatial locations, and each independent audio signal is used to drive a speaker at the corresponding spatial location; a 7.1 surround speaker system includes eight audio signals at different spatial locations, and each independent audio signal is used to drive a speaker at the corresponding spatial location.

[0004] However, if there is no effective way to connect the metadata file and the original audio file, it is difficult for them to work together in practical applications, so that the metadata cannot be accurately mapped to the corresponding part of the original audio file, which greatly reduces the efficiency and quality of audio production.

[0005] The present application provides a method for creating an audio production model file, so as to provide an audio production model that can solve the above-mentioned technical problems. Summary of the Invention

[0006] The purpose of this application is to provide a method and electronic device for creating an audio production model file to link metadata and original audio files.

[0007] To achieve the above objectives, the present application provides, in a first aspect, a method for creating an audio production model file, comprising:

[0008] Get the original audio file;

[0009] Analyzing the original audio file through a renderer to generate corresponding preset metadata elements; wherein the preset metadata elements include an audio program element, an audio content element, an audio object element, an audio track unique identifier element, an audio package format element, an audio channel format element, an audio stream format element, and an audio track format element;

[0010] Based on the relationship between the preset metadata elements, establishing a reference relationship between the corresponding elements;

[0011] Generating a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship;

[0012] creating a channel allocation file for connecting the metadata file and the audio track of the original audio file;

[0013] A reference relationship between the metadata file and the channel allocation file is established to generate an audio production model file.

[0014] A second aspect of the present application provides an electronic device, comprising: a memory and one or more processors;

[0015] The memory is used to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors execute the audio production model file creation method provided in any embodiment of the present application.

[0017] The audio production model file creation method and electronic device provided by the present application analyze the original audio file, determine the references between the metadata, and generate an audio production model. Each element of the model is used to describe various aspects of the audio, and these elements can be connected to each other through references to achieve the reproduction of three-dimensional sound. During the production process, the management and coordination of metadata format elements can be efficiently achieved, so that the metadata can accurately correspond to the corresponding part of the original audio file. On the other hand, the audio production model file of the present invention supports dynamic updating and real-time binding, avoids data redundancy, reduces the model volume accordingly, and improves processing efficiency. On the other hand, combined with room acoustic measurement data (such as impulse response), the reverberation parameters in the metadata can be dynamically adjusted to adapt the sound field to different spatial positions, thereby improving the user's immersion score and improving the efficiency and quality of audio production. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of a method for creating an audio production model file is provided in an embodiment of the present application;

[0019] Figure 2 A schematic diagram of an audio production model is provided in an embodiment of the present application;

[0020] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following examples are used to illustrate the present application but are not used to limit the scope of the present application.

[0022] Metadata is information that describes the structural characteristics of data. Metadata supports functions such as indicating storage location, historical data, resource search, or file records. Each stage of the multidimensional audio model produces metadata that describes the characteristics of that stage.

[0023] Among them, the content production part includes: audio program elements, audio content elements, audio object elements and audio track unique identification elements. Among them, it should be noted that: audio program represents a complete audio project (such as movie soundtracks, game sound effects), which contains a logical collection of multiple audio contents; audio content represents independent audio material units (such as background music, character dialogues), which can be associated with multiple audio objects and audio tracks; audio objects represent audio elements with independent spatial attributes (such as helicopter sound effects), which contain metadata such as position and motion trajectory; audio track unique identification represents the assignment of a unique ID to each audio track for accurate reference and management to avoid confusion.

[0024] The format production part includes: audio package format element, audio channel format element, audio stream format element and audio track format element, among which it should be noted that: the audio package format defines the container format of audio data (such as ADMBWF, MP4), supports multi-track encapsulation and metadata embedding; the audio channel format specifies the channel configuration (such as stereo L / R, 5.1 surround sound), and defines the mapping relationship between physical or virtual channels; the audio stream format describes the encoding parameters of the audio stream (such as PCM24bit / 48kHz, AAC-LC) to ensure decoding consistency; the audio track format contains detailed parameters of the audio track (such as sampling rate, dynamic range control), and guides the renderer processing strategy.

[0025] The audio program element may be used to reference at least one audio content element; the audio content element may be used to reference at least one audio object element; the audio object element may be used to reference the corresponding audio package format element and the corresponding audio track unique identification element; the audio track unique identification element may be used to reference the corresponding audio track format element and the corresponding audio package format element;

[0026] The audio package format element can be used to reference at least one of the audio channel format elements; the audio stream format element can be used to reference the corresponding audio channel format element and the corresponding audio package format element; the audio track format element and the corresponding audio stream format element reference each other. The reference relationship between elements is as follows: Figure 2 Indicated by arrows.

[0027] The original audio data is produced through the three-dimensional audio production model to generate a storage format containing metadata and audio data.

[0028] The metadata is information that describes data characteristics. Metadata supports functions including indicating storage location, historical data, resource search, or file records.

[0029] After the audio format containing metadata is transmitted to the remote end via communication, the remote end renders the synthesized audio data based on the metadata to restore the original sound scene.

[0030] Figure 2 The figure shows the division process between the content creation section, the format creation section, and the BW64 (Broadcast Wave 64bit) encapsulation file. Both the content creation section and the format creation section constitute XML-formatted metadata, which is contained in a block (the "axml" block) of the BW64 file. The BW64 file section includes a channel allocation block, which is a lookup table used to connect metadata to audio tracks in the original audio file. The channel allocation information in the channel allocation block is stored in the metadata block (the "chna" block), and the output mapping rules are defined by the channel allocation block.

[0031] The audio production model in the embodiment of the present application includes a content production part and a format production part. The format production part may not exist in the content production part, but not vice versa.

[0032] Assuming that an audio production model is used in a BW64 file, the actual audio tracks in the BW64 file need to be linked with the audio production model tracks.

[0033] Example

[0034] This application provides a method for creating an audio production model file based on a multi-dimensional sound audio model and provides a detailed description. Figure 1 The flowchart of the method for creating an audio production model file shown in FIG. 1 includes:

[0035] Step 110: Obtain the original audio file;

[0036] Step 120: Analyze the original audio file through a renderer to generate corresponding preset metadata elements; wherein the preset metadata elements include an audio program element, an audio content element, an audio object element, an audio track unique identifier element, an audio package format element, an audio channel format element, an audio stream format element, and an audio track format element;

[0037] Step 130: Based on the relationship between the preset metadata elements, establish a reference relationship between the corresponding elements;

[0038] Step 140: Generate a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship;

[0039] Step 150: Create a channel allocation file for connecting the metadata file and the audio track of the original audio file;

[0040] Step 160: Establish a reference relationship between the metadata file and the channel allocation file, and generate an audio production model file.

[0041] Optionally, based on the relationship between the preset metadata elements, a reference relationship between corresponding elements is established, including:

[0042] The reference relationship of the corresponding elements is established based on the connection between the following elements:

[0043] The audio program element's reference to at least one of the audio content elements; the audio content element's reference to at least one audio object element; the audio object element's reference to the corresponding audio package format element and the corresponding audio track unique identification element; the audio track unique identification element's reference to the corresponding audio track format element and the corresponding audio package format element; the audio package format element's reference to at least one of the audio channel format elements; the audio stream format element's reference to the corresponding audio channel format element and the corresponding audio package format element; and mutual references between the audio track format element and the corresponding audio stream format element.

[0044] Optionally, generating a metadata file from the preset metadata element according to the preset metadata element and the reference relationship includes:

[0045] According to the preset metadata elements and the reference relationship, the preset metadata elements are used to generate a metadata file in a first file format.

[0046] Optionally, creating a channel allocation file for connecting the metadata file and the audio track of the original audio file includes:

[0047] Creating a lookup table, wherein the lookup table records the corresponding relationship between the metadata file and the audio track of the original audio file;

[0048] The lookup table is used as the channel allocation file.

[0049] Optionally, establishing a reference relationship between the metadata file and the channel allocation file to generate an audio metadata file includes:

[0050] According to a second file format, the metadata file is placed in a first file block, and the channel allocation file is placed in a second file block;

[0051] The channel allocation file and the metadata file are combined to generate the audio production model file.

[0052] Optionally, the second file format is BW64.

[0053] Optionally, the first file format is XML.

[0054] The track unique identification element is used to create the track unique identification and generate metadata of the track unique identification for describing the structural characteristics of the track unique identification.

[0055] The audio content describes the content of a component of the audio content (such as background music) and refers to one or more audio objects to associate the content with its format. The audio content element is to produce audio content and generate metadata for the audio content to describe the structural characteristics of the audio content.

[0056] The audio program includes narration, sound effects, and background music. The audio program references one or more audio content, which are combined to form a complete audio object. The audio program elements are used to create audio objects and generate metadata for the audio program, which describes the structural characteristics of the audio program.

[0057] Specifically, this embodiment uses a symphony recording master as the original audio file to demonstrate how to create an audio production model file using the method of the present invention to achieve multi-track metadata management and three-dimensional sound field rendering. The original audio file contains the following tracks:

[0058] Track 1: Violin (stereo, PCM 24bit / 48kHz);

[0059] Track 2: Cello (mono, PCM 24bit / 48kHz);

[0060] Track 3: Piano (5.1 surround sound, AAC-LC encoding).

[0061] Use the FFmpeg library to read the file header, extract the track list and encoding format from the multi-track file encapsulated in BW64 format, and output the track information.

[0062] The renderer analyzes the audio track and generates metadata elements in XML format. The relationship between metadata elements is stored in a hash table. The renderer generates a three-dimensional sound field in real time through the spatial attributes and channel allocation in the metadata, with a delay of less than 10ms.

[0063] Establish a mapping relationship between the audio tracks in the metadata file and the physical channels of the original audio file to ensure that the renderer accurately distributes the audio signal. When audio tracks are added or removed or the channel configuration changes, only the lookup table needs to be modified without regenerating the metadata.

[0064] Encapsulates metadata (XML) and channel assignments (CSV) into a single file in BW64 format for easy storage, transfer, and rendering. Block storage enables fast retrieval and partial updates, allowing renderers to parse BW64 files directly without decompressing the metadata and audio streams.

[0065] Furthermore, in some embodiments, a graph-like association relationship is constructed, and each element contains a reference list and a referenced list, and the reference relationship between elements can be represented based on the unique identifier of the element. Exemplarily, a unique UUID (universally unique identifier) is assigned to each metadata element (such as program element / content element / object element / audio track element, etc.) as the core identifier for cross-file references. For example, the reference relationship between elements is as follows: the audio object element obj-001 establishes a reference with the audio track element trk-001 through the UUID. For example, a graph-like association relationship is constructed: the audio track element references the audio object element (trk-001→obj-001), and the audio track element is referenced by the program element (prg-001→trk-001). A non-redundant element association is established through the UUID, and changes are automatically propagated.

[0066] Each metadata element can also include a version number and modification timestamp to record change history. Incremental updates are used to store only changes to properties or references, rather than completely rewriting the file. When an element changes (such as the spatial position of an audio object), related elements (such as the associated audio track or parent program) are automatically notified through reference relationships.

[0067] When modifying a track's channel assignment, the channel assignment file is updated synchronously. File system monitoring (such as FSEvents or inotify) captures changes to metadata or channel files, triggering re-parsing. The channel assignment file supports dynamic expressions (such as all_surround, which automatically adapts to the current device channel count), and the audio engine's channel routing is immediately updated when changes are made. This solution supports runtime changes to channel assignments, adapting to diverse scenarios.

[0068] This embodiment fully demonstrates the generation process from multi-track original files to audio production model files. Through XML metadata, CSV channel allocation table and BW64 packaging format, it systematically solves the problems of multi-channel management, dynamic reference and cross-platform compatibility.

[0069] The audio production model of the embodiment of the present invention also supports encoding formats such as PCM, MP3, AAC, Dolby Atmos, etc., and automatically adapts the decoding strategy through metadata reference, thereby improving content production efficiency by 60%.

[0070] In a dual-speaker system, the virtual 7.1.4-channel rendering driven by the audio production model of the embodiment of the present invention can match and be compatible with a variety of speaker layouts (compared to traditional physical surround sound systems).

[0071] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device includes: a processor 30, a memory 31, an input device 32 and an output device 33. The number of processors 30 in the electronic device can be one or more. Figure 3 In the example, a processor 30 is used. The number of memories 31 in the electronic device can be one or more. Figure 3 In the example, a memory 31 is used. The processor 30, memory 31, input device 32 and output device 33 of the electronic device can be connected through a bus or other means. Figure 3 The example of the bus connection is shown in FIG. The electronic device may be a computer or a server. The embodiment of the present application is described in detail with the electronic device being a server, which may be an independent server or a cluster server.

[0072] The memory 31, as a computer-readable storage medium, can be used to store software programs, computer executable programs, and modules, such as the program instructions / modules of the audio production model file creation method described in any embodiment of the present application. The memory 31 may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created according to the use of the device, etc. In addition, the memory 31 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 31 may further include a memory remotely located relative to the processor 30, and these remote memories can be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0073] The input device 32 can be used to receive input digital or character information and generate key signals related to the user settings and function control of the electronic device. It can also be a camera for capturing images and a sound pickup device for capturing audio data. The output device 33 can include audio equipment such as speakers. It should be noted that the specific composition of the input device 32 and output device 33 can be set according to actual circumstances.

[0074] The processor 30 executes the software programs, instructions and modules stored in the memory 31 to perform various functional applications and data processing of the device, that is, to generate audio metadata.

[0075] In the description of this specification, reference to the terms "in one embodiment," "in another embodiment," "exemplary," or "in a specific embodiment" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.

[0076] Although the present application has been described in detail above using general explanations, specific embodiments, and experiments, it will be apparent to those skilled in the art that modifications or improvements may be made to the present application. Therefore, such modifications or improvements, without departing from the spirit of the present application, are within the scope of protection claimed in the present application.

Claims

1. A method for creating an audio production model file, characterized in that: include: Get the original audio file; Analyzing the original audio file through a renderer to generate corresponding preset metadata elements; wherein the preset metadata elements include an audio program element, an audio content element, an audio object element, an audio track unique identifier element, an audio package format element, an audio channel format element, an audio stream format element, and an audio track format element; Based on the relationship between the preset metadata elements, establishing a reference relationship between the corresponding elements; Generating a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship; creating a channel allocation file for connecting the metadata file and the audio track of the original audio file; A reference relationship between the metadata file and the channel allocation file is established to generate an audio production model file.

2. The method for creating an audio production model file according to claim 1, wherein: Based on the relationship between the preset metadata elements, a reference relationship between the corresponding elements is established, including: The reference relationship of the corresponding elements is established based on the connection between the following elements: The audio program element's reference to at least one audio content element; the audio content element's reference to at least one audio object element; the audio object element's reference to the corresponding audio package format element and the corresponding audio track unique identification element; the audio track unique identification element's reference to the corresponding audio track format element and the corresponding audio package format element; the audio package format element's reference to at least one audio channel format element; the audio stream format element's reference to the corresponding audio channel format element and the corresponding audio package format element; and mutual references between the audio track format element and the corresponding audio stream format element.

3. The method for creating an audio production model file according to claim 2, wherein: Generating a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship includes: According to the preset metadata elements and the reference relationship, the preset metadata elements are used to generate a metadata file in a first file format.

4. The method for creating an audio production model file according to claim 3, wherein: Creating a channel allocation file for connecting the metadata file and the audio track of the original audio file, including: Creating a lookup table, wherein the lookup table records the corresponding relationship between the metadata file and the audio track of the original audio file; The lookup table is used as the channel allocation file.

5. The method for creating an audio production model file according to claim 4, wherein: Establishing a reference relationship between the metadata file and the channel allocation file to generate an audio metadata file includes: According to a second file format, the metadata file is placed in a first file block, and the channel allocation file is placed in a second file block; The channel allocation file and the metadata file are combined to generate the audio production model file.

6. The method for creating an audio production model file according to claim 5, wherein: The second file format is BW64.

7. The method for creating an audio production model file according to claim 3, wherein: The first file format is XML.

8. An electronic device, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors execute the audio production model file creation method according to any one of claims 1 to 7.