Audio packet format metadata generation method and device, equipment and medium

By using an audio package format metadata generation method, the problems of format compatibility and channel recognition complexity in audio processing are solved, enabling unified management and rendering of multiple audio types and improving the efficiency and quality of audio production.

CN121600942APending Publication Date: 2026-03-03SINE MICRO (BEIJING) ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511500796.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In the existing audio processing workflow, audio files of different types and encoding states have problems such as poor format compatibility, complex channel recognition logic, and inconsistent metadata structure when generating metadata. As a result, the audio package format metadata cannot efficiently support the standardized production and flexible expansion of multiple types of audio, and it is difficult to meet the unified management and rendering needs of multi-track and multi-format audio in digital audio workstations.

Method used

By using audio package format metadata generation methods, including audio track format recognition, encoded audio track recognition, and channel recognition, an audio definition model is generated. This model supports unified conversion and parsing of multiple audio types, adopts a structured attribute design, allows nested audio package formats, and supports unlimited expansion from simple mono to complex audio scenarios.

Benefits of technology

It improves the accuracy of channel resolution, eliminates rendering errors, achieves compatibility and flexible expansion for multiple audio types, and improves the efficiency and quality of audio production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600942A_ABST
    Figure CN121600942A_ABST
Patent Text Reader

Abstract

The invention relates to an audio packet format metadata generation method and device, equipment and a medium, and belongs to the technical field of audio processing, and the method comprises the steps: converting an obtained audio file containing audio track data or audio samples of original equipment into audio file data in a preset format according to an audio track type, and storing the audio file data in a digital audio workstation, inputting the audio file into the renderer input end; performing audio track format identification on the audio file according to the audio data or the audio sample type; carrying out coding audio track identification on the audio track format of the audio data; performing sound channel identification according to the audio track format of the input audio data or the generated audio stream format metadata; and if the audio file or the audio stream is identified as a group of sound channels, the renderer performs identification and analysis processing on the audio stream track or metadata to generate audio packet format metadata for making an audio definition model. And during rendering, reproduction of three-dimensional sound can be realized in the channel, and the quality of a sound scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application filed on April 3, 2025, with application number 202510421774.X and invention title "A method, apparatus, device and medium for generating audio package format metadata". Technical Field

[0002] This disclosure relates to the technical field of audio processing, and in particular to a method, apparatus, device, and medium for generating audio packet format metadata. Background Technology

[0003] With the advancement of technology, audio has become increasingly complex. Early mono audio evolved into stereo, with a focus on the correct processing of the left and right channels. However, the advent of surround sound made the processing even more complex. Surround 5.1 speaker systems, in particular, constrain the order of multiple channels, leading to a wide variety of audio processing methods, such as surround 5.1 and surround 7.1 speaker systems, ensuring the correct signals are delivered to the appropriate speakers to create interconnected effects. Therefore, as sound becomes more immersive and interactive, the complexity of audio processing increases significantly.

[0004] An audio channel (or channel) refers to an independent audio signal that is captured or played back from different spatial locations during recording or playback. The number of channels corresponds to the number of audio sources during recording or the number of speakers during playback. For example, a 5.1 surround sound system includes six audio signals from different spatial locations, each driving a speaker at its corresponding location; a 7.1 surround sound system includes eight audio signals from different spatial locations, each driving a speaker at its corresponding location.

[0005] Therefore, the effects achieved by current speaker systems depend on the number and spatial location of the speakers. For example, a two-channel speaker system cannot achieve the effect of a 5.1 surround sound system. In existing audio processing workflows, different types (track data / audio samples) and different encoding states (encoded / unencoded) of audio files suffer from poor format compatibility, complex channel recognition logic, and inconsistent metadata structures during metadata generation. This results in audio package format metadata being unable to efficiently support the standardized production and flexible expansion of multiple types of audio, making it difficult to meet the unified management and rendering needs of multi-track, multi-format audio in digital audio workstations (DAWs).

[0006] This disclosure provides a method for generating audio package format metadata, so as to provide metadata that can solve the above-mentioned technical problems. Summary of the Invention

[0007] The purpose of this disclosure is to provide a method, apparatus, device, and medium for generating audio packet format metadata. This provides producers, copyright holders, and content operators with freedom of choice without affecting the integrity of the transmitted content or sound quality.

[0008] To achieve the above objectives, the first aspect of this disclosure provides a method for generating audio packet format metadata, comprising:

[0009] According to the audio track type, the acquired audio file is converted into an audio data storage format and stored in a digital audio workstation. The audio data is then read from the digital audio workstation and input into the renderer input terminal. The acquired audio file contains audio track data or audio samples.

[0010] The audio file is identified based on its audio track format or audio sample type.

[0011] Encoded audio track format or audio sample type of the audio data for audio track identification;

[0012] If the audio data is unencoded, channel identification is performed based on the audio track format of the input audio data or the generated audio stream format metadata;

[0013] If an audio file or audio stream format is identified as a set of channels, the renderer performs identification and parsing processing on multiple audio stream tracks or metadata of the non-encoded audio stream format to generate audio package format metadata for creating an audio definition model.

[0014] To achieve the above objectives, a second aspect of this disclosure provides an audio packet format metadata generation apparatus, comprising:

[0015] The acquisition module is used to convert the acquired audio file into an audio data storage format according to the audio track type and store it in a digital audio workstation, and read the audio data from the digital audio workstation and input it into the renderer input terminal; wherein, the acquired audio file contains audio track data or audio samples;

[0016] The audio track format recognition module is used to identify the audio track format of the audio file based on the audio track format or audio sample type of the audio data.

[0017] The encoded audio track recognition module is used to identify the encoded audio track format or audio sample type of audio data.

[0018] The audio channel recognition module is used to perform audio channel recognition based on the audio track format of the input audio data or the generated audio stream format metadata if the audio data is unencoded audio data.

[0019] The generation module is used to identify and parse multiple audio stream tracks or metadata of non-encoded audio stream formats if the audio file or audio stream format is identified as a set of channels, and generate audio package format metadata for creating an audio definition model.

[0020] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, including: a memory and one or more processors;

[0021] The memory is used to store one or more programs;

[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio package format metadata generation method provided in any embodiment.

[0023] To achieve the above objectives, the fourth aspect of this disclosure provides a storage medium containing computer-executable instructions that, when executed by a computer processor, implement the audio package format metadata generation method provided in any embodiment.

[0024] As can be seen from the above, the audio packet format metadata generation method disclosed herein has the following technical effects:

[0025] The metadata of the audio package format describes the audio package format of 5 audio types (such as sound bed type, matrix type, object type, scene type and binaural type). The audio type of the audio package format matches the type defined in the referenced audio channel format set. It allows references to other audio package formats and allows nested audio package formats. It supports more than 5 types of audio package formats. The nested reference mechanism and structured attribute design support unlimited expansion from simple mono to complex scene audio, improves project reusability and reduces configuration time for complex projects.

[0026] By using dynamic mono / multi-channel recognition, the accuracy of channel resolution is improved, and rendering errors caused by channel misjudgment are eliminated.

[0027] The metadata is categorized according to the number of audio channel formats. It can be a metadata packet consisting of a single audio channel format or a metadata packet grouped together from multiple audio channel formats. The metadata packet includes an attribute area and a sub-element area. The renderer uses standardized metadata to achieve "read and render instantly", which improves efficiency.

[0028] It solves the long-standing problems in the field of audio production, such as format heterogeneity, processing complexity, and scalability limitations. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating the method for generating audio packet format metadata in Embodiment 1 of this disclosure;

[0030] Figure 2 This is a schematic diagram of the framework for the audio packet format metadata generation method provided in Embodiment 1 of this disclosure;

[0031] Figure 3 This is a structural diagram of the audio packet format metadata generation apparatus provided in Embodiment 2 of this disclosure;

[0032] Figure 4 This is a structural diagram of the audio packet format metadata module of the audio packet format metadata generation apparatus described in this disclosure;

[0033] Figure 5 This is a structural diagram of the audio packet format metadata module type of the audio packet format metadata generation device described in this disclosure;

[0034] Figure 6 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this disclosure;

[0035] Figure 7 This is a schematic diagram illustrating the relationship between the encoding matrix and the decoding matrix audio packet format set, which provide matrix-type audio packet format metadata in Embodiment 1 of this disclosure;

[0036] Figure 8 This is a schematic diagram illustrating the relationship between the direct matrix audio packet format and the matrix type audio packet format metadata provided in Embodiment 1 of this disclosure. Detailed Implementation

[0037] The following examples are used to illustrate this disclosure, but are not intended to limit the scope of this disclosure.

[0038] Metadata is information that describes the structural characteristics of data, and the functions supported by metadata include indicating storage location, historical data, resource lookup, or file records.

[0039] The metadata of a 3D audio production model consists of a set of production elements. Each production element describes the structural characteristics of the data at the corresponding stage of audio production through metadata. The 3D audio production model includes content production and format production.

[0040] The elements for content creation include: audio program elements, audio content elements, audio object elements, and unique track identifier elements.

[0041] The elements for format creation include: audio package format elements, audio channel format elements, audio stream format elements, and audio track format elements.

[0042] Metadata describing the characteristics of each stage of the 3D audio-visual production model is generated.

[0043] The audio channel data created based on the above-mentioned 3D audio production model is transmitted to the remote end via communication. The remote end then renders the audio channel data in stages based on metadata to restore the produced sound scene.

[0044] Example 1

[0045] This disclosure provides metadata for an audio packet format in a three-dimensional acoustic audio model, specifically as follows: Figure 1 and Figure 7 As shown, and explained in detail.

[0046] In the metadata of a 3D audio production model, the audio packet format element divides the metadata of the audio object and the audio stream data into multiple data blocks based on channels; these data blocks are called audio packets. These audio packets are transmitted along different paths in one or more networks for recombination at their destination. Embodiments of this disclosure use audio packet format metadata to describe the structural information of the audio packet format. Figure 1 This is a flowchart of the method for generating metadata for the audio package format disclosed herein.

[0047] like Figure 2 As shown, the method for generating metadata for this audio packet format includes the following steps:

[0048] Step S11: According to the audio track type, the acquired audio file is converted into an audio data storage format and stored in a digital audio workstation. The audio data is then read from the digital audio workstation and input into the renderer input terminal. The acquired audio file contains audio track data or audio samples.

[0049] Step S12: Identify the audio track format of the audio file according to the audio track format or audio sample type of the audio data; wherein, the format feature extraction includes: identifying the audio track encapsulation format based on the audio file header information, and determining the audio sample encoding type through a spectrum analysis algorithm, wherein the encoding type includes at least one of PCM, MPEG-H, and Dolby Digital.

[0050] Step S13: Encode track identification for the audio track format or audio sample type of the audio data;

[0051] Step S14: If the audio data is unencoded audio data, channel identification is performed based on the audio track format of the input audio data or the generated audio stream format metadata.

[0052] Step S15: If the audio file or audio stream format is identified as a group of channels, the renderer performs identification and parsing processing on multiple audio stream tracks or metadata of the non-encoded audio stream format to generate audio package format metadata (audioPackFormat) for creating the audio definition model.

[0053] This digital audio workstation (DAW) is used to sample analog audio signals from input audio files, converting them into digital sound files that can be read by a computer. The computer then performs various processing steps on the sound to complete the audio processing. The DAW includes a renderer. The digital audio software automatically identifies the number of channels in the input audio file or stream. Since the audio file or stream has a fixed format, the output audio definition model metadata, combined with PCM audio data, is input into the renderer. After rendering, the renderer outputs a secondary audio definition model metadata and the actual PCM audio data, arranged for real-time playback on the sound terminal. The audio file format refers to the file format in which audio data is stored.

[0054] Optionally, in step S11, reading the audio file data from the digital audio workstation and inputting it into the renderer input includes: uncompressed PCM audio encoding format or compressed audio encoding format; if it is uncompressed PCM audio encoding format, generating audio track metadata and audio packet metadata; when the audio data is in a compressed audio encoding format (such as AAC, MP3), generating audio stream metadata. There are two types of audio data: raw audio data and encoded audio data.

[0055] Optionally, in step S13, the encoding track identification of the audio data track format or audio sample type includes: determining whether the audio data of the audio stream track formed by the track format and the encoded audio track can be decoded according to the track format of the input audio file, and performing identification processing on the audio stream track to generate metadata.

[0056] Optionally, in step S13, the audio track encoding identification of the audio data further includes: if the audio data is in a non-PCM format, calling an external decoder to decode the audio data before sending it to the renderer for processing. Renderer: The process of generating metadata and sending audio signals to traditional channels defined by a specific speaker layout, ultimately transmitting them to the speakers, is called rendering; the processor is called the renderer.

[0057] Optionally, in step S14, the channel identification based on the audio track format of the input audio data or the generated audio stream format metadata includes: performing mono identification on the unprocessed audio file with the encoded audio track; or performing mono identification on the decoded audio stream format metadata. Before identifying channels, it is also necessary to identify the audio tracks in the audio file. Identification includes recognizing the format of the audio track and whether the track is encoded. It should be noted that an audio track refers to multiple parallel tracks presented in DAW software. These multiple parallel tracks define their respective attributes, such as: number of channels, volume, sampling rate, bit rate, etc. Specifically, when an unencoded audio stream is identified, multi-dimensional channel identification is initiated to parse the channel layout descriptor of the audio frame, match the channel definition rules in the SMPTE ST 2116 standard, and generate metadata identifiers containing channel topology relationships.

[0058] Optionally, in step S15, after generating the audio pack format metadata for creating the audio definition model, the method further includes: identifying the audio type of the audio pack format metadata attributes.

[0059] Optionally, the audio type identification of the audio package format metadata attributes includes: analyzing and obtaining the channel type of the audio file or audio stream, and determining a preset audio type definition match as one of the following: audio bed type audio package, matrix type audio package, object type audio package, scene type audio package, and binaural type audio package; the audio type of the audio package format matches the type defined in the referenced audio channel format set; the audio channel data of each audio type is generated through the above-mentioned three-dimensional audio production model.

[0060] Optionally, the audio packet format metadata generated for creating the audio definition model includes: audio packet format metadata composed of one audio channel format, classified according to the number of audio channel formats, or audio packet format metadata grouped together by multiple audio channel formats (audioChannelFormat).

[0061] Specifically, during the renderer processing stage, dynamic track mapping is implemented on the audio stream with metadata identifiers. The dynamic track mapping includes: generating an audio object description set based on the channel topology relationship, and constructing an audio packet format metadata tree structure based on the spatial coordinate parameters of the audio objects in the three-dimensional sound field.

[0062] Optionally, the audio packet format is the format used when grouping and packaging audio objects and raw audio data according to channels, including: references to other audio packet formats and allowing nested audio packet formats. Examples of audio packet format sets include stereo and 5.1 channel formats for channel-based formats.

[0063] Optional, such as Figure 4 As shown, the audio packet format metadata 100 includes an attribute area 110 and a sub-element area 120.

[0064] The attribute area 110 includes an audio package format identifier 111 and an audio package format name 112. This ensures the uniqueness of the audio package format identifier and provides cross-referencing information between production elements. This reduces information storage and improves data processing efficiency.

[0065] For different audio types and different channel types, the general attributes of the generated audio packet format metadata conform to the specifications in Table 1.

[0066] Table 1

[0067]

[0068] The importance parameter of an audio pack format allows for prioritizing compromises made when reducing metadata size. When using an importance parameter in an audio pack format, it's used to reduce spatial audio quality. Nested audio pack formats also use this feature to reduce spatial audio quality. For example, an audio object with a main direct sound and an additional reverb sound might discard the reverb sound, thus preserving the main sound but at a lower quality. The main direct sound is contained within a parent audio pack format element with high importance. The reverb sound is contained within a child audio pack format with low importance.

[0069] The sub-element area 120 in the audioPackFormat depends on the type definition or type label of the audioPackFormat element, including: first reference information 121, second reference information 122, and absolute distance 123.

[0070] The first reference information 121 includes audio channel format information used by the audio channel associated with the audio package during rendering.

[0071] Because the audio played by the speaker of the sound bed type directly stimulates the listener's brain, creating the immersive or three-dimensional sound experience described in psychoacoustics, and without the spatial coupling effect of audio, the audio metadata generation does not require the generation of general attributes and sub-elements of the audio package format; instead, attributes or sub-elements are specifically defined for the sound bed type. The second reference information 122 includes the audio package format information used by the audio package associated with the audio package during rendering. Similarly, the absolute distance 123 indicates a preset invalid value, which indicates that the audio package of the sound bed type does not have a corresponding distance during rendering. For example, the preset invalid value is zero.

[0072] For different audio types and channel types, the sub-elements of the generated audio packet format metadata conform to the specifications in Table 2.

[0073] Table 2

[0074]

[0075] Optionally, the attribute area 110 may also include a channel type label indicating the audio channel format or audio package format used for downward referencing during rendering.

[0076] If the next production element of an audio package format element is an audio channel format element, the attribute area 110 of the metadata of this audio package format includes a channel type tag indicating that the audio channel format referenced downwards during rendering uses an audio channel. If the next production element of an audio package format element is also an audio package format element (i.e., the produced audio package format includes nested audio package formats), the attribute area 110 of the metadata of this audio package format includes a channel type tag indicating that the audio package format referenced downwards during rendering uses an audio channel.

[0077] Optionally, the attribute area 110 may also include an audio type indicating the audio object or audio package format referenced upwards during rendering, using an audio channel.

[0078] Optionally, the attribute area 110 also includes information indicating the importance of the metadata 100 of the audio package format in rendering. Based on the importance information, the metadata 100 of audio package formats with high importance can be rendered first, and even the metadata 100 of audio package formats with low importance can be discarded as needed, thereby adapting to the requirements of the rendering schedule.

[0079] The audio type definition will generate five different typeDefinition metadata that conform to the specifications in Table 3:

[0080] Table 3

[0081]

[0082] The metadata attribute list for the audiopacket format (audioPackFormat) type is the attribute list for the audiopacket format. The child element list for the audiopacket format (audioPackFormat) type metadata is the child element for the audiopacket format.

[0083] After identifying the audio type from the audio package format metadata attributes, the audio package format metadata for the audio bed type can be generated using the audio package format metadata generation method provided above.

[0084] The Matrix type audiopackFormat defines sub-elements in addition to general sub-elements: encoding (e.g., from left / right to center / side), decoding (e.g., from center / side to left / right), and direct (e.g., Lo / Ro) matrices. That is, the matrix includes encoding matrices, decoding matrices, or direct matrices, as detailed below. Figure 8 As shown.

[0085] In a 3D audio production model, the upward-referencing production element can be understood as the preceding production element of the audio package format element. If the preceding production element of the audio package format element is an audio object element, the metadata attribute area of ​​this audio package format includes an indication of the audio channel audio type used by the audio object referenced upward by this audio package format during rendering. If the preceding production element of the audio package format element is also an audio package format element (i.e., the produced audio package format includes nested audio package formats), the metadata attribute area of ​​this audio package format includes an indication of the audio channel audio type used by the audio package format referenced upward by this audio package format during rendering.

[0086] A downreferenced production element can be understood as the production element following the audio package format element. If the production element following the audio package format element is an audio channel format element, the metadata attribute area of ​​this audio package format includes a channel type tag indicating the audio channel used by the downreferenced audio channel format during rendering. If the production element following the audio package format element is also an audio package format element (i.e., the produced audio package format includes nested audio package formats), the metadata attribute area of ​​this audio package format includes a channel type tag indicating the audio channel used by the downreferenced audio package format during rendering.

[0087] Based on importance information, metadata for audio package formats with high importance can be rendered first, and metadata for audio package formats with low importance can be discarded as needed, thereby adapting to the requirements of the rendering schedule.

[0088] After identifying the audio type from the audio packet format metadata attributes, the method for generating audio packet format metadata for matrix-type audio packet formats, based on the audio packet format metadata generation method provided above, also includes:

[0089] Generate sub-elements of the metadata definition for the audio pack format (audioPackFormat) of the matrix type. These sub-elements are sub-elements of the audio pack format that conform to the specifications in Table 4 and also include the absolute distance used in the matrix mode.

[0090] Table 4

[0091]

[0092] The Objects type audio package format is used for object-based audio. The metadata definition of this object type audio package format defines a list of attributes for the audio package format.

[0093] After identifying the audio type from the audio package format metadata attributes, the method for generating audio package format metadata for object types, based on the audio package format metadata generation method provided above, also includes:

[0094] Generates a sub-element of the metadata definition for the audio package format of the object type. Its sub-element is the sub-element of the audio package format, including: first reference information, second reference information, and absolute distance. The absolute distance is used in the object mode in the form of simulated or virtual distance: that is, the distance from the simulated center point of the perceptible effect object coupled by the sound in space to the origin of the preset coordinates.

[0095] The SCENE type audio pack format is used for scene-based audio (such as higher-order ambient sounds). A set of scene components / signals will share the same normalization, NFC compensation, and / or screen relationships. When parameters are specified in the audio block format, these values ​​will override the values ​​given in the audio pack format. The list of attributes defined in the scene type audio pack format metadata is the attribute list for the audio pack format.

[0096] After identifying the audio type from the audio packet format metadata attributes, the method for generating audio packet format metadata for scene types, based on the audio packet format metadata generation method provided above, also includes:

[0097] The generation of scene type audio pack format metadata definition sub-elements, whose sub-elements include: first reference information, second reference information, absolute distance and scene component description information, wherein the scene component description information includes: normalization information, reference distance information and screen related information.

[0098] The sub-elements of the audio package format metadata definition for scene types conform to the specifications in Table 5:

[0099] Table 5

[0100]

[0101] After identifying the audio type from the audio packet format metadata attributes, for binaural audio packet format metadata, based on the audio packet format metadata generation method provided above, the binaural audio packet format metadata generation method also includes:

[0102] Generates a sub-element defining the metadata of the audio pack format for binaural audio, whose sub-elements do not contain the absoluteDistance and audioPackFormatIDRef sub-elements.

[0103] The Binaural audio pack format is used for binaural audio. The `absoluteDistance` and `audioPackFormatIDRef` child elements cannot be used in binaural mode. The list of attributes defined for this Binaural audio pack format is the same as the attributes of the audio pack format. The child elements defined for the Binaural audio pack format are child elements of the audio pack format; no other child elements are defined for the Binaural audio pack format (`audioPackFormat`).

[0104] This disclosure describes five types of audio package formats by generating metadata for the audio package formats. Through audio track format recognition and encoded audio track recognition processes, it achieves unified conversion and parsing of audio track data and audio samples, improving the compatibility of heterogeneous audio. Through the structured design of attribute area (format identifier, name) and sub-element area (type label), it achieves classified management of multiple types of audio, supports flexible expansion of complex audio projects, and enables the reproduction of three-dimensional sound in the channel, thereby improving the quality of the sound scene.

[0105] Example 2

[0106] This disclosure also provides method embodiments that follow the above embodiments, a method for generating metadata for audio package formats. The interpretation of the same names is the same as that of the above embodiments, and the same technical effects are achieved. Therefore, it will not be described again here.

[0107] like Figure 3 As shown, an audio packet format metadata generation device includes:

[0108] The acquisition module 21 is used to convert the acquired audio file into an audio data storage format according to the audio track type and store it in the digital audio workstation, and input the audio data read from the digital audio workstation into the renderer input terminal; wherein, the acquired audio file contains audio track data or audio samples;

[0109] The audio track format recognition module 22 is used to recognize the audio track format of the audio file according to the audio track format or audio sample type of the audio data;

[0110] The encoded audio track recognition module 23 is used to identify the encoded audio track format or audio sample type of the audio data.

[0111] The channel recognition module 24 is used to perform channel recognition based on the audio track format of the input audio data or the generated audio stream format metadata if the audio data is unencoded audio data.

[0112] The generation module 25 is used to generate an audio package format metadata module for creating an audio definition model if the audio file or audio stream format is identified as a group of channels. The renderer identifies and parses multiple audio stream tracks or metadata of the non-encoded audio stream format.

[0113] Optionally, in step S24, the channel identification based on the audio track format of the input audio data or the generated audio stream format metadata includes: performing mono identification on the unprocessed audio file of the encoded audio track; or performing mono identification on the decoded audio stream format metadata.

[0114] Optional, such as Figure 5 As shown, after generating the audio package format metadata module 26 for creating the audio definition model, it also includes: an audio type identification module for the audio package format metadata attributes.

[0115] Optionally, the audio type identification module for the audio package format metadata attributes includes: analyzing and obtaining the channel type of the audio file or audio stream, and obtaining a preset audio type definition to match and determine one of the following: audio bed type audio package, matrix type audio package, object type audio package, scene type audio package, and binaural type audio package; the audio type of the audio package format matches the type defined in the referenced audio channel format set, and the audio channel data of each audio type is generated through the above-mentioned three-dimensional audio production model.

[0116] Optionally, the audio packet format metadata module for generating the audio definition model includes: audio packet format metadata classified according to the number of audio channel formats, consisting of one audio channel format, or audio packet format metadata grouped together by multiple audio channel formats (audioChannelFormat).

[0117] Optionally, the audio packet format is the format used when grouping and packaging audio objects and raw audio data according to channels, including: references to other audio packet formats and allowing nested audio packet formats. Examples of audio packet format sets include stereo and 5.1 channel formats for channel-based formats.

[0118] Optional, such as Figure 4 As shown, the audio package format metadata 100 includes: an attribute area 110 and a sub-element area 120. The attribute area includes the audio package format identifier 111 and the audio package format name 112. The sub-element area depends on the type definition or type label of the audio package format element and includes: first reference information 121, second reference information 122 and absolute distance 123.

[0119] This disclosure embodiment achieves unified conversion and parsing of audio track data and audio samples through audio track format recognition and encoded audio track recognition processes. Through mono / multi-channel recognition rules, it dynamically parses the channel type of unencoded audio to ensure the accuracy of channel information. It allows audio package format to reference or nest other formats and supports flexible expansion of complex audio projects, such as multi-format layered production and cross-project metadata reuse.

[0120] Example 3

[0121] Figure 6 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this disclosure. Figure 6 As shown, the electronic device includes a processor 30, a memory 31, an input device 32, and an output device 33. The electronic device may have one or more processors 30. Figure 6 Taking a processor 30 as an example. The electronic device may contain one or more memory units 31. Figure 6 Taking a memory 31 as an example, the processor 30, memory 31, input device 32, and output device 33 of this electronic device can be connected via a bus or other means. Figure 6 Taking a bus connection as an example, the electronic device can be a computer or a server. This disclosure describes the embodiment using an electronic device as a server, which can be a standalone server or a cluster server.

[0122] The memory 31, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules of the audio package format metadata generation method described in any embodiment of this disclosure. The memory 31 may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 31 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 31 may further include memory remotely located relative to the processor 30, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0123] Input device 32 can be used to receive input digital or character information, and to generate key signal inputs related to viewer user settings and function control of electronic devices. It can also be a camera for acquiring images and a sound pickup device for acquiring audio data. Output device 33 may include audio devices such as speakers. It should be noted that the specific composition of input device 32 and output device 33 can be set according to actual conditions.

[0124] The processor 30 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 31, thereby generating audio package format metadata.

[0125] Example 4

[0126] Embodiment 4 of this disclosure also provides a storage medium containing computer-executable instructions, which are implemented by a computer processor using the audio package format metadata generation method provided in Embodiment 1.

[0127] Of course, the computer-executable instructions provided in the embodiments of this disclosure are not limited to the electronic method operations described above, but can also perform related operations in the electronic methods provided in any embodiment of this disclosure, and have corresponding functions and beneficial effects.

[0128] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this disclosure can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a robot, personal computer, server, or network device, etc.) to execute the electronic methods described in any embodiment of this disclosure.

[0129] It is worth noting that the various units and modules included in the above electronic device are only divided according to functional logic, but are not limited to the above division, as long as they can achieve the corresponding functions; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this disclosure.

[0130] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0131] In the description of this specification, references to terms such as "in one embodiment," "in yet another embodiment," "exemplary," or "in a particular embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0132] Although this disclosure has been described in detail above with general descriptions, specific embodiments, and experiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, such modifications or improvements made without departing from the spirit of this disclosure are all within the scope of protection claimed by this disclosure.

Claims

1. A method for generating audio packet format metadata, characterized in that, include: According to the audio track type, the acquired audio file is converted into an audio data storage format in a digital audio workstation, and the audio data is read from the digital audio workstation and input into the renderer input terminal; wherein, the acquired audio file contains audio track data or audio samples; The audio file is identified based on its audio track format or audio sample type. Encoded audio track format or audio sample type of the audio data for audio track identification; If the audio data is unencoded, channel identification is performed based on the audio track format of the input audio data or the generated audio stream format metadata; If the audio file or audio stream format is identified as a set of channels, the renderer performs identification and parsing processing on multiple audio stream tracks or metadata of the non-encoded audio stream format to generate audio package format metadata for creating an audio definition model; Audio type identification is performed on the metadata attributes of the audio packet format. The channel type of the obtained audio file or audio stream is analyzed, and the preset audio type definition is matched and determined to be one of the following: audio bed type audio package, matrix type audio package, object type audio package, scene type audio package, and binaural type audio package; the audio type of the audio package format matches the type defined in the referenced audio channel format set; The audio package format metadata includes an attribute area and a sub-element area. The attribute area includes the audio package format identifier and the audio package format name. The sub-element area depends on the type definition or type tag of the audio package format element. If the audio package format metadata is a matrix type audio package, the audio package format identifier includes information indicating that the audio type of the audio package is matrix type.

2. The method according to claim 1, characterized in that, The step of performing channel identification based on the audio track format of the input audio data or the generated audio stream format metadata includes: performing mono identification on an unprocessed audio file with an encoded audio track; Alternatively, mono recognition can be performed on the decoded audio stream format metadata.

3. The method according to claim 1, characterized in that, The audio package format metadata for generating the audio definition model includes: audio package format metadata composed of one audio channel format, or audio package format metadata grouped together by multiple audio channel formats belonging to each other, classified according to the number of audio channel formats.

4. The method according to claim 1, characterized in that, The audio package format is the format used when grouping and packaging audio objects and raw audio data according to channels, including: references to other audio package formats and allowing nested audio package formats.

5. The method according to claim 1, characterized in that, The sub-element region includes: first reference information, second reference information, absolute distance, and matrix reference information.

6. The method according to claim 5, characterized in that, The first reference information includes audio channel format information used by the audio channel associated with the audio package during rendering; the second reference information includes audio package format information used by the audio package associated with the audio package during rendering; the absolute distance indicator is a preset invalid value; the matrix reference information includes referencing the encoding matrix audio package format from the decoding matrix, referencing the decoding matrix audio package format from the encoding matrix, referencing the channel-based input audio package format, and referencing the channel-based matrix decoded audio package format.

7. An audio packet format metadata generation device, characterized in that, include: The acquisition module is used to convert the acquired audio file into an audio data of a preset format according to the audio track type and store it in a digital audio workstation, and read the audio data from the digital audio workstation and input it into the renderer input terminal; wherein, the acquired audio file contains audio track data or audio samples; The audio track format recognition module is used to identify the audio track format of the audio file based on the audio track format or audio sample type of the audio data. The encoded audio track recognition module is used to identify the encoded audio track format or audio sample type of the audio data. The audio channel recognition module is used to perform audio channel recognition based on the audio track format of the input audio data or the generated audio stream format metadata if the audio data is unencoded audio data. The generation module is used to generate audio package format metadata for creating an audio definition model if the audio file or audio stream format is identified as a set of channels. An audio type identification module that analyzes audio packet format metadata attributes; The channel type of the obtained audio file or audio stream is analyzed, and the preset audio type definition is matched and determined to be one of the following: audio bed type audio package, matrix type audio package, object type audio package, scene type audio package, and binaural type audio package; the audio type of the audio package format matches the type defined in the reference audio channel format set, and the audio channel data of each audio type is generated through the above three-dimensional audio production model; The audio package format metadata includes an attribute area and a sub-element area. The attribute area includes the audio package format identifier and the audio package format name. The sub-element area depends on the type definition or type tag of the audio package format element. If the audio package format metadata is a matrix type audio package, the audio package format identifier includes information indicating that the audio type of the audio package is matrix type.

8. An electronic device, characterized in that, include: Memory and one or more processors; The memory is used to store one or more programs; The program is adapted to be loaded and run by the processor to perform the audio package format metadata generation method according to any one of claims 1 to 6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are adapted to be loaded and run by the processor to perform the audio package format metadata generation method according to any one of claims 1 to 6.