Audio channel metadata generation method and device, equipment and storage medium
By automatically generating audio channel metadata, the problem of inefficiency in manual definition in the existing technology is solved, and adaptive rendering and immersive experience of multi-channel audio files are realized. It is suitable for multi-scene audio rendering such as car audio, VR equipment and theater systems.
Patent Information
- Application Number
- CN202510443663.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, audio channel metadata generation needs to be manually defined, resulting in inefficient and difficult to adapt to dynamically changing audio rendering requirements.
By confirming the audio file type, the audio channel format elements are generated, divided into multiple audio blocks, and common attributes and type attributes are identified to automatically generate metadata, including information such as frequency, audio type, and start and end time.
It realizes the automatic construction of audio channel format metadata, supports adaptive rendering in mono, 5.1 channel, 7.1 channel and binaural audio formats, improving the immersive experience of multi-channel audio, and adapts to the audio rendering needs of multi-scene such as car audio, VR equipment and theater systems.
Smart Images

Figure CN120544596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a method, apparatus, device and storage medium for generating audio channel metadata. Background Art
[0002] Audio technology has evolved through multiple stages, from mono to left and right channels to surround sound. The advent of surround sound, in particular, has made the processing of sound and the audio channels it belongs to increasingly complex. Sequencing multiple audio channels can enhance the interaction between speakers within these channels, improving the interactivity and immersion of the sound output.
[0003] Audio channels refer to the source or output of sound, and are typically used to describe the spatial location of sound. The number of channels corresponds to the number of sound sources during recording or the number of speakers during playback. For example, a 5.1 surround speaker system has six different audio signals, while a 7.1 surround speaker system has eight different audio signals. These different audio signals require different audio channels for output.
[0004] Audio comes in various types, such as direct speaker, matrix, object, high-order ambient, and binaural. These different audio channel types require different tagging operations, and these tags are metadata. Therefore, a method for generating audio channel metadata is necessary to meet the specific requirements of audio file rendering and playback. Existing metadata generation requires manual definition of audio channel types, resulting in inefficiency and difficulty adapting to dynamically changing audio rendering requirements. Summary of the Invention
[0005] The present invention provides an audio channel metadata generation method, device, equipment and storage medium to meet the rendering and playback needs of different types of audio files.
[0006] In a first aspect, the present application provides a method for generating audio channel metadata, which is applied to an audio renderer, comprising:
[0007] Confirm whether the received audio file is mono;
[0008] If so, generate the audio channel format element corresponding to the audio file;
[0009] Divide the audio channel into multiple audio blocks according to the audio channel format element;
[0010] Identify common attributes and type attributes of audio blocks;
[0011] Generate metadata corresponding to common attributes and metadata for type attributes.
[0012] Optionally, before confirming whether the received audio file is mono, the method further includes: identifying an audio track in the audio file.
[0013] Optionally, the audio channel format element further includes frequency, audio type, and start and end time.
[0014] Optionally, after generating the corresponding audio channel format element of the audio file, the method further includes:
[0015] Identify frequencies and generate metadata of cutoff frequencies;
[0016] Identify audio types and generate type attribute metadata;
[0017] According to the start and end times, the audio blocks in the audio channel and the duration metadata corresponding to the audio blocks are determined.
[0018] Optionally, common properties include gain, importance, head locking, and headphone virtualization.
[0019] Optionally, the type attribute of the audio block includes bed, matrix, object, scene, and binaural.
[0020] Optionally, generate metadata for type attributes, including:
[0021] Generate channel audio block metadata about the sound bed, including speaker position labels, sound horizontal angles, sound pitch angles, and normalized distances from the origin;
[0022] Generates matrix audio block metadata about the matrix, including channel output format, matrix sub-element gains, variable gain, phase, and variable phase;
[0023] Generate object audio block metadata about the object, including position coordinates, channel lock, and object divergence;
[0024] Generates scene audio block metadata about the scene, including decoding equations, order, degree, and normalization format.
[0025] A second aspect of the present application provides an audio metadata conversion device, which applies the method provided in the first aspect, including:
[0026] A mono confirmation module is used to confirm whether the received audio file is mono;
[0027] An audio channel format element generation module, used to generate an audio channel format element corresponding to an audio file;
[0028] An audio block division module, configured to divide the audio channel into multiple audio blocks according to the audio channel format element;
[0029] An attribute recognition module, used to recognize common attributes and type attributes of audio blocks;
[0030] The metadata generation module is used to generate metadata corresponding to common attributes and metadata corresponding to type attributes.
[0031] A third aspect of the present application provides an electronic device, comprising: a memory and one or more processors;
[0032] a memory for storing one or more programs;
[0033] When one or more programs are executed by one or more processors, the one or more processors implement the audio channel metadata generation method provided in any embodiment.
[0034] A fourth aspect of the present application provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used by a computer processor to implement the audio channel metadata generation method provided in any embodiment.
[0035] As can be seen from the above, this application provides a method for generating audio channel metadata. By determining whether the loaded audio file is a mono file, the audio channel format elements are automatically constructed to reduce dependence on manual configuration; the constructed audio channel format metadata also includes the frequency, audio type and start and end time corresponding to the audio file.
[0036] Generates metadata for common attributes (gain, importance, head lock) and type attributes (sound bed, matrix, object) using audio blocks as units, enabling adaptive rendering of multiple audio file types and compatible with mono, 5.1, 7.1, and binaural audio formats.
[0037] Furthermore, based on the audio type, spatial metadata such as speaker position labels, sound level angles, and origin normalized distance are generated for audio blocks of different types, such as sound beds and objects, enhancing the immersive experience of multi-channel audio.
[0038] The channel processing order can be optimized according to real-time rendering requirements to adapt to the audio rendering needs of multiple scenarios such as car audio, VR equipment, and cinema systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A flowchart of a method for generating audio channel metadata provided by an embodiment of the present invention;
[0040] Figure 2 A flowchart for processing audio channel format elements in a method for generating audio channel metadata provided by an embodiment of the present invention;
[0041] Figure 3A flowchart of metadata generation for type attributes in a method for generating metadata for an audio channel provided by an embodiment of the present invention;
[0042] Figure 4 A schematic structural diagram of an audio channel metadata generation device provided by an embodiment of the present invention;
[0043] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0045] Example 1
[0046] A method for generating audio channel metadata, applied to audio renderers, such as Figures 1 to 3 Shown, including:
[0047] S10: Confirm whether the received audio file is mono. Before confirming whether it is mono, step S00 needs to be executed: identify the audio track in the audio file. Identification includes identifying the format of the audio track and whether the audio track is an encoded audio track. It should be added here that audio tracks refer to multiple parallel tracks presented in a digital audio workstation (DAW). The above multiple parallel audio tracks respectively define the properties of their respective audio tracks, such as: number of channels, volume, sampling rate, bit rate, etc.
[0048] The purpose of identification is to obtain the properties of the audio track and generate metadata of the audio track, which is used for later processing.
[0049] As for identifying whether it is an encoded audio track, if it is, subsequent attribute metadata about the audio stream will be generated and decoding will be required; if it is not, a mono judgment will be made.
[0050] An audio file may contain multiple mono files or a multi-channel file. If the multi-channel file is transmitted in a multi-channel manner, it will be identified in the form of an audio package.
[0051] S20: If yes, generate the corresponding audio channel format elements of the audio file; the audio channel can be constructed through the above elements; the above audio channel format elements also include frequency, audio type, start and end time, and the recognition process does not have any order requirements.
[0052] The process of the renderer identifying the audio channel format elements is also the process of identifying the audio channel.
[0053] Specifically, when the renderer identifies the frequency element, it generates metadata of the cutoff frequency accordingly. The cutoff frequency includes a high cutoff frequency and a low cutoff frequency. The corresponding generated metadata also includes metadata of the high cutoff frequency and metadata of the low cutoff frequency.
[0054] When the renderer identifies the input audio type, it generates metadata for the corresponding type attributes. These metadata include, but are not limited to, bed, matrix, object, scene, and binaural. These type attributes define the type of audio channels and also define the type attributes of subsequent audio blocks.
[0055] When the renderer identifies the start and end time of the input audio block, it determines the metadata of the start time and duration of the audio block.
[0056] Specifically, the sampling rate is extracted from the audio file, the cutoff frequency is obtained according to the Nyquist sampling theorem, and the cutoff frequency value is recorded as metadata. Dynamic time warping (DTW) is used to align multi-frame frequency data to reduce noise interference, and the median of the frequency of 5 consecutive frames is taken to avoid instantaneous spikes.
[0057] Analyze the number of channels in the audio file, map it to a standard type based on the number of channels, use Mel spectrum analysis to distinguish speech from music, identify ambient sounds, instrument sounds, etc. through machine learning models (such as CNN), and record the channel layout and content type. Regularly update the classification model with new annotated data to adapt to new audio types (such as AI synthesized speech).
[0058] The sample positions corresponding to the start and end times are calculated based on the sampling rate of the audio file, and the waveform data of the corresponding sample interval is extracted from the original audio data to form an independent audio block. The duration of the independent audio block is calculated, and the duration metadata is obtained based on the duration of the independent audio block. Short blocks with intervals less than 200ms (such as coughs and interruptions) are merged to achieve millisecond-level audio block annotation and support interactive audio editing.
[0059] S30: Divide the audio channel into multiple audio blocks according to the audio channel format element; for example, when the type, cutoff frequency, or start and end time of the aforementioned audio channel are determined to be partially or entirely red, the audio channel may be divided into multiple audio blocks.
[0060] S40: Identify common attributes and type attributes of audio blocks. After the audio channel is divided into multiple audio blocks, it is necessary to identify the common attributes of the audio blocks and identify the special attributes of the audio blocks under different type attributes.
[0061] The audio block type attribute corresponds to the input audio type identification described above. Audio block type attributes include, but are not limited to, bed, matrix, object, scene, and binaural. Audio transmission also follows this correspondence. For example, scene-based audio types will also be transmitted through the audio blocks of the scene's audio channels.
[0062] S50: Generate metadata corresponding to the general attributes and metadata corresponding to the type attributes.
[0063] Common attributes include but are not limited to gain, importance, headLocked, and headphone Virtualize. The generated metadata is metadata corresponding to the above attributes.
[0064] As for generating metadata of type attributes, this process needs to go back to the process of the renderer generating the audio channel format audioChannelFormat element. In this process, the renderer obtains the metadata of the attributes of the audio channel format element by receiving the audio file: audio channel format name, audio channel identifier, audio channel type description information (audio channel type description information, including type tag and / or type definition), and as mentioned above, the audio channel format is divided into audio blocks (equivalent to sub-elements), and the obtained sub-element data includes the audio block format audioBlockFormat, frequency, and the audio block format attribute metadata includes the audio block format ID, start time, and duration;
[0065] The sub-elements in the audio block format are further read according to the aforementioned generated type definitions (sound bed, matrix, object, scene, and binaural), that is, the aforementioned five types are independently identified according to their respective audio block attributes.
[0066] For example:
[0067] S51: Generate channel audio block metadata about the sound bed, including the speaker position label speakerLabel, the sound horizontal angle azimuth, the sound pitch angle elevation, and the origin normalized distance distance;
[0068] It should be added here that the renderer reads the channel audio block metadata of the sound bed. The read content must include and can only include the above-mentioned speaker position label speakerLabel, sound horizontal angle azimuth, sound pitch angle elevation and origin normalized distance. Otherwise, the format of the channel audio block metadata of the sound bed is incorrect and the audio file cannot be played. In this process, the above-mentioned channel audio block metadata of the sound bed is stored in the renderer DAW cache during processing, and will be stored in the form of a document later.
[0069] S52: Generate matrix audio block metadata about the matrix, including channel output format outputChannelFormatDRef, matrix subelement gain MatrixGain, variable gain GainVar, phase Phase, and variable phase phaseVar;
[0070] S53: Generate object audio block metadata about the object, including position coordinates (including Cartesian and polar coordinates), channel lock channelLook, and object divergence objectDivergence.
[0071] S54: Generate scene audio block metadata about the scene, including decoding equations, orders, degrees, and normalized formats.
[0072] Among them, steps S51-S54 can be executed in parallel or selectively, specifically according to the type attributes of the audio block. And the reading, error reporting and storage of elements in the metadata of different types of audio blocks in steps S52-step S54 are similar to step S51, and will not be repeated here. Therefore, before steps S51-S54, it also includes identifying the type of audio file; determining whether the type belongs to one of sound bed, matrix, object, scene and binaural, and then executing one of steps S51-S54 according to the type. If the type is sound bed, execute step S51; if the type is matrix, execute step S52; if the type is object, execute step S53; if the type is scene, execute step S54; if the type is binaural, the attribute area and sub-element area of the audio channel metadata do not need to set other information, wherein the audio channel name is used to indicate the ear side of the binaural audio corresponding to the audio channel, and the ear side of the binaural audio includes the left ear and the right ear.
[0073] It should be added here that if the type attribute of the identified audio block is binaural, the metadata generated is the same as the metadata generated by the general attribute, and also includes gain, importance, head locking, and headphone virtualization.
[0074] This application provides a method for generating audio channel metadata. By determining whether the loaded audio file is a mono file, if so, it generates the elements of the audio channel required for the audio file transmission, and generates metadata of common attributes and type attributes in the form of audio blocks. In this way, metadata about the audio channel is established, thereby meeting the rendering and playback requirements of different types of audio files. The renderer can automatically generate metadata for up to 128 channels of audio input and corresponding output metadata.
[0075] Example 2
[0076] Figure 4 A schematic structural diagram of an audio channel metadata generation device provided by an embodiment of the present invention includes:
[0077] Mono confirmation module 01, used to confirm whether the received audio file is mono;
[0078] The audio channel format element generation module 02 is used to generate the audio channel format element corresponding to the audio file;
[0079] An audio block division module 03 is configured to divide the audio channel into multiple audio blocks according to the audio channel format element;
[0080] Attribute recognition module 04, used to identify the general attributes and type attributes of the audio block;
[0081] The metadata generation module 05 is used to generate metadata corresponding to common attributes and metadata corresponding to type attributes.
[0082] The audio channel metadata generation device provided in this application applies the audio channel metadata generation method provided in the aforementioned embodiment, has the same technical concept and achieves the same technical effect, which will not be repeated here.
[0083] Example 3
[0084] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 5 As shown, the electronic device includes: a processor 510, a memory 520, an input device 530, and an output device 540. The number of processors 510 in the electronic device can be one or more. Figure 5 In the example, a processor 510 is used. The number of memories 520 in the electronic device can be one or more. Figure 5 A memory 520 is taken as an example. The processor 510, memory 520, input device 530 and output device 540 of the electronic device can be connected through a bus or other means. Figure 5The example of the bus connection is shown in FIG. The electronic device may be a computer or a server. The embodiment of the present application is described in detail with the electronic device being a server, which may be an independent server or a cluster server.
[0085] Memory 520, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules of the audio channel metadata generation device in any embodiment of the present application. Memory 520 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on device usage. Furthermore, memory 520 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, memory 520 may further include memory located remotely from processor 510, which may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0086] The input device 530 can be used to receive input digital or character information and generate key signals related to the user settings and function control of the electronic device. It can also be a camera for capturing images and a sound pickup device for capturing audio data. The output device 540 can include audio equipment such as speakers. It should be noted that the specific composition of the input device 530 and the output device 540 can be set according to actual circumstances.
[0087] The processor 510 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 520, that is, implements the audio channel metadata generation method.
[0088] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present application is not limited to the above-mentioned audio channel metadata generation method, and can also execute related operations of the audio channel metadata generation method provided in any embodiment of the present application, and has corresponding functions and beneficial effects.
[0089] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0090] It is worth noting that in the above-mentioned embodiment of the audio channel metadata generation device, the various units and modules included are divided only according to functional logic, but are not limited to the above-mentioned division, as long as they can achieve the corresponding functions; in addition, the specific names of the functional units are only for the convenience of distinguishing them from each other and are not intended to limit the scope of protection of the present invention.
[0091] Although the present invention has been described in detail above using general explanations, specific embodiments, and experiments, it will be apparent to those skilled in the art that modifications and improvements may be made based on the present invention. Therefore, such modifications and improvements, which do not depart from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A method for generating audio channel metadata, applied to an audio renderer, characterized in that: include: Confirm whether the received audio file is mono; If yes, generating an audio channel format element corresponding to the audio file; Divide the audio channel into a plurality of audio blocks according to the audio channel format element; Identifying general attributes and type attributes of the audio block; Metadata corresponding to the general attribute and metadata corresponding to the type attribute are generated.
2. The method for generating audio channel metadata according to claim 1, wherein: Before confirming whether the received audio file is mono, the method further includes: An audio track in the audio file is identified.
3. The method for generating audio channel metadata according to claim 1, wherein: The audio channel format element also includes frequency, audio type, and start and end time.
4. The method for generating audio channel metadata according to claim 3, wherein: After generating the audio channel format element corresponding to the audio file, the method further includes: identifying the frequencies of the audio file and generating metadata of the cutoff frequencies; Identify the audio type of the audio file and generate type attribute metadata; The audio block in the audio channel and duration metadata corresponding to the audio block are determined according to the start and end times.
5. The method for generating audio channel metadata according to claim 1, wherein: The common properties include gain, importance, head locking, and headphone virtualization.
6. The method for generating audio channel metadata according to claim 1, wherein: The type attributes of the audio block include bed, matrix, object, scene, and binaural.
7. The method for generating audio channel metadata according to claim 6, wherein: The generating of metadata of the type attribute includes: generating channel audio block metadata about the sound bed, including speaker position labels, sound horizontal angles, sound elevation angles, and origin normalized distances; generating matrix audio block metadata about the matrix, including channel output format, matrix subelement gain, variable gain, phase, and variable phase; generating object audio block metadata about the object, including position coordinates, channel lock, and object divergence; Scene audio block metadata about the scene is generated, including decoding equations, orders, degrees, and normalized formats.
8. An audio channel metadata generation device, applying the method according to any one of claims 1 to 7, characterized in that: include: A mono confirmation module is used to confirm whether the received audio file is mono; An audio channel format element generation module, configured to generate an audio channel format element corresponding to the audio file; an audio block division module, configured to divide the audio channel into a plurality of audio blocks according to the audio channel format element; An attribute identification module, configured to identify general attributes and type attributes of the audio block; The metadata generation module is used to generate metadata corresponding to the general attributes and metadata corresponding to the type attributes.
9. An electronic device, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A storage medium containing computer-executable instructions, characterized in that: The computer executable instructions are implemented by a computer processor to implement the method according to any one of claims 1 to 7.