Audio processing methods and systems based on three-dimensional premium sound

By standardizing the format conversion and processing the metadata of the Audio Vivid audio format, the problem of streaming audio content not being able to be directly reproduced in professional cinemas has been solved, achieving efficient format compatibility and audio quality fidelity, thus enhancing the viewing experience.

CN122093731APending Publication Date: 2026-05-26CHINA RES INST OF FILM SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RES INST OF FILM SCI & TECH
Filing Date
2026-01-29
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technology cannot directly convert Audio Vivid audio format to meet the standards of domestic immersive sound processing systems for digital cinemas, resulting in streaming audio content being unable to be reproduced in high quality in professional cinemas, causing resource waste and hindering the improvement of the viewing experience.

Method used

By parsing the Audio Vivid audio definition model file, performing standardized format conversion, fixed-duration block processing, interpolation calculation, coordinate system transformation, and coordinate encoding processing, the data is converted into standardized audio data and metadata that conform to the domestic immersive sound processing system for digital cinema.

Benefits of technology

It achieves seamless integration between the Audio Vivid audio format and the domestic immersive sound processing system for digital cinemas, ensuring audio fidelity and spatial reproduction effects, and meeting the high-quality sound reproduction needs of professional cinemas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093731A_ABST
    Figure CN122093731A_ABST
Patent Text Reader

Abstract

This application relates to the field of audio processing technology, and discloses an audio processing method and system based on 3D immersive sound. The method includes: acquiring an ADM audio definition model file exported from the Audio Vivid audio production tool, parsing it to obtain raw monotrack audio data and corresponding raw metadata; performing standardized format conversion and independent storage on the raw monotrack audio data to obtain standardized audio data; and performing fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing, and coordinate encoding processing on the raw metadata to obtain standardized metadata. This application solves the problem that the Audio Vivid audio format cannot be directly used in domestic immersive sound processing systems for digital cinemas, and achieves seamless integration between streaming media 3D audio content and professional cinema sound reproduction systems, ensuring audio fidelity and spatial reproduction effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, specifically to an audio processing method and system based on three-dimensional immersive sound. Background Technology

[0002] To promote the independent development and standardization of film audio technology in my country, the China Film Science and Technology Research Institute has developed a domestic digital cinema immersive sound processing system based on the SMPTE ST 2098-2 immersive audio bitstream specification standard. This system integrates standard-compliant immersive sound production, IAB encoding, decoding, and rendering functions, providing key technical support for the production and reproduction of immersive sound in Chinese films.

[0003] Meanwhile, in the broadcast television and streaming media fields, the Audio Vivid standard, developed by the Ultra HD Video Industry Alliance (UWA), is progressing rapidly. However, the AVS3-Audio encoding format is primarily designed for lossy compression and transmission scenarios with a limited number of channels (e.g., a maximum of 16 channels), making it difficult to meet the high-quality audio playback requirements of at least 32 channels in cinemas. Furthermore, existing cinema servers cannot directly decode and render the Audio Vivid format. This technical bottleneck severely limits its application in the film industry, preventing the direct integration of rich streaming 3D audio content into professional cinema sound systems, resulting in a waste of audio resources and hindering the improvement of the viewing experience.

[0004] Therefore, how to construct an efficient and reliable audio format conversion solution to adapt object-based Audio Vivid audio content to a domestic digital cinema immersive sound processing system that conforms to the SMPTE ST 2098-2 standard, and to open up the technical link from streaming media audio content to cinema immersive sound playback system, has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] To address the aforementioned issues, this application provides an audio processing method and system based on 3D immersive sound, which solves the problem that the Audio Vivid audio format cannot be directly used in domestic digital cinema immersive sound processing systems. It achieves seamless integration between streaming media 3D audio content and professional cinema sound systems, ensuring audio fidelity and spatial reproduction effects.

[0006] The embodiments of this application adopt the following technical solutions: Firstly, this application provides an audio processing method based on three-dimensional high-resolution audio, including: Obtain the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata; The original single-track audio data is subjected to standardized format conversion and independent storage to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema. The original metadata is processed by fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing and coordinate encoding processing to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

[0007] Secondly, this application also provides an audio processing system based on three-dimensional high-fidelity sound, comprising: The data acquisition module is used to acquire the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata. The audio data standardization module is used to perform standardized format conversion and independent storage on the original single-track audio data to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema. The metadata standardization module is used to perform fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing, and coordinate encoding processing on the original metadata to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

[0008] Thirdly, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described audio processing method based on three-dimensional immersive sound.

[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described audio processing method based on three-dimensional sound.

[0010] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: This application addresses the technical bottleneck and achieves format compatibility by resolving the issue that the Audio Vivid audio format cannot be directly used in domestic immersive sound processing systems for digital cinemas. Through a series of precise conversion and processing steps, the Audio Vivid ADM audio definition model file is converted into standardized audio data and metadata that conform to the domestic immersive sound processing system for digital cinemas (SMPTE ST 2098-2 standard). This establishes a technical link from streaming audio content to the cinema immersive sound playback system, achieving efficient compatibility between the two formats.

[0011] Ensuring audio quality and spatial reproduction: This application strictly adheres to relevant standards during the conversion process, employing scientifically sound interpolation calculations, coordinate system transformations, and coordinate encoding methods. Experimental verification shows that the converted audio achieves excellent levels in terms of sound quality fidelity, static object attribute overlap, and dynamic object change overlap, fully preserving the original audio's sound quality characteristics and spatial information, providing viewers with a high-quality, immersive listening experience.

[0012] It exhibits good applicability and scalability: The conversion method in this application is designed strictly in accordance with international standards and the requirements of domestically developed systems, thus possessing strong applicability. Furthermore, the various processing steps of this method are relatively independent, facilitating adjustments and optimizations based on subsequent updates to technical standards or new application requirements, thereby demonstrating good scalability. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an audio processing method based on three-dimensional immersive sound according to an embodiment of this application is shown. Figure 2 A schematic diagram illustrating the principle of an audio processing method based on three-dimensional immersive sound according to an embodiment of this application is shown; Figure 3 A schematic diagram of the structure of an audio processing system based on three-dimensional immersive sound according to an embodiment of this application is shown; Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] To more comprehensively illustrate the technical solution of this application, the details of each technical aspect are explained below to ensure that those skilled in the art can fully understand and implement this application.

[0016] (I) Basic introduction to ADM audio definition model files.

[0017] The ADM audio definition model file uses XML as its specification language, based on the ITU-R BS.2076-2 standard. Its XML data can be embedded into specific data blocks within the audio file (e.g., ...). <axml>(in the block).

[0018] The structure of an ADM (Audio Definition Model) file is mainly divided into two parts: content and format. The content part describes the audio's creation information, such as the language of the dialogue and loudness; the format part describes the technical characteristics of the audio to ensure that it can be correctly decoded or rendered. An ADM file consists of a series of elements describing different aspects of the audio. Each element is represented by an XML element, containing several attributes and sub-elements, and they are linked together through mutual references.

[0019] ADM audio definition model files can support content and metadata descriptions based on various audio formats such as channels, objects, and scenes.

[0020] In the object-based ADM audio definition model file, the sound bed metadata is defined in the tag. <audio channelformat type definition=""DirectSpeakers”">The system is located within a Cartesian coordinate system, with X, Y, and Z axes ranging from [-1, 1], corresponding to the left-right, front-back, and up-down dimensions, respectively. The position of the audio source in the sound bed is static, and the specific coordinate values ​​corresponding to the positions of each speaker are shown in Table 1.

[0021] Table 1 shows the specific coordinates of each speaker location:

[0022] Object metadata is defined in tags. <audio channel format type definition=""Objects”">In Audio Vivid, the position, size, diffusion, and consistency of an audio object can all change dynamically. Its metadata is organized by time blocks, including audio block ID, start time (rtime) of each audio block, duration of each audio block, position coordinates of the audio object, jump position, and interpolation length threshold. The jump position controls the temporal interpolation method of the position value, ensuring that the audio object can complete a spatial jump within the time specified by the jump interpolation length threshold, rather than moving smoothly within the entire audio block. It's important to note that the size of audio blocks in Audio Vivid is not fixed.

[0023] Audio Vivid uses a Cartesian coordinate system to describe the location of audio data, with its origin located at the listening reference point (i.e., the center of the room, facing the screen). The XY plane coincides with the horizontal plane of the sound center of the main channel and surround channel speaker system. The coordinate system is defined as follows: the X-axis represents the left-right direction, with the right being positive; the Y-axis represents the front-back direction, with the front being positive; and the Z-axis represents the up-down direction, with the top being positive. The coordinate values ​​of the X, Y, and Z axes all range from [-1, 1].

[0024] (II) Basic introduction to the domestic immersive sound processing system for digital cinema.

[0025] The domestically developed immersive sound processing system for digital cinema provides object-based audio IAB encoding, decoding, and audio rendering functions. The object-based audio IAB encoding tool is developed according to the SMPTE ST 2098-2 standard, which has specific requirements for the input audio data and metadata, as follows: Audio data requirements: 48kHz sampling rate, mono format, 32-bit integer; Metadata requirements: Organized at 24 frames per second at a 48kHz sampling rate, each frame contains 2000 sampling points, and further divided into 8 audio blocks, each audio block corresponding to 250 sampling points and a duration of 5.2ms, as the basic unit for audio and metadata synchronization processing.

[0026] The object location metadata is represented using a right-handed Cartesian coordinate system, defining its position in space relative to the origin through three orthogonal axes (x, y, z). The x-axis corresponds to the horizontal (left-right) direction of the auditorium, the y-axis to the vertical (front-back) direction, and the z-axis to the vertical (up-down) direction. The audio object location coordinates are normalized, ranging from (0,0,0) to (1,1,1). The origin of the coordinate system is set at the front left corner of the auditorium; x=0 corresponds to the left wall, x=1 to the right wall; y=0 to the front wall, y=1 to the rear wall; z=0 corresponds to the sound radiation center of the main channel and surround channel speaker systems, and z=1 corresponds to the auditorium ceiling.

[0027] Based on the above coordinate system definition, typical positions in a movie theater playback environment are as follows: (0,0,0) represents the front left corner of the theater, located at the center of the left channel speaker system's sound radiation; (1,0,0) represents the front right corner of the theater, located at the center of the right channel speaker system's sound radiation; and (0.5,0.5,1) represents the center of the theater ceiling.

[0028] Because Audio Vivid's ADM audio definition model file differs from the input requirements of the domestic immersive sound processing system for digital cinema in terms of audio block duration and coordinate system definition, it is necessary to convert it into standardized data that meets the requirements through a series of conversion processing steps in this application.

[0029] Figure 1 A flowchart illustrating an audio processing method based on three-dimensional immersive sound according to an embodiment of this application is shown. Figure 2 A schematic diagram illustrating the principle of an audio processing method based on three-dimensional high-resolution audio according to an embodiment of this application is shown. (Refer to...) Figure 1 and Figure 2 As shown, this embodiment includes steps S110 to S130: Step S110: Obtain the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata.

[0030] Obtain the ADM audio definition model file exported by the Audio Vivid audio production tool. The ADM audio definition model file uses XML as its specification language, conforming to the ITU-R BS.2076-2 standard, and its XML data can be embedded within the audio file. <axml>In the block.

[0031] The ADM audio definition model file was parsed using an XML parsing tool to extract the single-track PCM (pulse code modulation) audio data, single-track object PCM audio data, and the corresponding raw metadata.

[0032] PCM is an encoding method that converts analog audio signals into digital signals, preserving the original sound quality information of the audio.

[0033] The raw metadata includes at least one non-fixed-duration audio block and the start time, duration, location coordinates of the audio data (especially object audio), jump position parameters, and jump interpolation length threshold for each non-fixed-duration audio block.

[0034] The location coordinates of the audio data further include: the location coordinates of the audio data at the start time of a non-fixed-duration audio block and the location coordinates of the audio data at the end time of a non-fixed-duration audio block.

[0035] Therefore, in some optional implementations, parsing the ADM audio definition model file to obtain the original monotrack audio data and the corresponding original metadata includes: parsing the ADM audio definition model file to extract the monotrack sound bed PCM audio data, the monotrack object PCM audio data and the corresponding original metadata; wherein, the original metadata includes: at least one non-fixed duration audio block and the start time, duration, position coordinates of the audio data, jump position parameters and jump interpolation length threshold of each non-fixed duration audio block.

[0036] Step S120: Perform standardized format conversion and independent storage on the original single-track audio data to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema.

[0037] The extracted audio data of the two types should be stored as WAV files respectively, and strictly in accordance with the input requirements of the domestic immersive sound processing system for digital cinema, set to a sampling rate of 48kHz, mono, and 32-bit integer.

[0038] As a lossless audio file format, WAV files can effectively avoid audio quality loss during storage, ensuring the audio quality of subsequent processing.

[0039] The parameter settings are based on the following: The IAB encoding tool of the domestic immersive sound processing system for digital cinemas is developed based on the SMPTE ST2098-2 standard. Its input audio sampling rate must be consistent with the standard sampling rate (48kHz) of the cinema sound reproduction system. The mono format can ensure the independence of each audio object, and the 32-bit integer format can ensure high-precision storage of audio data and avoid quantization distortion, thereby meeting the high-quality sound reproduction requirements of professional cinemas.

[0040] Therefore, in some alternative implementations, the original mono audio data is subjected to standardized format conversion and independent storage, including: storing the mono audio bed PCM audio data and the mono audio object PCM audio data as WAV files respectively; wherein the parameters of the WAV file are: sampling rate 48kHz, mono, 32-bit integer.

[0041] Step S130 involves performing fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing, and coordinate encoding processing on the original metadata to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

[0042] The standardization of raw metadata is the core technical aspect of this application, which includes four steps: fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing, and coordinate encoding processing.

[0043] 1. Fixed duration block processing.

[0044] In the ADM audio definition model, the duration of non-fixed-duration audio blocks is not fixed. However, the IAB encoding tool of the domestic immersive sound processing system for digital cinema requires the metadata to be organized at 24 frames per second at a sampling rate of 48kHz, with each frame containing 2000 sampling points, and further divided into 8 audio blocks, that is, each audio block corresponds to 250 sampling points and a duration of 5.2ms.

[0045] Therefore, it is necessary to perform fixed-duration block processing on the non-fixed-duration audio blocks in the original metadata, dividing each non-fixed-duration audio block into several 5.2ms fixed-duration audio blocks. During the block division process, the fixed-duration audio blocks inherit the jump position parameters and jump interpolation length threshold of the corresponding non-fixed-duration audio blocks.

[0046] In addition, since fixed-duration segmentation is performed on non-fixed-duration audio blocks, the timestamp of the start point (start time), the timestamp of the end point (end time), the position coordinates of the audio data at the start point of the fixed-duration audio block, and the position coordinates of the audio data at the end point of the fixed-duration audio block can be determined based on the start time, duration, and position coordinates of the audio data at the end point of the fixed-duration audio block.

[0047] Therefore, in some optional implementations, fixed-duration block processing includes: performing fixed-duration block processing on non-fixed-duration audio blocks according to a fixed duration of 5.2ms to obtain fixed-duration audio blocks.

[0048] 2. Interpolation calculation processing.

[0049] For each fixed-duration audio block, based on its inherited jump position parameters and jump interpolation length threshold, the corresponding interpolation algorithm is used to calculate the position coordinates of the audio data (especially the object audio) in each target frame within the fixed-duration audio block, ensuring the smoothness and accuracy of dynamic object audio position changes.

[0050] The specific logic of interpolation calculation is as follows: First, calculate the timestamp of the target frame. The timestamp of the start of this fixed-duration audio block The difference Calculated based on the following formula (1): , formula (1).

[0051] Then, based on the values ​​of the jump position parameters and With the jump interpolation length threshold ( Based on the size relationship, select the corresponding difference formula.

[0052] When jump Position=1 and When using the following jump interpolation formulas (2) to (4): , formula (2); , formula (3); , formula (4).

[0053] The jump interpolation formula is suitable for scenarios that require spatial jumping of the object's audio, ensuring that the object's audio... Complete a rapid switch in spatial location within a specified time.

[0054] When jump Position 1 or When using linear interpolation formulas (5) to (7), the following formulas are employed: , formula (5); , formula (6); , formula (7).

[0055] Linear interpolation formulas are suitable for scenarios where the audio of an object moves smoothly, ensuring the continuity of positional changes.

[0056] in, This indicates the coordinates of the end point of the audio data within the fixed-duration audio block. This indicates the coordinates of the starting point of the audio data within the fixed-duration audio block. This indicates the position coordinates of the target frame within the fixed-duration audio block. , The timestamp indicates the end of the fixed-duration audio block.

[0057] It should be noted that, , , All values ​​are coordinates in the Audio Vivid coordinate system.

[0058] Therefore, in some optional implementations, the interpolation calculation process includes: for a fixed-duration audio block, performing interpolation calculations on the fixed-duration audio block based on the jump position parameters inherited by the fixed-duration audio block and the jump interpolation length threshold; when jump Position=1 and When using this method, the following jump interpolation formula is employed: , , When jump position 1 or When using linear interpolation, the following formula is employed: , , Wherein, jump Position represents the jump position parameter inherited by this fixed-duration audio block. , This represents the timestamp of the target frame within the fixed-duration audio block. The timestamp indicating the start of a fixed-duration audio block. This indicates the coordinates of the end point of the audio data within the fixed-duration audio block. This indicates the coordinates of the starting point of the audio data within the fixed-duration audio block. This indicates the position coordinates of the target frame within the fixed-duration audio block. This indicates the threshold for the jump interpolation length. , The timestamp indicates the end of the fixed-duration audio block.

[0059] 3. Coordinate system transformation processing.

[0060] The Audio Vivid coordinate system differs from the cinema coordinate system used in domestic digital cinema immersive sound processing systems. Therefore, it is necessary to convert the position coordinates in the Audio Vivid coordinate system to the position coordinates in the cinema coordinate system using a coordinate transformation formula.

[0061] The following compares the differences between the Audio Vivid coordinate system and the cinema coordinate system.

[0062] The Audio Vivid coordinate system uses a Cartesian coordinate system, with the origin located at the listening reference point (the center of the room, facing the screen). The X-axis represents the left-right direction (right is positive), the Y-axis represents the front-back direction (front is positive), and the Z-axis represents the up-down direction (up is positive). The coordinate values ​​all range from [-1, 1].

[0063] The cinema coordinate system adopts a right-handed Cartesian coordinate system, with the origin located at the front left corner of the auditorium. The x-axis corresponds to the horizontal (left-right) direction of the auditorium (x=0 corresponds to the left wall, x=1 corresponds to the right wall), the y-axis corresponds to the vertical (front-back) direction of the auditorium (y=0 corresponds to the front wall, y=1 corresponds to the rear wall), and the z-axis corresponds to the vertical (up-down) direction of the auditorium (z=0 corresponds to the sound radiation center of the main channel and surround channel speaker system, z=1 corresponds to the auditorium ceiling). The coordinate values ​​are all in the range of [0,1].

[0064] According to Appendix B of T / CSMPTE 31–2024 "Subjective Evaluation Method for the Quality of Immersive Audio Rendering Technology in Digital Cinema", the following formulas (8) to (10) are used for coordinate transformation: , formula (8); , formula (9); , formula (10); in, This represents the position coordinates of the audio data in the Audio Vivid coordinate system. This indicates the position coordinates of the audio data in the cinema coordinate system.

[0065] The basis for the above conversion formula is as follows: Xx axis transformation: The range of values ​​[-1,1] in the Audio Vivid coordinate system is linearly mapped to the range [0,1] in the cinema coordinate system. Through the linear transformation of formula (8), the accurate correspondence of the left and right coordinates is achieved.

[0066] Yy axis transformation: Since the front of the Y axis is positive in the Audio Vivid coordinate system and the back of the Y axis is positive in the cinema coordinate system, the transformation of formula (9) is used to realize the reverse mapping of the front and back coordinates, so as to ensure that the front and back positions of the audio data are consistent with the actual space of the cinema.

[0067] Zz-axis transformation: Considering that the sound radiation center of a cinema speaker system is usually located in the lower middle part of the auditorium, and there are no speaker layouts below, therefore, in the Audio Vivid coordinate system... When converting it to the cinema coordinate system To avoid rendering errors caused by invalid lower position coordinates; when At this time, the original coordinate values ​​are directly retained to ensure accurate mapping of the position above.

[0068] Therefore, in some alternative implementations, the coordinate system transformation process includes: converting the position coordinates of the audio data in the AudioVivid coordinate system to the position coordinates in the cinema coordinate system based on the following transformation formula; , , ,in, This represents the position coordinates of the audio data in the Audio Vivid coordinate system. This indicates the position coordinates of the audio data in the cinema coordinate system.

[0069] 4. Coordinate encoding processing.

[0070] After completing the coordinate system transformation, the position coordinates within the range [0,1] in the cinema coordinate system need to be encoded as 16-bit unsigned integers as specified in the SMPTE ST2098-2 standard, so that the IAB encoding tool of the domestic immersive sound processing system for digital cinema can perform subsequent processing. The encoding formulas are formulas (11) to (13): , formula (11); , formula (12); , formula (13); Where n=16 (i.e., a 16-bit unsigned integer). This represents the encoded digital coordinate parameters.

[0071] Therefore, in some alternative implementations, the coordinate encoding process includes: encoding the position coordinates of the audio data in the cinema coordinate system based on the following encoding formula; , , ,in, , This represents the digital coordinate parameters of the audio data in the cinema coordinate system.

[0072] After completing the above processing, standardized metadata that meets the requirements of the domestic immersive sound processing system for digital cinema is obtained.

[0073] After standardizing audio processing (48kHz, mono, 32-bit integer WAV file) and metadata processing (16-bit unsigned integer encoded coordinate parameters and other related information), the standardized audio data and metadata are input into the domestic immersive sound processing system for digital cinema. This process ultimately generates an immersive audio stream (IAB) conforming to the SMPTE ST 2098-2 standard. The immersive audio playback tool decodes and renders the IAB stream, and after equalization and delay adjustments, it reproduces immersive sound in the theater through the cinema speaker system, achieving high-quality playback of Audio Vivid audio content in professional cinemas.

[0074] To verify the effectiveness of the audio format conversion method proposed in this application, a compliant evaluation environment was built according to the T / CSMPTE 31-2024 standard, "Subjective Evaluation Method for the Quality of Immersive Audio Rendering Technology in Digital Cinema," and corresponding test program sources were prepared. A subjective evaluation experiment was conducted on the adaptation results of object-based Audio Vivid audio with the domestic immersive sound processing system for digital cinema. The evaluation content mainly includes three aspects: sound quality, static objects (sound bed audio), and dynamic objects (object audio), specifically including: Sound quality: the degree of sound quality damage to the sound bed and the degree of sound quality damage to the target; Static objects: positional overlap, loudness overlap, size overlap; Dynamic objects: overlap due to changes in position, overlap due to changes in loudness, and overlap due to changes in distance.

[0075] The experimental results are shown in Table 2. All evaluation indicators reached the "excellent" level.

[0076] Table 2, Experimental Results:

[0077] The experimental results show that the adaptation results reached the "excellent" level in all evaluation indicators, specifically as follows: In terms of sound quality, the damage to the sound quality of both the sound bed and the object is imperceptible; For static objects, the position, loudness, and size of the rendered output are consistent with the description information of the reference source; Regarding dynamic objects, the changes in the object's position, loudness, and distance all match the description of the reference source.

[0078] In summary, the adaptation results demonstrate excellent performance across all subjective evaluation metrics for immersive audio rendering in digital cinema, meeting the technical requirement of "not lower than good" in the T / CSMPTE 31-2024 standard. This verifies the correctness of the conversion results and their effectiveness and reliability in audio quality, spatial positioning, and dynamic tracking.

[0079] Figure 3 A schematic diagram of the structure of an audio processing system based on three-dimensional high-fidelity sound according to an embodiment of this application is shown. (Refer to...) Figure 3 As shown, the audio processing system 300 based on three-dimensional sound enhancement includes: Data acquisition module 310 is used to acquire the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata. The audio data standardization module 320 is used to perform standardized format conversion and independent storage on the original single-track audio data to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema. The metadata standardization module 330 is used to perform fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing and coordinate encoding processing on the original metadata to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

[0080] In some optional implementations, in the above system, the data acquisition module 310 is used to: parse the ADM audio definition model file, extract single-track sound bed PCM audio data, single-track object PCM audio data and corresponding raw metadata; wherein, the raw metadata includes: at least one non-fixed duration audio block and the start time, duration, position coordinates of the audio data, jump position parameters and jump interpolation length threshold of each non-fixed duration audio block.

[0081] In some alternative implementations, in the above system, the audio data standardization module 320 is used to: store the mono audio bed PCM audio data and the mono audio object PCM audio data as WAV files respectively; wherein the parameters of the WAV file are: sampling rate 48kHz, mono, 32-bit integer.

[0082] In some alternative implementations, in the above system, the metadata standardization module 330 is used to: perform fixed-duration block processing on non-fixed-duration audio blocks according to a fixed duration of 5.2ms to obtain fixed-duration audio blocks.

[0083] In some optional implementations, in the above system, the metadata standardization module 330 is used to: for a fixed-duration audio block, perform interpolation operations on the fixed-duration audio block based on the jump position parameters and jump interpolation length threshold inherited by the fixed-duration audio block; when jump Position=1 and When using this method, the following jump interpolation formula is employed: , , When jump position 1 or When using linear interpolation, the following formula is employed: , , Wherein, jump Position represents the jump position parameter inherited by this fixed-duration audio block. , This represents the timestamp of the target frame within the fixed-duration audio block. The timestamp indicating the start of a fixed-duration audio block. This indicates the coordinates of the end point of the audio data within the fixed-duration audio block. This indicates the coordinates of the starting point of the audio data within the fixed-duration audio block. This indicates the position coordinates of the target frame within the fixed-duration audio block. This indicates the threshold for the jump interpolation length. , The timestamp indicates the end of the fixed-duration audio block.

[0084] In some alternative implementations, in the above system, the metadata standardization module 330 is used to: convert the position coordinates of the audio data in the Audio Vivid coordinate system to the position coordinates in the cinema coordinate system based on the following transformation formula; , , ,in, This represents the position coordinates of the audio data in the Audio Vivid coordinate system. This indicates the position coordinates of the audio data in the cinema coordinate system.

[0085] In some alternative implementations, in the above system, the metadata standardization module 330 is used to: encode the position coordinates of the audio data in the cinema coordinate system based on the following encoding formula; , , ,in, , This represents the digital coordinate parameters of the audio data in the cinema coordinate system.

[0086] It should be noted that the audio processing system 300 based on three-dimensional sound can implement the aforementioned audio processing methods based on three-dimensional sound, which will not be elaborated further.

[0087] Figure 4 This invention illustrates a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 4 As shown, the electronic device includes a processor, internal memory, a network interface, and a non-volatile storage medium connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external devices via a network connection. When executed by the processor, the computer program implements the functions or steps of a three-dimensional audio processing method based on holographic sound.

[0088] In one embodiment, the electronic device provided in this application includes a memory and a processor. The memory stores a database and a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the audio processing method based on three-dimensional sound.

[0089] The above is as stated in this application. Figure 3 The method executed by the audio processing system based on three-dimensional crystal sound disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0090] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of an audio processing method based on three-dimensional sound.

[0091] It should be noted that the functions or steps that the above-mentioned electronic devices or computer-readable storage media can achieve can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0092] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0093] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0094] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.< / axml> < / audio> < / audio> < / axml>

Claims

1. An audio processing method based on three-dimensional immersive sound, characterized in that, include: Obtain the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata; The original single-track audio data is subjected to standardized format conversion and independent storage to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema. The original metadata is processed by fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing and coordinate encoding processing to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

2. The method according to claim 1, characterized in that, The parsing of the ADM audio definition model file yields the original single-track audio data and the corresponding original metadata, including: Parse the ADM audio definition model file to extract the PCM audio data of a single audio track sound bed, the PCM audio data of a single audio track object, and the corresponding raw metadata. The raw metadata includes at least one non-fixed duration audio block and the start time, duration, position coordinates of the audio data, jump position parameters, and jump interpolation length threshold of each non-fixed duration audio block.

3. The method according to claim 2, characterized in that, The process of performing standardized format conversion and independent storage on the original single-track audio data includes: The PCM audio data of the mono audio bed and the PCM audio data of the mono audio object are stored as WAV files respectively; the parameters of the WAV file are: sampling rate 48kHz, mono, 32-bit integer.

4. The method according to claim 2, characterized in that, The fixed-duration block processing includes: By performing fixed-duration block processing on non-fixed-duration audio blocks with a fixed duration of 5.2ms, a fixed-duration audio block is obtained.

5. The method according to claim 4, characterized in that, The interpolation calculation process includes: For a fixed-duration audio block, interpolation is performed on the fixed-duration audio block based on the jump position parameters and jump interpolation length threshold inherited by the fixed-duration audio block; When jump Position=1 and When using this method, the following jump interpolation formula is employed: , , , When jump Position 1 or When using linear interpolation, the following formula is employed: , , , Here, "jump Position" represents the jump position parameter inherited by this fixed-duration audio block. , This represents the timestamp of the target frame within the fixed-duration audio block. The timestamp indicating the start of a fixed-duration audio block. This indicates the coordinates of the end point of the audio data within the fixed-duration audio block. This indicates the coordinates of the starting point of the audio data within the fixed-duration audio block. This indicates the position coordinates of the target frame within the fixed-duration audio block. This indicates the threshold for the jump interpolation length. , The timestamp indicates the end of the fixed-duration audio block.

6. The method according to claim 5, characterized in that, The coordinate system transformation process includes: The position coordinates of the audio data in the Audio Vivid coordinate system are converted to the position coordinates in the cinema coordinate system using the following transformation formula; , , , in, This indicates the position coordinates of the audio data in the Audio Vivid coordinate system. This indicates the position coordinates of the audio data in the cinema coordinate system.

7. The method according to claim 6, characterized in that, The coordinate encoding process includes: The position coordinates of the audio data in the cinema coordinate system are encoded based on the following encoding formula; , , , in, , This represents the digital coordinate parameters of the audio data in the cinema coordinate system.

8. An audio processing system based on three-dimensional immersive sound, characterized in that, include: The data acquisition module is used to acquire the ADM audio definition model file exported by the Audio Vivid audio production tool, and parse the ADM audio definition model file to obtain the original single-track audio data and the corresponding original metadata. The audio data standardization module is used to perform standardized format conversion and independent storage on the original single-track audio data to obtain standardized audio data that conforms to the domestic immersive sound processing system for digital cinema. The metadata standardization module is used to perform fixed-duration block processing, interpolation calculation processing, coordinate system transformation processing, and coordinate encoding processing on the original metadata to obtain standardized metadata that conforms to the domestic immersive sound processing system for digital cinema.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio processing method based on three-dimensional immersive sound as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the audio processing method based on three-dimensional immersive sound as described in any one of claims 1 to 7.