3D sound rendering methods, devices, terminals and computer program products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0015]本申请实施例中,为了实现系统级的三维声渲染,终端的系统中设置三维声渲染器,并由系统首先对获取到的原始三维声元数据进行处理,使处理得到的目标三维声元数据符合三维声渲染器的渲染需求,从而通过三维声渲染器对音频数据以及目标三维声元数据进行三维声渲染,得到三维声音频信号。
Smart Images

Figure CN122575381A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to a three-dimensional sound rendering method, apparatus, terminal, and computer program product. Background Technology
[0002] As an audio format, 3D sound carries multiple audio signals that constitute complete audio content through multiple channels. These signals are then directly reproduced by multiple speakers located at different heights around the listener, or reproduced after rendering or mapping. This provides higher sound image spatial resolution and gives the listener an immersive sound field experience. Summary of the Invention
[0003] This application provides a three-dimensional sound rendering method, apparatus, terminal, and computer program product. The technical solution is as follows:
[0004] On one hand, embodiments of this application provide a three-dimensional sound rendering method, the method being executed by a terminal system, the system being equipped with a three-dimensional sound renderer, the method comprising:
[0005] Acquire audio data and the corresponding original three-dimensional acoustic data;
[0006] The original 3D acoustic data is processed to obtain target 3D acoustic data, which meets the rendering requirements of the 3D acoustic renderer.
[0007] Based on the audio data and the target 3D sound data, the sound is rendered using the 3D sound renderer to obtain a 3D sound audio signal.
[0008] On the other hand, embodiments of this application provide a three-dimensional sound rendering apparatus, the apparatus comprising:
[0009] The acquisition module is used to acquire audio data and the original three-dimensional acoustic data corresponding to the audio data;
[0010] The processing module is used to process the original three-dimensional acoustic data to obtain target three-dimensional acoustic data, which meets the rendering requirements of the three-dimensional acoustic renderer.
[0011] The rendering module is used to render sound using a 3D sound renderer based on the audio data and the target 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound renderer is set in the system.
[0012] On the other hand, embodiments of this application provide a terminal, the terminal including a processor and a memory, the memory storing at least one computer instruction, the at least one computer instruction being loaded and executed by the processor to implement the three-dimensional sound rendering method as described above.
[0013] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the three-dimensional sound rendering method as described above.
[0014] On the other hand, embodiments of this application provide a computer program product, the computer program product including computer instructions stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the three-dimensional sound rendering method as described above.
[0015] In this embodiment of the application, in order to achieve system-level three-dimensional sound rendering, a three-dimensional sound renderer is set in the system of the terminal. The system first processes the acquired original three-dimensional sound data to make the processed target three-dimensional sound data meet the rendering requirements of the three-dimensional sound renderer. Then, the three-dimensional sound renderer performs three-dimensional sound rendering on the audio data and the target three-dimensional sound data to obtain a three-dimensional sound audio signal.
[0016] Since the rendering process is completed on the system side, the system can still provide 3D sound rendering and playback functionality for applications even if they do not natively support 3D sound rendering, thus expanding the application scenarios for 3D sound playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this application's embodiments helps improve the rendering effect because the target 3D sound metadata and audio data contain more information than the rendered channel data. Attached Figure Description
[0017] Figure 1 This is a schematic diagram illustrating the implementation of the three-dimensional audio-visual processing process in related technologies;
[0018] Figure 2 A flowchart illustrating a three-dimensional sound rendering method provided in an exemplary embodiment of this application is shown;
[0019] Figure 3 This is a schematic diagram illustrating a spatial position completion process in an exemplary embodiment of this application;
[0020] Figure 4 This is a flowchart illustrating a sound rendering process based on an audio track object, as shown in an exemplary embodiment of this application.
[0021] Figure 5 This is an exemplary embodiment of the present application illustrating the process of writing audio data and three-dimensional acoustic data into an audio track object;
[0022] Figure 6 This is an embodiment of another exemplary embodiment of the present application illustrating the process of writing audio data and three-dimensional acoustic data into an audio track object;
[0023] Figure 7 This is a flowchart illustrating the process of acquiring and rendering three-dimensional acoustic data and audio data, as shown in an exemplary embodiment of this application.
[0024] Figure 8 This is a schematic diagram illustrating an exemplary embodiment of the three-dimensional audio / video decoding and playback process of this application;
[0025] Figure 9 This is a schematic diagram of an exemplary embodiment of the sound object editing interface shown in this application;
[0026] Figure 10 This is a schematic diagram illustrating an embodiment of the three-dimensional audio-visual encoding process of this application;
[0027] Figure 11 A structural block diagram of a three-dimensional sound rendering apparatus provided in another exemplary embodiment of this application is shown;
[0028] Figure 12 A structural block diagram of a terminal provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0030] In this article, "multiple" refers to two or more. "And / or," describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0031] For ease of understanding, the terms used in the embodiments of this application will be explained below.
[0032] Three-dimensional audio bitstream: also known as three-dimensional audio coded bitstream, refers to the bitstream obtained by encoding audio data and its corresponding three-dimensional audio metadata. The audio data may include at least one of the following: channel-based audio data, sound object-based audio data, or HOA (High-Order Ambisonics) based audio data. In some embodiments, the channel-based audio data may be mono data, stereo data, or multi-channel surround sound data.
[0033] Audio data can be encoded using general bitrate audio coding or lossless audio coding, while 3D acoustic data is encoded using metadata coding. The encoded audio data and 3D acoustic data are then multiplexed through a 3D audio bitstream to obtain the 3D audio bitstream.
[0034] 3D audio decoding, the reverse process of 3D audio encoding, involves decoding the 3D audio bitstream using general-rate audio decoding or lossless audio decoding to obtain channel signals, object signals, or HOA signals. This is then followed by metadata decoding to obtain 3D audio metadata. The decoded audio data and the 3D audio metadata are then used for 3D audio rendering to produce a 3D audio signal. This 3D audio signal can then be used for speaker playback or headphone playback.
[0035] Three-dimensional acoustic metadata: In a three-dimensional acoustic system, three-dimensional acoustic metadata is used to describe spatial information such as the position, size, direction, and trajectory of a sound object, as well as technical parameters such as the encoding method, sampling rate, and number of channels of the audio signal. This application's embodiments limit the specific data content included in the three-dimensional acoustic metadata.
[0036] AudioTrack Object: A class used for playing decoded PCM (Pulse Code Modulation) audio data. It provides a low-level audio playback interface, suitable for low-latency playback scenarios and real-time audio applications. It contains audio data and other information related to the audio data, such as the audio data's sampling rate, bit width, length, type, etc. In this embodiment, the AudioTrack Object includes not only the audio data but also the corresponding three-dimensional acoustic data.
[0037] In related technologies, such as Figure 1 As shown, for applications that support 3D audio playback, the application decodes the 3D audio stream 11 and performs metadata processing to obtain dual-channel or multi-channel data, and then sends the data of each channel to the system.
[0038] The system uses an audio mixer 12 (for mixing audio from different applications) and a renderer 13 (for rendering special sound effects) to mix and render the data of each channel, and finally plays it through an audio playback device 14 (such as headphones or speakers).
[0039] Clearly, the decoding and rendering of the 3D audio stream are completed on the application side. The system only involves mixing, rendering, and outputting the channel data, and does not involve the direct processing of the sound object data.
[0040] Although the above solution simplifies the system-side design, it will not be possible to achieve 3D sound playback if the application does not natively support it, thus limiting the application scenarios of 3D sound playback.
[0041] In this embodiment, to expand the application scenarios of 3D sound playback, a 3D sound renderer is set on the system side of the terminal, and the system supports processing the acquired raw 3D sound data to obtain target 3D sound data that meets the rendering requirements of the 3D sound renderer. The system performs 3D sound rendering on the audio data and target 3D sound data through the 3D sound renderer to obtain a 3D sound audio signal. Since the system's target 3D sound data and audio data contain more information than the channel data rendered on the application side, it helps to improve the rendering effect.
[0042] The solution provided in this application can be executed by a terminal, which may be a smartphone, tablet, wearable device, computer, audio playback device (such as a speaker), etc. Furthermore, the terminal's system side is equipped with a 3D sound renderer, which enables the system to provide 3D sound rendering and playback services for the application. In some embodiments, the terminal can implement 3D audio processing through a processor or a separately configured audio processing chip.
[0043] In the following embodiments, for ease of description, the three-dimensional audio-visual processing method is described using an example of execution by a terminal (specifically, a terminal system), but this does not constitute a limitation.
[0044] Please refer to Figure 2 This document illustrates a flowchart of a three-dimensional sound rendering method provided in an exemplary embodiment of this application. This embodiment uses the method applied to a terminal as an example for illustration, and the method may include the following steps:
[0045] Step 201: Obtain the audio data and the corresponding original three-dimensional acoustic data.
[0046] The audio data and the original 3D acoustic data are decoded from the 3D audio stream. This 3D audio stream can be a real-time 3D audio stream, or an audio stream contained within a 3D audio file.
[0047] The audio data can be channel-based, sound object-based, or HOA (including FOA)-based. Channel-based audio data can be mono, dual-channel stereo, or multi-channel surround sound.
[0048] In some embodiments, the audio data and the original three-dimensional acoustic data can be obtained by the application decoding the three-dimensional acoustic audio stream, or by the three-dimensional acoustic decoder set on the system side.
[0049] In some embodiments, when an application has a 3D audio stream decoding function but no 3D audio rendering function, and when there is a need for 3D audio playback and the system's 3D audio rendering function is enabled, the application sends the decoded audio data and the original 3D audio data to the system. Correspondingly, the system receives the audio data and the original 3D audio data sent by the application.
[0050] In some embodiments, when the application does not have the functions of 3D audio stream decoding and 3D audio rendering, but has the requirement of 3D audio playback and the system's 3D audio rendering function is enabled, the application sends a 3D audio stream to the system. Correspondingly, the system receives the 3D audio stream sent by the application and decodes the 3D audio stream through the 3D audio decoder set on the system side to obtain audio data and original 3D audio data.
[0051] In some embodiments, when the application has 3D sound rendering capabilities, and there is a need for 3D sound playback, and the system's 3D sound rendering function is enabled, the user can choose whether the application or the system performs the 3D sound rendering. Specifically, if the user chooses the system to perform the 3D sound rendering, the application sends the decoded audio data and the original 3D sound data to the system, and the system receives the audio data and the original 3D sound data sent by the application.
[0052] In one possible implementation, a three-dimensional audio decoder is provided on the system side. This three-dimensional audio decoder can be integrated into the system's existing audio / video decoder, or it can be set up independently of the system's existing audio / video decoder.
[0053] In some embodiments, the 3D audio decoder comprises an audio decoder and a metadata decoder. The audio decoder performs audio decoding on the 3D audio stream to obtain audio data (PCM data); the metadata decoder performs metadata decoding on the 3D audio stream to obtain 3D audio metadata.
[0054] Optionally, the audio decoder supports general bitrate audio decoding and / or lossless audio decoding, and can use the corresponding decoding method to perform audio decoding according to the encoding method used in the three-dimensional audio bitstream.
[0055] It should be noted that before decoding the three-dimensional audio and video stream, the system needs to perform preprocessing such as decapsulation on the stream, which will not be elaborated here in this embodiment.
[0056] Of course, the audio data and the original 3D acoustic data can also be pre-decoded and transmitted to the application, which will then send them to the system for rendering. This embodiment does not limit the specific method of obtaining the audio data and the original 3D acoustic data.
[0057] Step 202: Process the original 3D acoustic data to obtain the target 3D acoustic data, which meets the rendering requirements of the 3D acoustic renderer.
[0058] Since the obtained raw 3D acoustic data may not meet the rendering requirements of the system's 3D acoustic renderer, the system needs to process the raw 3D acoustic data to obtain target 3D acoustic data that meets the rendering requirements of the 3D acoustic renderer in order to ensure the feasibility of subsequent 3D acoustic rendering.
[0059] The rendering requirements include, but are not limited to, requirements for the format of 3D audio data, requirements for the integrity of 3D audio data, and requirements for the effects of 3D audio data (such as the need for transition effects when 3D audio data changes).
[0060] In some embodiments, the system detects whether the original 3D acoustic data meets the rendering requirements. If it does, the original 3D acoustic data is determined as the target 3D acoustic data; if it does not, the original 3D acoustic data is processed. The specific method of processing the original 3D acoustic data is described in detail in the following embodiments.
[0061] Step 203: Based on the audio data and the target 3D acoustic data, perform sound rendering using a 3D sound renderer to obtain a 3D audio signal.
[0062] In this embodiment, a 3D sound renderer is provided on the system side. After obtaining the audio data and the target 3D sound data, the system uses the audio data and the target 3D sound data as input to the 3D sound renderer, and performs 3D sound rendering to obtain 3D sound audio signals.
[0063] In some embodiments, the system can select the appropriate rendering method for 3D sound rendering based on the device type of the audio playback device.
[0064] In some embodiments, the rendered 3D audio signal can be a mono audio signal, a stereo audio signal, or a multi-channel audio signal.
[0065] In some embodiments, after the system completes the 3D sound rendering, it can further mix and render the 3D sound audio signal using a mixer and a renderer, and output it to an audio playback device for 3D sound playback. The audio playback device can be the terminal's own playback device, such as the terminal's speaker, or an external playback device, such as headphones, speakers, etc. This embodiment does not limit the specific device used.
[0066] In summary, in this embodiment of the application, in order to achieve system-level three-dimensional sound rendering, a three-dimensional sound renderer is set in the system of the terminal. The system first processes the acquired original three-dimensional sound data to make the processed target three-dimensional sound data meet the rendering requirements of the three-dimensional sound renderer. Thus, the three-dimensional sound renderer performs three-dimensional sound rendering on the audio data and the target three-dimensional sound data to obtain a three-dimensional sound audio signal.
[0067] Since the rendering process is completed on the system side, the system can still provide 3D sound rendering and playback functionality for applications even if they do not natively support 3D sound rendering, thus expanding the application scenarios for 3D sound playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this application's embodiments helps improve the rendering effect because the target 3D sound metadata and audio data contain more information than the rendered channel data.
[0068] 3D acoustic data processing process
[0069] Since the 3D audio data carried in the 3D audio bitstream may have some missing metadata, and the 3D audio data may not conform to the specific protocol acquisition format specification, the preprocessing methods for 3D audio data include at least one of metadata completion and metadata conversion.
[0070] Metadata completion is used to supplement missing metadata in the 3D audio metadata. In one possible implementation, the terminal completes the missing necessary metadata in the 3D audio metadata based on metadata completion rules. This necessary metadata is essential for the 3D audio rendering process.
[0071] In some embodiments, the 3D audio bitstream contains at least one complete 3D audio metadata. The terminal supplements the missing metadata in the currently decoded 3D audio metadata based on the complete 3D audio metadata and / or the 3D audio metadata obtained from previous decoding.
[0072] Metadata conversion is used to convert 3D audio-visual data to a specific data format. In some embodiments, if the 3D audio-visual data does not conform to a standard metadata format, the terminal system converts the 3D audio-visual data into a standard data format. This standard data format may include Dolby format, AudioVivid format, etc., and this embodiment does not limit this to any particular format.
[0073] In one possible implementation, the system processes the raw 3D acoustic data from both data integrity and effect requirements, and may include at least one of the following methods:
[0074] 1. Perform data completion processing on the original 3D acoustic data to obtain the target 3D acoustic data.
[0075] In one possible implementation, the terminal system performs a data integrity check on the original 3D acoustic data to determine if any data is missing. If data is missing, the terminal system performs data completion processing on the original 3D acoustic data to obtain the target 3D acoustic data.
[0076] Regarding the method of data integrity verification, in some embodiments, the terminal system acquires the complete full-scale 3D acoustic data at least once, and then checks whether the current original 3D acoustic data is complete based on the full-scale 3D acoustic data. The missing original 3D acoustic data is a subset of the full-scale 3D acoustic data.
[0077] Optionally, the full 3D audio data can be distributed by the application side, or obtained by the terminal system by decoding the 3D audio stream.
[0078] In the event of missing data, in one possible implementation, the terminal performs data completion processing on the original 3D acoustic data based on the full amount of historically acquired 3D acoustic data to obtain the target 3D acoustic data.
[0079] In some embodiments, the terminal determines the missing metadata type in the original 3D acoustic metadata by the difference between the metadata type contained in the full 3D acoustic metadata and the metadata type contained in the original 3D acoustic metadata, and then completes the specific data of the metadata type so that the metadata type contained in the completed target 3D acoustic metadata is consistent with the metadata type contained in the full 3D acoustic metadata.
[0080] Regarding the specific data source for the metadata type to be supplemented, in one possible implementation, the data can be a default value or historical 3D acoustic metadata. For example, when the original 3D acoustic metadata lacks spatial coordinates, the terminal can use the default Cartesian coordinate system for supplementation; when the original 3D acoustic metadata lacks sound bed information, the terminal can use the sound bed information from the most recently acquired historical 3D acoustic metadata for supplementation.
[0081] 2. Generate transition metadata based on the original 3D acoustic metadata; generate target 3D acoustic metadata containing the original 3D acoustic metadata and transition metadata.
[0082] For performance and other reasons, 3D audio metadata will not be configured for the packaged audio data. For example, if the 3D audio metadata remains unchanged for a long time (i.e., the 3D sound effect remains unchanged), duplicate 3D audio metadata will not be issued. To ensure normal 3D sound rendering even when 3D audio metadata is not obtained, the terminal needs to have a 3D audio metadata caching function. When no new 3D audio metadata is obtained, the system caches the 3D audio metadata and configures it to the 3D sound renderer, allowing the 3D sound renderer to perform 3D sound rendering based on the old 3D audio metadata. When new 3D audio metadata is obtained, the system needs to parse the new 3D audio metadata and configure it to the 3D sound renderer.
[0083] Because the attributes of channels, sound beds, and sound objects in the old and new 3D audio metadata may change—for example, the spatial location of a sound object may change, and the gain of a channel may change—to avoid obvious sound effect gaps before and after the change in 3D audio metadata, such as a sudden change in the spatial orientation of a sound object or the sudden disappearance of the sound of a sound object, in one possible implementation, after obtaining the original 3D audio metadata, the terminal needs to generate transition metadata based on the original 3D audio metadata, and jointly determine the original 3D audio metadata and transition metadata as the target 3D audio metadata, and configure it to the 3D audio renderer, so that the 3D audio renderer can smoothly transition the 3D audio signal rendered based on the target 3D audio metadata.
[0084] Optionally, the transition metadata is metadata located between the historical 3D acoustic metadata and the currently acquired original 3D acoustic metadata, used to achieve the transition of sound effects from the 3D sound effects represented by the historical 3D acoustic metadata to the 3D sound effects represented by the original 3D acoustic metadata.
[0085] In some embodiments, when the original three-dimensional acoustic metadata includes at least one of the channel gain values of the sound channels and the spatial coordinates of the sound object, generating transition metadata may include at least one of the following two cases:
[0086] Case 1: Based on the channel gain values in the original 3D acoustic metadata, generate transition metadata, which includes the transition gain values of the channels.
[0087] In some embodiments, for the same channel, when the channel gain value contained in the original 3D acoustic data is inconsistent with the channel gain value contained in the historical 3D acoustic data, the terminal generates transition metadata containing transition gain values based on the channel gain values in the original 3D acoustic data and the historical 3D acoustic data. The process of determining the transition metadata may include the following steps:
[0088] 1. Based on the first channel gain value of the channel in the original three-dimensional acoustic data and the second channel gain value of the channel in the historical three-dimensional acoustic data, determine the gain change of the channel.
[0089] In one possible implementation, for the same channel, the terminal determines the gain change of that channel as the difference between the first channel gain value in the original 3D audio data and the second channel gain value in the historical 3D audio data. A negative gain change indicates a decrease in the gain of that channel; a positive gain change indicates an increase in the gain of that channel.
[0090] In some embodiments, when the channel is a newly added channel, the second channel gain of the channel is 0, and when the channel is a channel that is about to disappear, the first channel gain of the channel is 0. Accordingly, by generating transition metadata, a fade-in / fade-out effect can be achieved.
[0091] In some embodiments, in order to reduce the computational load of the 3D sound rendering process, if the absolute value of the gain change is greater than a threshold (i.e., the gain change of the same channel is significant), the terminal executes the subsequent process of generating transition metadata; if the absolute value of the gain change is less than or equal to the threshold (i.e., the gain change of the same channel is not significant), the terminal does not need to determine the transition metadata.
[0092] 2. Based on the gain change and transition duration, determine the transition gain value, which is located between the first channel gain value and the second channel gain value.
[0093] The transition duration is the time it takes for the gain value of the second channel to change to the gain value of the first channel. Optionally, the transition duration can be a preset duration.
[0094] In some embodiments, the terminal determines at least one transition gain value between the first channel gain value and the second channel gain value based on the gain change and the transition duration. For example, when n transition gain values need to be determined, the i-th transition gain value = the second channel gain value + i * (gain change / n).
[0095] 3. Generate transition metadata based on the original 3D acoustic data and transition gain value.
[0096] Furthermore, the terminal generates transition metadata based on the transition gain value and metadata other than the gain value in the original three-dimensional acoustic metadata to ensure the integrity of the transition metadata.
[0097] Case 2: Based on the spatial coordinates of the acoustic object in the original 3D acoustic metadata, generate transition metadata, which contains the transition spatial coordinates of the acoustic object.
[0098] In some embodiments, for the same acoustic object, when the spatial coordinates of the acoustic object contained in the original 3D acoustic data are inconsistent with the spatial coordinates of the acoustic object contained in the historical 3D acoustic data, the terminal generates transitional metadata containing transitional spatial coordinates based on the spatial coordinates of the acoustic object in the original 3D acoustic data and the historical 3D acoustic data. The process of determining the transitional metadata may include the following steps:
[0099] 1. Based on the first spatial coordinates of the acoustic object in the original 3D acoustic data and the second spatial coordinates of the acoustic object in the historical 3D acoustic data, determine the coordinate change of the acoustic object.
[0100] In one possible implementation, for the same acoustic object, the terminal determines the coordinate change of the acoustic object as the difference between the first spatial coordinate of the acoustic object in the original three-dimensional acoustic data and the second spatial coordinate of the acoustic object in the historical three-dimensional acoustic data. This coordinate change is used to characterize the magnitude of the spatial position change of the acoustic object.
[0101] In some embodiments, to reduce the computational load of the 3D sound rendering process, the terminal detects whether the spatial position change amplitude represented by the coordinate change is greater than an amplitude threshold. If the spatial position change amplitude is greater than the amplitude threshold, the terminal executes the subsequent process of generating transition metadata; if the spatial position change amplitude is less than or equal to the amplitude threshold, the terminal does not execute the subsequent process of generating transition metadata. The spatial position change amplitude can be represented by the absolute value or the sum of squares of the coordinate changes. For example, when the coordinate change is (x, y, z), the represented spatial position change amplitude can be |x|+|y|+|z|, or, x... 2 +y 2 +z 2 .
[0102] 2. Based on the coordinate change and transition time, determine the transition space coordinates, which are located on the change path between the first space coordinates and the second space coordinates.
[0103] The transition duration is the time required to change from the second spatial coordinate to the first spatial coordinate. Optionally, the transition duration can be a preset duration.
[0104] In some embodiments, the path of change from the second spatial coordinates to the first spatial coordinates can be a preset path. This preset path can be a straight line, a curve, a polyline, etc., and this embodiment does not limit this.
[0105] In some embodiments, the terminal determines at least one transition space coordinate located on the change path based on the coordinate change amount and the transition duration. For example, when it is necessary to determine n transition space coordinates and the change path is a straight line, the i-th transition space coordinate = the second space coordinate + i * (coordinate change amount / n).
[0106] Indicative, such as Figure 3 As shown, when the historical three-dimensional acoustic data indicates that the acoustic object 32 in the virtual space 31 is located at the first position, and the currently acquired original three-dimensional acoustic data indicates that the acoustic object 32 is located at the second position, the terminal determines three transition positions (corresponding to their respective transition space coordinates) based on the transition duration and the change path between the first and second positions.
[0107] 3. Generate transition metadata based on the original 3D acoustic data and transition space coordinates.
[0108] Furthermore, the terminal generates transition metadata based on the transition spatial coordinates and metadata other than spatial coordinates in the original three-dimensional acoustic metadata, in order to ensure the integrity of the transition metadata.
[0109] In this embodiment, the terminal completes the 3D audio metadata to ensure that the subsequent 3D audio renderer can render based on the complete 3D audio metadata, thus ensuring the correct execution of the rendering process. The terminal generates transition metadata based on historical and current 3D audio metadata, enabling the subsequent 3D audio renderer to achieve fade-in and fade-out of the audio channels and gradual changes in the position of the audio objects based on the transition metadata, thereby improving the smoothness of the 3D audio effect change process.
[0110] Definition of 3D acoustic data
[0111] In one possible implementation, the definition of three-dimensional acoustic data can be as shown in Table 1.
[0112] Table 1
[0113]
[0114] It should be noted that the raw 3D acoustic data obtained by the terminal system may include all or part of the data in Table 1.
[0115] The timestamp and spatial coordinates of the sound object are used to characterize the spatial location of the sound object at different times.
[0116] In some embodiments, the initial value of the rendering flag in the 3D audio metadata corresponding to the audio data is false. For audio data that has undergone 3D audio rendering, the rendering flag in its corresponding 3D audio metadata is set to true. Subsequently, when other devices obtain the audio data (which may be rendered or unrendered) and the 3D audio metadata, they can determine whether further 3D audio rendering of the audio signal is needed based on the rendering flag.
[0117] Optionally, when the target 3D acoustic data includes channel information and at least one of sound bed information and sound object information, the 3D acoustic rendering process may include the following steps:
[0118] 1. Based on the vocal tract information and the vocal bed information, determine the vocal bed audio data in the audio data.
[0119] In some embodiments, when the sound bed information indicates the presence of a sound bed in the audio data, the terminal determines the channel to which the sound bed belongs based on the sound bed information and the channel information, and then determines the sound bed audio data in the audio data.
[0120] 2. Based on the channel information and the sound object information, determine the sound object audio data in the audio data.
[0121] In some embodiments, when the sound object information indicates the presence of a sound object in the audio data, for each sound object, the terminal determines the channel to which each sound object belongs based on the channel information and the correspondence between the sound object and the channel indicated by the sound object information, and then determines the sound object audio data of each sound object in the audio data.
[0122] 3. Based on the audio data and information of the sound bed, perform sound bed rendering using a 3D sound renderer, and / or, based on the audio data and information of the sound object, perform sound object rendering using a 3D sound renderer to obtain 3D sound and audio signals.
[0123] In some embodiments, the 3D sound renderer performs sound bed rendering based on the relevant sound bed parameters and sound bed audio data in the sound bed information; and / or performs sound object rendering (rendering the sound object to a specified spatial location) based on the relevant sound object parameters (such as spatial coordinates) in the sound object information. Further, the 3D sound renderer generates 3D sound and audio signals based on the sound bed audio signal and the sound object audio signal.
[0124] 3D sound rendering process
[0125] In typical non-3D sound rendering scenarios, the system usually writes audio data into an audio track object, which is then rendered by the sound renderer. However, in 3D sound rendering scenarios, because it is necessary to reconstruct the spatial information of the sound object, such as its position, size, direction, and motion trajectory, the audio track object needs to be modified to enable it to carry both audio data and 3D sound data simultaneously.
[0126] In some embodiments, such as Figure 4 As shown, based on audio data and 3D sound data, sound rendering using a 3D sound renderer to obtain a 3D sound audio signal can include the following steps:
[0127] Step 203A: Write the audio data and the target 3D acoustic data into the audio track object.
[0128] In order for the subsequent 3D sound renderer to obtain the 3D sound metadata corresponding to the audio data when rendering the sound track object, the system needs to write the audio data and the processed target 3D sound metadata together into the sound track object, that is, it needs to extend the ability of the audio object to carry 3D sound metadata.
[0129] In some embodiments, applications can invoke the system-side 3D audio decoder for decoding and preprocessing in various ways. Correspondingly, the way the system writes audio data and target 3D audio data to the audio track object differs under different invocation methods. The following describes the writing process of audio data and target 3D audio data under different invocation methods.
[0130] Invocation method 1: The application invokes the system's 3D sound decoder by sending a playback request.
[0131] In one possible implementation, when an application requires 3D audio / video playback, it invokes the system-side media player (MediaPlayer) interface by sending a playback request. During this invocation, the application transmits the 3D audio / video stream to the media player.
[0132] Accordingly, in response to the application's playback request, the terminal calls the 3D audio decoder through the system's built-in player to decode the 3D audio stream and obtain the audio data and the original 3D audio metadata.
[0133] Optionally, if multiple decoders are set on the system side, when the playback request is detected to indicate the playback of 3D audio, the system's built-in player calls the 3D audio decoder to decode the 3D audio stream.
[0134] Furthermore, after decoding the audio data and the original 3D audio data by calling the 3D audio decoder, the terminal system processes the original 3D audio data to obtain the target 3D audio data, and writes the audio data and the target 3D audio data into the audio track object through the system's built-in player.
[0135] It should be noted that when it is necessary to play 3D sound and the corresponding video at the same time, the terminal also calls the video decoder through the system's built-in player to decode the video stream, which will not be described in detail here.
[0136] In an illustrative example, such as Figure 5 As shown, when an application at the application (APP) layer requires 3D audio playback (AudioPlay), it calls the MediaPlayer interface at the Java Native Interface (JNI) layer. In response to the call to the MediaPlayer interface, the terminal, through the system's built-in player (Nuplayer) at the Framework layer, calls the 3D Audio decoder (which also includes MP3, AAC, and other audio decoders) in the MediaCodec to decode the 3D audio stream, obtaining audio data and 3D audio metadata. The system's built-in player further writes the decoded audio data and 3D audio metadata into an AudioTrack object for subsequent 3D audio rendering.
[0137] Method 2: The application calls the system's 3D audio decoder by sending a decoding request.
[0138] In one possible implementation, when an application requires 3D audio / video playback, it invokes the system-side media codec interface by sending a decoding request, directly calling the system-side 3D audio decoder for decoding. During this call to the media codec interface, the application transmits the 3D audio / video bitstream to the media decoder.
[0139] In response to the application's decoding request, the terminal calls the 3D audio decoder to decode the 3D audio stream and obtain the audio data and the original 3D audio metadata.
[0140] Optionally, if multiple decoders are set up on the system side, when the decoding request is detected to decode the three-dimensional audio and video stream, the terminal calls the three-dimensional audio decoder to decode the three-dimensional audio and video stream.
[0141] Since the application has only requested decoding so far, the decoded audio data and the processed target 3D acoustic data need to be sent back to the application so that the application can instruct on further processing.
[0142] When an application needs to further utilize the system's 3D sound rendering capabilities for 3D sound rendering, the application sends audio data and target 3D sound data to the system's audio service system. Correspondingly, the terminal writes the audio data and target 3D sound data sent by the application into an audio track object. For example, the application can write the audio data and target 3D sound data into the audio track object using the `write` method of the audio track object.
[0143] It should be noted that when it is necessary to play 3D sound and the corresponding video at the same time, the terminal also calls the video decoder to decode the video stream, which will not be described in detail in this embodiment.
[0144] In an illustrative example, such as Figure 6 As shown, when an application at the application (APP) layer requires 3D audio playback (AudioPlay), it first calls the MediaCodec interface at the Java Native Interface (JNI) layer. In response to the call to the MediaCodec interface, the terminal calls the 3D Audio decoder (which also includes MP3, AAC, and other audio decoders) in the MediaCodec layer at the Framework layer to decode the 3D audio stream, obtaining audio data and 3D audio metadata, which are then sent back to the application. The application further writes the decoded audio data and 3D audio metadata into an AudioTrack object for subsequent 3D audio rendering.
[0145] Step 203B: Render the audio track object using a 3D sound renderer to obtain a 3D audio signal.
[0146] In some embodiments, the 3D sound renderer obtains audio data and corresponding target 3D sound data from the audio track object, and then performs sound rendering on the audio data based on the target 3D sound data to obtain a 3D sound audio signal.
[0147] In this embodiment, by modifying the data structure of the audio track object, the audio track object can carry both audio data and three-dimensional audio data simultaneously. When the three-dimensional sound renderer performs sound rendering on the audio track object, it can accurately obtain the audio data and its corresponding three-dimensional audio data, thereby ensuring the correct rendering of the three-dimensional audio.
[0148] In addition, for different calling methods of the application side calling the system side 3D audio decoder, specific data decoding and writing schemes are set to ensure the normal decoding of 3D audio data and the correct writing of data in the audio track object.
[0149] Considering that the speed at which the 3D audio decoder decodes and obtains the 3D audio metadata may not be consistent with the speed at which the audio track object consumes the 3D audio metadata, in one possible implementation, the 3D audio metadata is cached in the metadata buffer queue of the audio track object, while the audio data is cached in the audio data buffer of the audio track object.
[0150] In one possible implementation, when creating an audio track object, the system allocates a shared memory block (metadata buffer queue) for the reading and writing of 3D audio data, thereby enabling data transfer between processes (decoding process and rendering process). This shared memory can be in the form of a circular queue to ensure normal operation even when the buffer overflows.
[0151] In addition, the size of the shared memory can be determined by the system based on the processing time required for the three-dimensional audio data carried by the audio data within the same time period. For example, when processing 1920 frames of audio data with a period of 20ms, and carrying 2 three-dimensional audio data, the time required is also 20ms, then the size of the shared memory is the size of 2 three-dimensional audio data.
[0152] Of course, the terminal can also implement the metadata buffer queue in other ways, and this embodiment does not limit this.
[0153] The terminal then retrieves audio data and 3D audio metadata from the audio buffer and metadata buffer queue, respectively, and performs sound rendering on the retrieved audio data and 3D audio metadata using a 3D sound renderer. For example... Figure 7 As shown, the process may include the following steps:
[0154] Step 701: Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object.
[0155] Since 3D audio metadata may not be continuously effective during audio playback, but rather effective at a certain point in the audio playback process, or effective from a certain point in time during audio playback, it is necessary to determine the effective 3D audio metadata at the current moment when retrieving the 3D audio metadata for sound rendering from the metadata buffer queue.
[0156] In some embodiments, the 3D acoustic metadata written to the metadata buffer queue carries a start timestamp, which is the timestamp when the 3D acoustic metadata begins to take effect.
[0157] In some embodiments, the start timestamp is a timestamp that starts from the audio playback start time (i.e., the duration relative to the audio playback start time). For example, when the start timestamp is 10, it indicates that the three-dimensional acoustic data is effective from the 10th second after the audio playback starts, that is, from the 10th second after the audio playback starts, the subsequently played audio has the spatial effect represented by the three-dimensional acoustic data.
[0158] To identify effective 3D audio metadata based on its start timestamp, the audio track object in this embodiment maintains an object timestamp. This object timestamp is relative to the start time of audio playback, indicating the duration of audio playback; that is, the object timestamp is continuously updated during audio playback. Accordingly, the terminal determines the 3D audio metadata currently applied to sound rendering based on this object timestamp and the start timestamp of the 3D audio metadata.
[0159] In some embodiments, the terminal obtains the effective 3D audio metadata from the metadata buffer queue of the audio track object based on the object timestamp maintained by the audio track object and the start timestamp of the 3D audio metadata in the metadata buffer queue.
[0160] In one possible implementation, the terminal determines the effective 3D acoustic data as those whose start timestamp is less than or equal to the object timestamp by comparing the start timestamp and the object timestamp.
[0161] As an illustration, when the object timestamp maintained by the audio track object is 10, the terminal will determine the three-dimensional audio metadata with a starting timestamp of 10 in the metadata buffer queue as the effective three-dimensional audio metadata.
[0162] Step 702: Obtain audio data from the audio buffer of the audio track object.
[0163] In one possible implementation, the terminal retrieves the decoded audio data from the audio buffer in a first-in-first-out (FIFO) order.
[0164] It should be noted that there is no strict sequential order between steps 701 and 702 above, that is, steps 701 and 702 can be executed synchronously. This embodiment does not impose any restrictions on the execution order of the two.
[0165] Step 703: Use a 3D sound renderer to render the audio data and the effective 3D sound data to obtain a 3D sound audio signal.
[0166] The audio data and effective 3D audio metadata extracted from the audio buffer and metadata buffer queue are sent to the 3D audio renderer, which performs 3D audio rendering to obtain 3D audio signals.
[0167] In this embodiment, by caching the 3D audio metadata to the metadata buffer queue of the audio track object, the decoding and consumption speed of the 3D audio metadata is kept consistent. Furthermore, the terminal system determines the effective 3D audio metadata based on the object timestamp maintained by the audio track object and the start timestamp carried by the 3D audio metadata, and sends it to the 3D audio renderer for sound rendering, ensuring the accuracy of the timing of 3D audio rendering.
[0168] Optionally, the 3D audio metadata can also carry a termination timestamp, which is the timestamp when the 3D audio metadata stops being effective. Similarly, this termination timestamp is a timestamp starting from the audio playback start time.
[0169] Of course, 3D acoustic data may also not carry an end timestamp. For 3D acoustic data without an end timestamp, the 3D acoustic data ceases to be effective from the start timestamp of the next 3D acoustic data.
[0170] In some embodiments, when the effective 3D audio metadata includes a termination timestamp, during 3D audio rendering based on the effective 3D audio metadata, the 3D audio renderer stops rendering audio based on the object timestamp maintained by the audio track object and the termination timestamp of the effective 3D audio metadata.
[0171] Choosing a 3D sound renderer
[0172] In one possible implementation, the terminal's system side is equipped with multiple 3D sound renderers, each corresponding to a different 3D sound rendering method. These 3D sound rendering methods include speaker rendering, binaural rendering, and so on.
[0173] Accordingly, when rendering sound based on audio data and 3D audio data, the terminal renders the sound using a 3D audio renderer that matches the device type of the audio playback device, thus obtaining a 3D audio signal. This audio playback device can be the terminal's own playback device (such as a mobile phone speaker) or an external playback device (such as a speaker or headphones).
[0174] In one possible implementation, the terminal system has a mapping relationship between the device type of the audio playback device and the 3D sound renderer. The system then determines the 3D sound renderer that matches the device type of the current audio playback device based on this mapping relationship.
[0175] Indicative, such as Figure 8As shown, the application sends the 3D audio stream 81 to the 3D audio decoder 82 on the system side, where the decoder 82 decodes the audio data and 3D audio metadata. After metadata preprocessing, the decoded 3D audio metadata, along with the audio data, is sent to the 3D audio renderer 83 on the system side for sound rendering. Specifically, when the audio playback device is a multi-channel speaker, the 3D audio renderer 83 performs VectorBase Amplitude Panning (VBAP) to obtain a multi-channel audio signal, which is then mixed by a multi-channel mixer before being output. When the audio playback device is a headphone or speaker, the 3D audio renderer performs stereo rendering to obtain a stereo audio signal, which is then mixed by a stereo mixer before being output. For speakers, crosstalk cancellation is performed before output; for headphones, head rotation compensation is performed based on head rotation.
[0176] 3D acoustic data generation
[0177] In one possible scenario, the application or system provides a sound object adjustment function. Using this function, users can adjust the spatial position of sound objects during 3D audio playback, or control the sound objects to move along a specific trajectory. To reproduce the adjusted 3D sound effect, in one possible implementation, in response to a metadata generation operation, the terminal generates 3D sound metadata and writes the audio data and the generated 3D sound metadata into the audio track object.
[0178] In some embodiments, the terminal displays a sound object adjustment interface, which includes a virtual space and sound object identifiers for each sound object located within that virtual space. Users can adjust the sound objects using these sound object identifiers.
[0179] The initial position of the acoustic object identifier in the virtual space is determined based on the spatial position information of the acoustic object in the initial three-dimensional acoustic metadata.
[0180] Indicative, such as Figure 9 As shown, the sound object adjustment interface includes a virtual control 91, and a first sound object identifier 92 for the first sound object, a second sound object identifier 93 for the second sound object, and a third sound object identifier 94 for the third sound object, all located in the virtual space 91.
[0181] Optionally, this metadata generation operation can be an adjustment operation for the sound object identifier. For example, adjusting the position of the sound object identifier, moving the sound object identifier along a movement trajectory, adjusting the size of the sound object, adjusting the direction of the sound object, etc.
[0182] Of course, in addition to adjusting the sound object identifier, new three-dimensional sound metadata can also be generated in other ways (i.e., the metadata generation operation can also be other types of operation), such as directly modifying the parameter values of the three-dimensional sound metadata of the sound object. This application embodiment does not limit this.
[0183] Accordingly, the three-dimensional acoustic data generated by the terminal may include the spatial location information of the adjusted acoustic object, the motion trajectory of the acoustic object, the volume of the adjusted acoustic object, the direction of the adjusted acoustic object, etc. This application embodiment does not limit the specific content included in the generated three-dimensional acoustic data.
[0184] Indicative, such as Figure 9 As shown, when the user drags the third sound object identifier 94 along a specific motion trajectory, the terminal generates three-dimensional sound data corresponding to the third sound object, which contains the motion trajectory of the third sound object.
[0185] After the terminal system writes the audio data and the generated 3D sound data into the audio track object, the 3D sound renderer on the system side renders the sound based on the audio track object, thus restoring the 3D sound effect after the sound object is adjusted.
[0186] The specific process of writing audio data and generated 3D sound data into the audio track object, and performing sound rendering based on the audio track object, can be referred to in the above embodiments, and will not be repeated here.
[0187] For example, such as Figure 9 As shown, the 3D sound renderer renders sound based on the 3D sound data containing the motion trajectory of the third sound object and the corresponding audio data of the third sound object, thus reproducing the acoustic effects produced when the third sound object moves along a specific motion trajectory in space.
[0188] To ensure the accuracy of the rendering timing for sound based on the generated 3D audio metadata, when generating the 3D audio metadata, the terminal determines the start timestamp of the generated 3D audio metadata based on the object timestamp maintained by the audio track object. The object timestamp is a timestamp relative to the start time of audio playback, and the start timestamp is the timestamp when the 3D audio metadata begins to take effect.
[0189] In some embodiments, when generating three-dimensional audio metadata, the terminal determines the current object timestamp maintained by the audio track object as the starting timestamp of the generated three-dimensional audio metadata.
[0190] In some embodiments, the terminal can also determine the end timestamp of the generated 3D audio metadata based on the object timestamp maintained by the audio track object. For example, when the audio object is moved along a specific motion trajectory, the terminal determines the object timestamp maintained by the audio track object at the start of the movement as the start timestamp of the 3D audio metadata, and determines the object timestamp maintained by the audio track object at the end of the movement as the end timestamp of the 3D audio metadata. That is, the generated 3D audio metadata includes the start timestamp, the end timestamp, and the motion trajectory.
[0191] In this embodiment, the terminal system has a metadata generation function, which can dynamically generate three-dimensional sound metadata according to the received metadata generation operation, and supports writing the generated three-dimensional sound metadata and audio data into the audio track object for the three-dimensional sound renderer to perform sound rendering, so that users can hear the changed three-dimensional sound effect in real time.
[0192] 3D sound recording
[0193] To facilitate subsequent reproduction, a 3D audio encoder can also be set on the system side. In audio recording mode, the terminal system performs 3D audio rendering based on the generated 3D audio metadata and audio data, and simultaneously encodes the audio data and the generated 3D audio metadata through the 3D audio encoder to obtain a 3D audio encoded bitstream.
[0194] The 3D acoustic encoder includes an audio encoder and a metadata encoder. The audio encoder is used to encode audio data, while the metadata encoder is used to encode 3D acoustic metadata. Furthermore, the audio encoder can be a general-purpose bitrate audio encoder or a lossless audio encoder; this embodiment does not limit this.
[0195] In some embodiments, the 3D acoustic encoded bitstream can be saved as a file and provided to the application.
[0196] Indicative, in Figure 8 On the basis of, such as Figure 10 As shown, the application side is also equipped with a 3D sound encoder 84. In response to the metadata generation operation, the generated 3D sound metadata and the decoded audio data are sent together to the 3D sound renderer 83 for sound rendering; on the other hand, the generated 3D sound metadata and audio data are copied (the copied 3D sound metadata and audio data are not processed by sound rendering) and sent to the 3D sound encoder 84 for encoding.
[0197] Besides adjusting the 3D audio data in real time during decoding and playback, other possible methods include... Figure 10As shown, PCM audio data can also be provided by the application, and the system can generate three-dimensional acoustic metadata for the PCM audio data (the user can manually adjust the sound object to trigger the generation of three-dimensional acoustic metadata), and then send the PCM audio data and the generated three-dimensional acoustic metadata together into the three-dimensional sound encoder 84 for encoding.
[0198] In this embodiment, the terminal system provides a three-dimensional sound coding function. With the help of this function, the terminal system can encode the real-time generated three-dimensional sound metadata and audio data, which facilitates the subsequent three-dimensional sound reproduction based on the three-dimensional sound coding bitrate obtained by encoding.
[0199] Please refer to Figure 11 This illustration shows a structural block diagram of a three-dimensional sound rendering apparatus provided in an exemplary embodiment of this application. The apparatus includes:
[0200] The acquisition module 1101 is used to acquire audio data and the original three-dimensional acoustic data corresponding to the audio data;
[0201] Processing module 1102 is used to process the original three-dimensional acoustic data to obtain target three-dimensional acoustic data, which meets the rendering requirements of the three-dimensional acoustic renderer.
[0202] The rendering module 1103 is used to render sound using a 3D sound renderer based on the audio data and the target 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound renderer is set in the system.
[0203] Optionally, the processing module 1102 is used for:
[0204] The original three-dimensional acoustic data is subjected to data completion processing to obtain the target three-dimensional acoustic data.
[0205] Transitional metadata is generated based on the original 3D acoustic metadata; target 3D acoustic metadata containing the original 3D acoustic metadata and the transitional metadata is generated.
[0206] Optionally, when performing data completion processing on the original three-dimensional acoustic data to obtain the target three-dimensional acoustic data, the processing module 1102 is used to:
[0207] Based on the full set of historical 3D acoustic data, the original 3D acoustic data is supplemented to obtain the target 3D acoustic data.
[0208] Optionally, the original three-dimensional acoustic data includes at least one of the channel gain value of the audio channel and the spatial coordinates of the acoustic object;
[0209] When generating transition metadata based on the original three-dimensional acoustic metadata, the processing module 1102 is used to:
[0210] Based on the channel gain values in the original three-dimensional acoustic data, the transition metadata is generated, and the transition metadata includes the transition gain values of the channels;
[0211] Based on the spatial coordinates of the acoustic object in the original three-dimensional acoustic metadata, the transition metadata is generated, and the transition metadata includes the transition spatial coordinates of the acoustic object.
[0212] Optionally, the processing module 1102 is used for:
[0213] Based on the first channel gain value of the channel in the original three-dimensional acoustic data and the second channel gain value of the channel in the historical three-dimensional acoustic data, the gain change of the channel is determined.
[0214] Based on the gain change and the transition duration, a transition gain value is determined, which is located between the first channel gain value and the second channel gain value.
[0215] The transition metadata is generated based on the original three-dimensional acoustic metadata and the transition gain value.
[0216] Optionally, the processing module 1102 is used for:
[0217] Based on the first spatial coordinates of the acoustic object in the original three-dimensional acoustic data and the second spatial coordinates of the acoustic object in the historical three-dimensional acoustic data, the coordinate change of the acoustic object is determined.
[0218] Based on the coordinate change and transition duration, the transition space coordinates are determined, and the transition space coordinates are located on the change path between the first space coordinates and the second space coordinates;
[0219] The transition metadata is generated based on the original three-dimensional acoustic metadata and the transition space coordinates.
[0220] Optionally, the rendering module 1103 is used for:
[0221] Write the audio data and the target 3D acoustic data into the audio track object;
[0222] The audio track object is rendered using the 3D audio renderer to obtain the 3D audio signal.
[0223] Optionally, the three-dimensional acoustic metadata is cached in the metadata buffer queue of the audio track object, and the audio data is cached in the audio data buffer of the audio track object;
[0224] The rendering module 1103 is used for:
[0225] Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object;
[0226] The audio data is obtained from the audio buffer of the audio track object;
[0227] The audio data and the effective 3D audio data are rendered using the 3D audio renderer to obtain the 3D audio signal.
[0228] Optionally, the target three-dimensional acoustic data includes a start timestamp, which is the timestamp when the target three-dimensional acoustic data begins to take effect, and the start timestamp is a timestamp relative to the audio playback start time;
[0229] The rendering module 1103 is used for:
[0230] Based on the object timestamp maintained by the audio track object and the start timestamp of the target 3D audio metadata in the metadata buffer queue, the effective 3D audio metadata is obtained from the metadata buffer queue of the audio track object, where the object timestamp is a timestamp relative to the audio playback start time.
[0231] Optionally, the effective 3D audio data includes a termination timestamp, which is the timestamp when the target 3D audio data stops being effective, and the termination timestamp is a timestamp relative to the audio playback start time;
[0232] The rendering module 1103 is also used for:
[0233] Based on the object timestamp maintained by the audio track object and the termination timestamp of the effective 3D audio metadata, stop sound rendering based on the effective 3D audio metadata.
[0234] Optionally, the device further includes:
[0235] A generation module is used to determine the start timestamp of the generated three-dimensional audio metadata in response to the metadata generation operation, based on the object timestamp maintained by the audio track object.
[0236] The rendering module 1103 is also used to write the audio data and the generated three-dimensional audio data into the audio track object.
[0237] Optionally, the target three-dimensional acoustic data includes vocal tract information, and the target three-dimensional acoustic data includes at least one of acoustic bed information and acoustic object information;
[0238] The rendering module 1103 is used for:
[0239] Based on the vocal tract information and the acoustic bed information, determine the acoustic bed audio data in the audio data;
[0240] Based on the channel information and the sound object information, determine the sound object audio data in the audio data;
[0241] Based on the sound bed audio data and the sound bed information, sound bed rendering is performed through the three-dimensional sound renderer, and / or, based on the sound object audio data and the sound object information, sound object rendering is performed through the three-dimensional sound renderer to obtain the three-dimensional sound audio signal.
[0242] Optionally, the acquisition module 1101 is used for:
[0243] Obtain the audio data sent by the application and the original three-dimensional acoustic data corresponding to the audio data;
[0244] or,
[0245] Obtain the three-dimensional audio stream issued by the application; decode the three-dimensional audio stream using the three-dimensional audio decoder set by the system to obtain the audio data and the original three-dimensional audio data corresponding to the audio data.
[0246] In summary, in this embodiment of the application, in order to achieve system-level three-dimensional sound rendering, a three-dimensional sound renderer is set in the system of the terminal. The system first processes the acquired original three-dimensional sound data to make the processed target three-dimensional sound data meet the rendering requirements of the three-dimensional sound renderer. Thus, the three-dimensional sound renderer performs three-dimensional sound rendering on the audio data and the target three-dimensional sound data to obtain a three-dimensional sound audio signal.
[0247] Since the rendering process is completed on the system side, the system can still provide 3D sound rendering and playback functionality for applications even if they do not natively support 3D sound rendering, thus expanding the application scenarios for 3D sound playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this application's embodiments helps improve the rendering effect because the target 3D sound metadata and audio data contain more information than the rendered channel data.
[0248] It should be noted that the apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.
[0249] See Figure 12 , Figure 12 This is a schematic diagram of the structure of a terminal provided in an exemplary embodiment of this application. The terminal may also include one or more of the following components: a processor 1210 and a memory 1220.
[0250] Optionally, the processor 1210 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1220, and by calling data stored in the memory 1220. Optionally, the processor 1210 can be implemented in at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA).
[0251] The processor 1210 can integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), and a baseband chip. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content displayed on the touchscreen; the NPU implements artificial intelligence (AI) functions; and the baseband chip handles wireless communication. It is understood that the baseband chip can also be implemented as a separate chip without being integrated into the processor 1210.
[0252] The memory 1220 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1220 may include a non-transitory computer-readable storage medium. The memory 1220 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1220 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the electronic device, etc.
[0253] In addition, those skilled in the art will understand that the structure of the terminal shown in the above figures does not constitute a limitation on the terminal. The terminal may include more (e.g., power supply components, display components, sensor components) or fewer components than shown, or combine certain components, or have different component arrangements.
[0254] This application provides a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the three-dimensional sound rendering method as described in the above embodiments.
[0255] On the other hand, embodiments of this application provide a computer program product, the computer program product including computer instructions stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the three-dimensional sound rendering method as described in the above embodiments.
[0256] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0257] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A three-dimensional sound rendering method, characterized in that, The method is executed by a terminal system equipped with a 3D sound renderer, and the method includes: Acquire audio data and the corresponding original three-dimensional acoustic data; The original 3D acoustic data is processed to obtain target 3D acoustic data, which meets the rendering requirements of the 3D acoustic renderer. Based on the audio data and the target 3D sound data, the sound is rendered using the 3D sound renderer to obtain a 3D sound audio signal.
2. The method according to claim 1, characterized in that, The process of processing the original three-dimensional acoustic data to obtain the target three-dimensional acoustic data includes at least one of the following: The original three-dimensional acoustic data is subjected to data completion processing to obtain the target three-dimensional acoustic data. Transition metadata is generated based on the original three-dimensional acoustic metadata. Generate the target 3D acoustic metadata containing the original 3D acoustic metadata and the transition metadata.
3. The method according to claim 2, characterized in that, The step of performing data completion processing on the original three-dimensional acoustic data to obtain the target three-dimensional acoustic data includes: Based on the full set of historical 3D acoustic data, the original 3D acoustic data is supplemented to obtain the target 3D acoustic data.
4. The method according to claim 2, characterized in that, The original three-dimensional acoustic data includes at least one of the channel gain value of the acoustic channel and the spatial coordinates of the acoustic object; The generation of transition metadata based on the original three-dimensional acoustic metadata includes at least one of the following: Based on the channel gain values in the original three-dimensional acoustic data, the transition metadata is generated, and the transition metadata includes the transition gain values of the channels; Based on the spatial coordinates of the acoustic object in the original three-dimensional acoustic metadata, the transition metadata is generated, and the transition metadata includes the transition spatial coordinates of the acoustic object.
5. The method according to claim 4, characterized in that, The process of generating the transition metadata based on the channel gain values in the original three-dimensional acoustic metadata includes: Based on the first channel gain value of the channel in the original three-dimensional acoustic data and the second channel gain value of the channel in the historical three-dimensional acoustic data, the gain change of the channel is determined. Based on the gain change and the transition duration, a transition gain value is determined, which is located between the first channel gain value and the second channel gain value. The transition metadata is generated based on the original three-dimensional acoustic metadata and the transition gain value.
6. The method according to claim 4, characterized in that, The generation of transition metadata based on the spatial coordinates of the acoustic object in the original three-dimensional acoustic metadata includes: Based on the first spatial coordinates of the acoustic object in the original three-dimensional acoustic data and the second spatial coordinates of the acoustic object in the historical three-dimensional acoustic data, the coordinate change of the acoustic object is determined. Based on the coordinate change and transition duration, the transition space coordinates are determined, and the transition space coordinates are located on the change path between the first space coordinates and the second space coordinates; The transition metadata is generated based on the original three-dimensional acoustic metadata and the transition space coordinates.
7. The method according to any one of claims 1 to 6, characterized in that, The process of rendering sound using the 3D sound renderer based on the audio data and the target 3D sound data to obtain a 3D sound audio signal includes: Write the audio data and the target 3D acoustic data into the audio track object; The audio track object is rendered using the 3D audio renderer to obtain the 3D audio signal.
8. The method according to claim 7, characterized in that, The three-dimensional acoustic metadata is cached in the metadata buffer queue of the audio track object, and the audio data is cached in the audio data buffer of the audio track object. The step of rendering the audio track object using the 3D audio renderer to obtain the 3D audio signal includes: Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object; The audio data is obtained from the audio buffer of the audio track object; The audio data and the effective 3D audio data are rendered using the 3D audio renderer to obtain the 3D audio signal.
9. The method according to claim 8, characterized in that, The target 3D acoustic data includes a start timestamp, which is the timestamp when the target 3D acoustic data begins to take effect, and the start timestamp is a timestamp relative to the start time of audio playback; The step of retrieving effective 3D audio metadata from the metadata buffer queue of the audio track object includes: Based on the object timestamp maintained by the audio track object and the start timestamp of the target 3D audio metadata in the metadata buffer queue, the effective 3D audio metadata is obtained from the metadata buffer queue of the audio track object, where the object timestamp is a timestamp relative to the audio playback start time.
10. The method according to claim 9, characterized in that, The effective 3D audio data includes a termination timestamp, which is the timestamp when the target 3D audio data stops being effective, and the termination timestamp is a timestamp relative to the audio playback start time; The method further includes: Based on the object timestamp maintained by the audio track object and the termination timestamp of the effective 3D audio metadata, stop sound rendering based on the effective 3D audio metadata.
11. The method according to claim 9, characterized in that, The method further includes: In response to the metadata generation operation, the start timestamp of the generated three-dimensional acoustic metadata is determined based on the object timestamp maintained by the audio track object; The audio data and the generated 3D acoustic data are written into the audio track object.
12. The method according to any one of claims 1 to 11, characterized in that, The target three-dimensional acoustic data includes vocal tract information, and the target three-dimensional acoustic data includes at least one of acoustic bed information and acoustic object information; The process of rendering sound using the 3D sound renderer based on the audio data and the target 3D sound data to obtain a 3D sound audio signal includes: Based on the vocal tract information and the acoustic bed information, determine the acoustic bed audio data in the audio data; Based on the channel information and the sound object information, determine the sound object audio data in the audio data; Based on the sound bed audio data and the sound bed information, sound bed rendering is performed through the three-dimensional sound renderer, and / or, based on the sound object audio data and the sound object information, sound object rendering is performed through the three-dimensional sound renderer to obtain the three-dimensional sound audio signal.
13. The method according to any one of claims 1 to 12, characterized in that, The acquisition of audio data and the corresponding original three-dimensional acoustic data includes: Obtain the audio data sent by the application and the original three-dimensional acoustic data corresponding to the audio data; or, Obtain the three-dimensional audio stream issued by the application; decode the three-dimensional audio stream using the three-dimensional audio decoder set by the system to obtain the audio data and the original three-dimensional audio data corresponding to the audio data.
14. A three-dimensional sound rendering device, characterized in that, The device includes: The acquisition module is used to acquire audio data and the original three-dimensional acoustic data corresponding to the audio data; The processing module is used to process the original three-dimensional acoustic data to obtain target three-dimensional acoustic data, which meets the rendering requirements of the three-dimensional acoustic renderer. The rendering module is used to render sound using a 3D sound renderer based on the audio data and the target 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound renderer is set in the system.
15. A terminal, characterized in that, The terminal includes a processor and a memory, the memory storing at least one computer instruction, which is loaded and executed by the processor to implement the three-dimensional sound rendering method as described in any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer instruction, which is executed by a processor to implement the three-dimensional sound rendering method as described in any one of claims 1 to 13.
17. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the three-dimensional sound rendering method as described in any one of claims 1 to 13.