Three-dimensional audio and video processing methods, devices, terminals, and computer program products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0015]本申请实施例中,为了实现系统级的三维声解码渲染,终端的系统中设置三维声解码器以及三维声渲染器,由系统通过三维声解码器对应用下发的三维声音频码流进行解码,并通过三维声渲染器对解码得到的音频数据以及三维声元数据进行三维声渲染,得到三维声音频信号。由于解码以及渲染均在系统侧完成,因此即便在应用原生不支持三维声播放的情况下,系统也能够为应用提供三维声播放功能,扩大了三维声播放的应用场景。并且,相较于由应用将渲染后的声道数据交由系统进行渲染播放,采用本申请实施例提供的方案,系统获取到原始的三维声元数据以及音频数据相较于渲染后的声道数据包含更多信息,因此有助于提高渲染效果。
Smart Images

Figure CN122575380A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to a three-dimensional audio processing method, apparatus, terminal, and computer program product. Background Technology
[0002] As an audio format, 3D sound carries multiple audio signals that constitute complete audio content through multiple channels. These signals are then directly reproduced by multiple speakers located at different heights around the listener, or reproduced after rendering or mapping. This provides higher sound image spatial resolution and gives the listener an immersive sound field experience. Summary of the Invention
[0003] This application provides a three-dimensional audio / video processing method, apparatus, terminal, and computer program product. The technical solution is as follows:
[0004] On one hand, embodiments of this application provide a three-dimensional audio processing method, which is executed by a terminal system. The system is equipped with a three-dimensional audio decoder and a three-dimensional audio renderer. The method includes:
[0005] Obtain the 3D audio and video streams issued by the application;
[0006] The three-dimensional audio stream is decoded by the three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio data.
[0007] Based on the audio data and the three-dimensional sound data, the sound is rendered using the three-dimensional sound renderer to obtain a three-dimensional sound audio signal.
[0008] On the other hand, embodiments of this application provide a three-dimensional audio-visual processing apparatus, the apparatus comprising:
[0009] The acquisition module is used to acquire the 3D audio and video bitstreams issued by the application;
[0010] The decoding module is used to decode the three-dimensional audio stream using a three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio data.
[0011] The rendering module is used to render sound using a 3D sound renderer based on the audio data and the 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound decoder and the 3D sound renderer are set in the system.
[0012] On the other hand, embodiments of this application provide a terminal, the terminal including a processor and a memory, the memory storing at least one computer instruction, the at least one computer instruction being loaded and executed by the processor to implement the three-dimensional audio-visual processing method as described above.
[0013] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the three-dimensional audio-visual processing method as described above.
[0014] On the other hand, embodiments of this application provide a computer program product, the computer program product including computer instructions stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the three-dimensional audio-visual processing method as described above.
[0015] In this embodiment, to achieve system-level 3D audio decoding and rendering, the terminal system is equipped with a 3D audio decoder and a 3D audio renderer. The system decodes the 3D audio stream sent by the application using the 3D audio decoder, and renders the decoded audio data and 3D audio metadata using the 3D audio renderer to obtain a 3D audio signal. Since both decoding and rendering are completed on the system side, the system can provide 3D audio playback functionality even if the application does not natively support 3D audio playback, thus expanding the application scenarios for 3D audio playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this embodiment allows the system to obtain more information from the original 3D audio metadata and audio data compared to the rendered channel data, thus improving the rendering effect. Attached Figure Description
[0016] Figure 1 This is a schematic diagram illustrating the implementation of the three-dimensional audio-visual processing process in related technologies;
[0017] Figure 2 A flowchart illustrating a three-dimensional audio-visual processing method provided in an exemplary embodiment of this application is shown;
[0018] Figure 3 This is a flowchart illustrating a sound rendering process based on an audio track object, as shown in an exemplary embodiment of this application.
[0019] Figure 4 This is an exemplary embodiment of the present application illustrating the process of writing audio data and three-dimensional acoustic data into an audio track object;
[0020] Figure 5 This is an embodiment of another exemplary embodiment of the present application illustrating the process of writing audio data and three-dimensional acoustic data into an audio track object;
[0021] Figure 6 This is a flowchart illustrating the process of acquiring and rendering three-dimensional acoustic data and audio data, as shown in an exemplary embodiment of this application.
[0022] Figure 7 This is a schematic diagram illustrating an exemplary embodiment of the three-dimensional audio / video decoding and playback process of this application;
[0023] Figure 8 This is a schematic diagram of an exemplary embodiment of the sound object editing interface shown in this application;
[0024] Figure 9 This is a schematic diagram illustrating an embodiment of the three-dimensional audio-visual encoding process of this application;
[0025] Figure 10 A structural block diagram of a three-dimensional audio-visual processing apparatus provided in another exemplary embodiment of this application is shown;
[0026] Figure 11 A structural block diagram of a terminal provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0028] In this article, "multiple" refers to two or more. "And / or," describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0029] For ease of understanding, the terms used in the embodiments of this application will be explained below.
[0030] Three-dimensional audio bitstream: also known as three-dimensional audio coded bitstream, refers to the bitstream obtained by encoding audio data and its corresponding three-dimensional audio metadata. The audio data may include at least one of the following: channel-based audio data, sound object-based audio data, or HOA (High-Order Ambisonics) based audio data. In some embodiments, the channel-based audio data may be mono data, stereo data, or multi-channel surround sound data.
[0031] Audio data can be encoded using general bitrate audio coding or lossless audio coding, while 3D acoustic data is encoded using metadata coding. The encoded audio data and 3D acoustic data are then multiplexed through a 3D audio bitstream to obtain the 3D audio bitstream.
[0032] 3D audio decoding, the reverse process of 3D audio encoding, involves decoding the 3D audio bitstream using general-rate audio decoding or lossless audio decoding to obtain channel signals, object signals, or HOA signals. This is then followed by metadata decoding to obtain 3D audio metadata. The decoded audio data and the 3D audio metadata are then used for 3D audio rendering to produce a 3D audio signal. This 3D audio signal can then be used for speaker playback or headphone playback.
[0033] Three-dimensional acoustic metadata: In a three-dimensional acoustic system, three-dimensional acoustic metadata is used to describe spatial information such as the position, size, direction, and trajectory of a sound object, as well as technical parameters such as the encoding method, sampling rate, and number of channels of the audio signal. This application's embodiments limit the specific data content included in the three-dimensional acoustic metadata.
[0034] AudioTrack Object: A class used for playing decoded PCM (Pulse Code Modulation) audio data. It provides a low-level audio playback interface, suitable for low-latency playback scenarios and real-time audio applications. It contains audio data and other information related to the audio data, such as the audio data's sampling rate, bit width, length, type, etc. In this embodiment, the AudioTrack Object includes not only the audio data but also the corresponding three-dimensional acoustic data.
[0035] In related technologies, such as Figure 1 As shown, for applications that support 3D audio playback, the application decodes the 3D audio stream 11 and performs metadata processing to obtain dual-channel or multi-channel data, and then sends the data of each channel to the system.
[0036] The system uses an audio mixer 12 (for mixing audio from different applications) and a renderer 13 (for rendering special sound effects) to mix and render the data of each channel, and finally plays it through an audio playback device 14 (such as headphones or speakers).
[0037] Clearly, the decoding and rendering of the 3D audio stream are completed on the application side. The system only involves mixing, rendering, and outputting the channel data, and does not involve the direct processing of the sound object data.
[0038] Although the above solution simplifies the system-side design, it will not be possible to achieve 3D sound playback if the application does not natively support it, thus limiting the application scenarios of 3D sound playback.
[0039] In this embodiment, to expand the application scenarios of 3D sound playback, a 3D sound decoder and a 3D sound renderer are set on the system side of the terminal. The system decodes the 3D sound audio stream sent by the application using the 3D sound decoder, and renders the decoded audio data and 3D sound data using the 3D sound renderer to obtain a 3D sound audio signal. Since the 3D sound data and audio data obtained by the system through 3D sound decoding contain more information than the channel data rendered by the application side, it helps to improve the rendering effect.
[0040] The solution provided in this application can be executed by a terminal, which may be a smartphone, tablet, wearable device, computer, audio playback device (such as a speaker), etc. Furthermore, the terminal's system side is equipped with a 3D audio decoder and a 3D audio renderer. With the help of these three-dimensional audio decoders and renderers, the application only needs to provide the system with a 3D audio stream, and the system can provide 3D audio playback services for the application. In some embodiments, the terminal can implement 3D audio processing through a processor or a separately configured audio processing chip.
[0041] In the following embodiments, for ease of description, the three-dimensional audio-visual processing method is described using an example of execution by a terminal (specifically, a terminal system), but this does not constitute a limitation.
[0042] Please refer to Figure 2 This document illustrates a flowchart of a three-dimensional audio / video processing method provided in an exemplary embodiment of this application. This embodiment uses the method applied to a terminal as an example for illustration, and the method may include the following steps:
[0043] Step 201: Obtain the 3D audio and video streams issued by the application.
[0044] The 3D audio stream can be a real-time 3D audio stream, or the 3D audio file can contain an audio stream.
[0045] Optionally, the application can be either a first-type application that does not have 3D audio / video stream decoding and rendering capabilities, or a second-type application that does have 3D audio / video stream decoding and rendering capabilities.
[0046] In some embodiments, when the application belongs to the first type of application, has a requirement for 3D sound playback, and the system's 3D sound decoding and rendering function is enabled, the application sends a 3D sound audio stream to the system, and the system receives the sent 3D sound audio stream accordingly.
[0047] Optionally, if the application belongs to the first category of applications, has a requirement for 3D sound playback, and the system's 3D sound decoding and rendering function is not enabled, the system may prompt the user to enable the system's 3D sound decoding and rendering function.
[0048] In other embodiments, when the application belongs to the second type of application, has a requirement for 3D sound playback, and the system's 3D sound decoding and rendering function is enabled, the application sends a 3D sound audio stream to the system, and the system receives the sent 3D sound audio stream accordingly.
[0049] Optionally, if the application belongs to the second category of applications, has a requirement for 3D sound playback, and the system's 3D sound decoding and rendering function is not enabled, the application can perform 3D sound decoding and rendering, and send the rendered 3D sound audio signal to the system for mixing and rendering output.
[0050] Of course, in other possible implementations, for the second type of application, a 3D sound decoding and rendering selection function can be provided, allowing the user to choose whether the application or the system performs the 3D sound decoding and rendering.
[0051] Step 202: Decode the three-dimensional audio stream using a three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio metadata.
[0052] In this embodiment, a three-dimensional audio decoder is provided on the system side. This three-dimensional audio decoder can be integrated into the system's existing audio / video decoder, or it can be set up independently of the system's existing audio / video decoder.
[0053] In some embodiments, the 3D audio decoder comprises an audio decoder and a metadata decoder. The audio decoder performs audio decoding on the 3D audio stream to obtain audio data (PCM data); the metadata decoder performs metadata decoding on the 3D audio stream to obtain 3D audio metadata.
[0054] The audio data can be channel-based, sound object-based, or HOA (including FOA)-based. Channel-based audio data can be mono, dual-channel stereo, or multi-channel surround sound.
[0055] Optionally, the audio decoder supports general bitrate audio decoding and / or lossless audio decoding, and can use the corresponding decoding method to perform audio decoding according to the encoding method used in the three-dimensional audio bitstream.
[0056] It should be noted that before decoding the three-dimensional audio and video stream, the system needs to perform preprocessing such as decapsulation on the stream, which will not be elaborated here in this embodiment.
[0057] Step 203: Based on the audio data and the three-dimensional sound data, the sound is rendered using a three-dimensional sound renderer to obtain a three-dimensional sound audio signal.
[0058] In this embodiment, a 3D sound renderer is provided on the system side. After decoding the audio data and 3D sound data metadata through the 3D sound decoder, the system uses the audio data and 3D sound data metadata as input to the 3D sound renderer, and performs 3D sound rendering to obtain 3D sound audio signals.
[0059] In some embodiments, the system can select the appropriate rendering method for 3D sound rendering based on the device type of the audio playback device.
[0060] In some embodiments, the rendered 3D audio signal can be a mono audio signal, a stereo audio signal, or a multi-channel audio signal.
[0061] In some embodiments, after the system completes the 3D sound rendering, it can further mix and render the 3D sound audio signal using a mixer and a renderer, and output it to an audio playback device for 3D sound playback. The audio playback device can be the terminal's own playback device, such as the terminal's speaker, or an external playback device, such as headphones, speakers, etc. This embodiment does not limit the specific device used.
[0062] In summary, in this embodiment of the application, to achieve system-level 3D audio decoding and rendering, a 3D audio decoder and a 3D audio renderer are configured in the terminal system. The system decodes the 3D audio stream sent by the application using the 3D audio decoder, and renders the decoded audio data and 3D audio metadata using the 3D audio renderer to obtain a 3D audio signal. Since both decoding and rendering are completed on the system side, the system can provide 3D audio playback functionality even if the application does not natively support 3D audio playback, thus expanding the application scenarios for 3D audio playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this embodiment of the application allows the system to obtain more information from the original 3D audio metadata and audio data compared to the rendered channel data, thereby improving the rendering effect.
[0063] 3D sound rendering process
[0064] In typical non-3D sound rendering scenarios, the system usually writes audio data into an audio track object, which is then rendered by the sound renderer. However, in 3D sound rendering scenarios, because it is necessary to reconstruct the spatial information of the sound object, such as its position, size, direction, and motion trajectory, the audio track object needs to be modified to enable it to carry both audio data and 3D sound data simultaneously.
[0065] In some embodiments, such as Figure 3 As shown, based on audio data and 3D sound data, sound rendering using a 3D sound renderer to obtain a 3D sound audio signal can include the following steps:
[0066] Step 203A: Write the audio data and 3D audio data into the audio track object.
[0067] In order for the subsequent 3D sound renderer to obtain the 3D sound metadata corresponding to the audio data when rendering the sound track object, the system needs to write the decoded audio data and the 3D sound metadata together into the sound track object, that is, it needs to extend the ability of the audio object to carry 3D sound metadata.
[0068] In some embodiments, applications can invoke the system-side 3D audio decoder for decoding in various ways. Correspondingly, the way the system writes audio data and 3D audio data to the audio track object differs under different invocation methods. The following describes the writing process of audio data and 3D audio data under different invocation methods.
[0069] Invocation method 1: The application invokes the system-side 3D sound decoder by sending a playback request.
[0070] In one possible implementation, when an application requires 3D audio / video playback, it invokes the system-side media player (MediaPlayer) interface by sending a playback request. During this invocation, the application transmits the 3D audio / video stream to the media player.
[0071] Accordingly, in response to the application's playback request, the terminal calls the 3D audio decoder through the system's built-in player to decode the 3D audio stream and obtain audio data and 3D audio metadata.
[0072] Optionally, if multiple decoders are set on the system side, when the playback request is detected to indicate the playback of 3D audio, the system's built-in player calls the 3D audio decoder to decode the 3D audio stream.
[0073] Furthermore, after obtaining the audio data and 3D audio metadata by calling the 3D audio decoder, the terminal writes the audio data and 3D audio metadata into the audio track object through the system's built-in player.
[0074] It should be noted that when it is necessary to play 3D sound and the corresponding video at the same time, the terminal also calls the video decoder through the system's built-in player to decode the video stream, which will not be described in detail here.
[0075] In an illustrative example, such as Figure 4 As shown, when an application at the application (APP) layer requires 3D audio playback (AudioPlay), it calls the MediaPlayer interface at the Java Native Interface (JNI) layer. In response to the call to the MediaPlayer interface, the terminal, through the system's built-in player (Nuplayer) at the Framework layer, calls the 3D Audio decoder (which also includes MP3, AAC, and other audio decoders) in the MediaCodec to decode the 3D audio stream, obtaining audio data and 3D audio metadata. The system's built-in player further writes the decoded audio data and 3D audio metadata into an AudioTrack object for subsequent 3D audio rendering.
[0076] Method 2: The application calls the system's 3D audio decoder by sending a decoding request.
[0077] In one possible implementation, when an application requires 3D audio / video playback, it invokes the system-side media codec interface by sending a decoding request, directly calling the system-side 3D audio decoder for decoding. During this call to the media codec interface, the application transmits the 3D audio / video bitstream to the media decoder.
[0078] In response to the application's decoding request, the terminal calls the 3D audio decoder to decode the 3D audio stream and obtain audio data and 3D audio metadata.
[0079] Optionally, if multiple decoders are set up on the system side, when the decoding request is detected to decode the three-dimensional audio and video stream, the terminal calls the three-dimensional audio decoder to decode the three-dimensional audio and video stream.
[0080] Since the application has only requested decoding so far, the decoded audio data and 3D acoustic data need to be sent back to the application so that the application can instruct on further processing.
[0081] When an application needs to further utilize the system's 3D sound rendering capabilities for 3D sound rendering, the application sends audio data and 3D sound data to the system's audio service system. Correspondingly, the terminal writes the audio data and 3D sound data sent by the application to an audio track object. For example, the application can write the audio data and 3D sound data to the audio track object using the `write` method.
[0082] It should be noted that when it is necessary to play 3D sound and the corresponding video at the same time, the terminal also calls the video decoder to decode the video stream, which will not be described in detail in this embodiment.
[0083] In an illustrative example, such as Figure 5 As shown, when an application at the application (APP) layer requires 3D audio playback (AudioPlay), it first calls the MediaCodec interface at the Java Native Interface (JNI) layer. In response to the call to the MediaCodec interface, the terminal calls the 3D Audio decoder (which also includes MP3, AAC, and other audio decoders) in the MediaCodec layer at the Framework layer to decode the 3D audio stream, obtaining audio data and 3D audio metadata, which are then sent back to the application. The application further writes the decoded audio data and 3D audio metadata into an AudioTrack object for subsequent 3D audio rendering.
[0084] Step 203B: Render the audio track object using a 3D sound renderer to obtain a 3D audio signal.
[0085] In some embodiments, the 3D sound renderer obtains audio data and corresponding 3D sound metadata from the audio track object, and then performs sound rendering on the audio data based on the 3D sound metadata to obtain a 3D sound audio signal.
[0086] In this embodiment, by modifying the data structure of the audio track object, the audio track object can carry both audio data and three-dimensional audio data simultaneously. When the three-dimensional sound renderer performs sound rendering on the audio track object, it can accurately obtain the audio data and its corresponding three-dimensional audio data, thereby ensuring the correct rendering of the three-dimensional audio.
[0087] In addition, for different calling methods of the application side calling the system side 3D audio decoder, specific data decoding and writing schemes are set to ensure the normal decoding of 3D audio data and the correct writing of data in the audio track object.
[0088] Considering that the speed at which the 3D audio decoder decodes and obtains the 3D audio metadata may not be consistent with the speed at which the audio track object consumes the 3D audio metadata, in one possible implementation, the 3D audio metadata is cached in the metadata buffer queue of the audio track object, while the audio data is cached in the audio data buffer of the audio track object.
[0089] In one possible implementation, when creating an audio track object, the system allocates a shared memory block (metadata buffer queue) for the reading and writing of 3D audio metadata, thereby enabling data transfer between processes (decoding process and rendering process). This shared memory can be in the form of a circular queue to ensure normal operation even when the buffer overflows.
[0090] In addition, the size of the shared memory can be determined by the system based on the processing time required for the three-dimensional audio data carried by the audio data within the same time period. For example, when processing 1920 frames of audio data with a period of 20ms, and carrying 2 three-dimensional audio data, the time required is also 20ms, then the size of the shared memory is the size of 2 three-dimensional audio data.
[0091] Of course, the terminal can also implement the metadata buffer queue in other ways, and this embodiment does not limit this.
[0092] The terminal then retrieves audio data and 3D audio metadata from the audio buffer and metadata buffer queue, respectively, and performs sound rendering on the retrieved audio data and 3D audio metadata using a 3D sound renderer. For example... Figure 6 As shown, the process may include the following steps:
[0093] Step 601: Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object.
[0094] Since 3D audio metadata may not be continuously effective during audio playback, but rather effective at a certain point in the audio playback process, or effective from a certain point in time during audio playback, it is necessary to determine the effective 3D audio metadata at the current moment when retrieving the 3D audio metadata for sound rendering from the metadata buffer queue.
[0095] In some embodiments, the 3D acoustic metadata written to the metadata buffer queue carries an effective timestamp, which is the timestamp when the 3D acoustic metadata begins to take effect.
[0096] In some embodiments, the effective timestamp is a timestamp starting from the audio playback start time (i.e., the duration relative to the audio playback start time). For example, when the effective timestamp is 10, it indicates that the three-dimensional acoustic data is effective from the 10th second after the audio playback starts, that is, from the 10th second after the audio playback starts, the subsequently played audio has the spatial effect represented by the three-dimensional acoustic data.
[0097] Optionally, the 3D audio metadata can also carry an expiration timestamp, which is the timestamp when the 3D audio metadata ceases to be effective. Similarly, this expiration timestamp is a timestamp starting from the audio playback start time.
[0098] Of course, 3D acoustic data may also not carry an expiration timestamp. For 3D acoustic data without an expiration timestamp, the 3D acoustic data will cease to be effective from the effective timestamp of the next 3D acoustic data.
[0099] To identify effective 3D audio metadata based on its effective timestamp, the audio track object in this embodiment maintains an object timestamp. This object timestamp is relative to the start time of audio playback, indicating the duration of audio playback; that is, the object timestamp is continuously updated during audio playback. Accordingly, the terminal determines the 3D audio metadata currently applied to sound rendering based on this object timestamp and the effective timestamp of the 3D audio metadata.
[0100] In some embodiments, the terminal obtains the effective 3D audio metadata from the metadata buffer queue of the audio track object based on the object timestamp maintained by the audio track object and the effective timestamp of the 3D audio metadata in the metadata buffer queue.
[0101] In one possible implementation, the terminal determines the effective 3D acoustic data as effective 3D acoustic data by comparing the effective timestamp and the object timestamp.
[0102] As an illustration, when the object timestamp maintained by the audio track object is 10, the terminal will determine the three-dimensional audio metadata with an effective timestamp of 10 in the metadata buffer queue as the effective three-dimensional audio metadata.
[0103] Step 602: Obtain audio data from the audio buffer of the audio track object.
[0104] In one possible implementation, the terminal retrieves the decoded audio data from the audio buffer in a first-in-first-out (FIFO) order.
[0105] It should be noted that there is no strict sequential order between steps 601 and 602 above, that is, steps 601 and 602 can be executed synchronously. This embodiment does not impose any restrictions on the execution order of the two.
[0106] Step 603: Use a 3D sound renderer to render the audio data and the effective 3D sound data to obtain a 3D sound audio signal.
[0107] The audio data and effective 3D audio metadata extracted from the audio buffer and metadata buffer queue are sent to the 3D audio renderer, which performs 3D audio rendering to obtain 3D audio signals.
[0108] In this embodiment, by caching the decoded 3D audio metadata to the metadata buffer queue of the audio track object, the decoding and consumption speed of the 3D audio metadata is kept consistent. Furthermore, the terminal system determines the effective 3D audio metadata based on the object timestamp maintained by the audio track object and the effective timestamp carried by the 3D audio metadata, and sends it to the 3D audio renderer for sound rendering, ensuring the accuracy of the timing of 3D audio rendering.
[0109] Metadata preprocessing
[0110] To ensure that the subsequent 3D sound renderer can correctly parse and use the 3D sound metadata, in one possible implementation, the terminal preprocesses the decoded 3D sound metadata before sound rendering, and then writes the preprocessed 3D sound metadata into the audio track object.
[0111] Since the 3D audio data carried in the 3D audio bitstream may have some missing metadata, and the 3D audio data may not conform to the specific protocol acquisition format specification, the preprocessing methods for 3D audio data include at least one of metadata completion and metadata conversion.
[0112] Metadata completion is used to supplement missing metadata in the 3D audio metadata. In one possible implementation, the terminal supplements the missing necessary metadata in the 3D audio metadata based on metadata completion rules. This necessary metadata is essential for the 3D audio rendering process.
[0113] In some embodiments, the 3D audio bitstream contains at least one complete 3D audio metadata. The terminal supplements the missing metadata in the currently decoded 3D audio metadata based on the complete 3D audio metadata and / or the 3D audio metadata obtained from previous decoding.
[0114] Metadata conversion is used to convert 3D audio-visual data to a specific data format. In some embodiments, if the 3D audio-visual data does not conform to a standard metadata format, the terminal system converts the 3D audio-visual data into a standard data format. This standard data format may include Dolby format, AudioVivid format, etc., and this embodiment does not limit this to any particular format.
[0115] Choosing a 3D sound renderer
[0116] In one possible implementation, the terminal's system side is equipped with multiple 3D sound renderers, each corresponding to a different 3D sound rendering method. These 3D sound rendering methods include speaker rendering, binaural rendering, and so on.
[0117] Accordingly, when rendering sound based on audio data and 3D audio data, the terminal renders the sound using a 3D audio renderer that matches the device type of the audio playback device, thus obtaining a 3D audio signal. This audio playback device can be the terminal's own playback device (such as a mobile phone speaker) or an external playback device (such as a speaker or headphones).
[0118] In one possible implementation, the terminal system has a mapping relationship between the device type of the audio playback device and the 3D sound renderer. The system then determines the 3D sound renderer that matches the device type of the current audio playback device based on this mapping relationship.
[0119] Indicative, such as Figure 7 As shown, the application sends the 3D audio stream 71 to the 3D audio decoder 72 on the system side, where the decoder 72 decodes the audio data and 3D audio metadata. After metadata preprocessing, the decoded 3D audio metadata, along with the audio data, is sent to the 3D audio renderer 73 on the system side for sound rendering. Specifically, when the audio playback device is a multi-channel speaker, the 3D audio renderer 73 performs VectorBase Amplitude Panning (VBA P) to obtain a multi-channel audio signal, which is then mixed by a multi-channel mixer before output. When the audio playback device is a headphone or speaker, the 3D audio renderer performs stereo rendering to obtain a stereo audio signal, which is then mixed by a stereo mixer before output. For speakers, crosstalk cancellation is performed before output; for headphones, head rotation compensation is performed based on head rotation.
[0120] 3D acoustic data generation
[0121] In one possible scenario, the application or system provides a sound object adjustment function. Using this function, users can adjust the spatial position of sound objects during 3D audio playback, or control the sound objects to move along a specific trajectory. To reproduce the adjusted 3D sound effect, in one possible implementation, in response to a metadata generation operation, the terminal generates 3D sound metadata and writes the audio data and the generated 3D sound metadata into the audio track object.
[0122] In some embodiments, the terminal displays a sound object adjustment interface, which includes a virtual space and sound object identifiers for each sound object located within that virtual space. Users can adjust the sound objects using these sound object identifiers.
[0123] The initial position of the acoustic object identifier in the virtual space is determined based on the spatial position information of the acoustic object in the initial three-dimensional acoustic metadata.
[0124] Indicative, such as Figure 8 As shown, the sound object adjustment interface includes a virtual control 81, and a first sound object identifier 82 for the first sound object, a second sound object identifier 83 for the second sound object, and a third sound object identifier 84 for the third sound object, all located in the virtual space 81.
[0125] Optionally, this metadata generation operation can be an adjustment operation for the sound object identifier. For example, adjusting the position of the sound object identifier, moving the sound object identifier along a movement trajectory, adjusting the size of the sound object, adjusting the direction of the sound object, etc.
[0126] Of course, in addition to adjusting the sound object identifier, new three-dimensional sound metadata can also be generated in other ways (i.e., the metadata generation operation can also be other types of operation), such as directly modifying the parameter values of the three-dimensional sound metadata of the sound object. This application embodiment does not limit this.
[0127] Accordingly, the three-dimensional acoustic data generated by the terminal may include the spatial location information of the adjusted acoustic object, the motion trajectory of the acoustic object, the volume of the adjusted acoustic object, the direction of the adjusted acoustic object, etc. This application embodiment does not limit the specific content included in the generated three-dimensional acoustic data.
[0128] Indicative, such as Figure 8 As shown, when the user drags the third sound object identifier 84 along a specific motion trajectory, the terminal generates three-dimensional sound data corresponding to the third sound object, which contains the motion trajectory of the third sound object.
[0129] After the terminal system writes the audio data and the generated 3D sound data into the audio track object, the 3D sound renderer on the system side renders the sound based on the audio track object, thus restoring the 3D sound effect after the sound object is adjusted.
[0130] The specific process of writing audio data and generated 3D sound data into the audio track object, and performing sound rendering based on the audio track object, can be referred to in the above embodiments, and will not be repeated here.
[0131] For example, such as Figure 8 As shown, the 3D sound renderer renders sound based on the 3D sound data containing the motion trajectory of the third sound object and the corresponding audio data of the third sound object, thus reproducing the acoustic effects produced when the third sound object moves along a specific motion trajectory in space.
[0132] To ensure the accuracy of the rendering timing for sound based on the generated 3D audio metadata, when generating the 3D audio metadata, the terminal determines the effective timestamp of the generated 3D audio metadata based on the object timestamp maintained by the audio track object. The object timestamp is a timestamp relative to the start time of audio playback, and the effective timestamp is the timestamp when the 3D audio metadata begins to take effect.
[0133] In some embodiments, when generating 3D audio metadata, the terminal determines the current object timestamp maintained by the audio track object as the effective timestamp of the generated 3D audio metadata.
[0134] In some embodiments, the terminal can also determine the expiration timestamp of the generated 3D audio data based on the object timestamp maintained by the audio track object. For example, when the audio object is moved along a specific motion trajectory, the terminal determines the object timestamp maintained by the audio track object at the start of the movement as the effective timestamp of the 3D audio data, and determines the object timestamp maintained by the audio track object at the end of the movement as the expiration timestamp of the 3D audio data. That is, the generated 3D audio data includes the effective timestamp, the expiration timestamp, and the motion trajectory.
[0135] In this embodiment, the terminal system has a metadata generation function, which can dynamically generate three-dimensional sound metadata according to the received metadata generation operation, and supports writing the generated three-dimensional sound metadata and audio data into the audio track object for the three-dimensional sound renderer to perform sound rendering, so that users can hear the changed three-dimensional sound effect in real time.
[0136] 3D sound recording
[0137] To facilitate subsequent reproduction, a 3D audio encoder can also be set on the system side. In audio recording mode, the terminal system performs 3D audio rendering based on the generated 3D audio metadata and audio data, and simultaneously encodes the audio data and the generated 3D audio metadata through the 3D audio encoder to obtain a 3D audio encoded bitstream.
[0138] The 3D acoustic encoder includes an audio encoder and a metadata encoder. The audio encoder is used to encode audio data, while the metadata encoder is used to encode 3D acoustic metadata. Furthermore, the audio encoder can be a general-purpose bitrate audio encoder or a lossless audio encoder; this embodiment does not limit this.
[0139] In some embodiments, the 3D acoustic encoded bitstream can be saved as a file and provided to the application.
[0140] Indicative, in Figure 7 On the basis of, such as Figure 9As shown, the application side is also equipped with a 3D sound encoder 74. In response to the metadata generation operation, the generated 3D sound metadata and the decoded audio data are sent together to the 3D sound renderer 73 for sound rendering; on the other hand, the generated 3D sound metadata and audio data are copied (the copied 3D sound metadata and audio data are not processed by sound rendering) and sent to the 3D sound encoder 74 for encoding.
[0141] Besides adjusting the 3D audio data in real time during decoding and playback, other possible methods include... Figure 9 As shown, PCM audio data can also be provided by the application, and the system can generate three-dimensional acoustic metadata for the PCM audio data (the user can manually adjust the sound object to trigger the generation of three-dimensional acoustic metadata), and then send the PCM audio data and the generated three-dimensional acoustic metadata together into the three-dimensional sound encoder 74 for encoding.
[0142] In this embodiment, the terminal system provides a three-dimensional sound coding function. With the help of this function, the terminal system can encode the real-time generated three-dimensional sound metadata and audio data, which facilitates the subsequent three-dimensional sound reproduction based on the three-dimensional sound coding bitrate obtained by encoding.
[0143] Please refer to Figure 10 This illustration shows a structural block diagram of a three-dimensional audio-visual processing apparatus provided in an exemplary embodiment of this application. The apparatus includes:
[0144] The acquisition module 1001 is used to acquire the three-dimensional audio and video bitstreams issued by the application;
[0145] Decoding module 1002 is used to decode the three-dimensional audio stream using a three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio data.
[0146] The rendering module 1003 is used to render sound using a 3D sound renderer based on the audio data and the 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound decoder and the 3D sound renderer are set in the system.
[0147] Optionally, the rendering module 1003 is used for:
[0148] Write the audio data and the three-dimensional acoustic data into the audio track object;
[0149] The audio track object is rendered using the 3D audio renderer to obtain the 3D audio signal.
[0150] Optionally, the decoding module 1002 is used for:
[0151] In response to the application's playback request, the system's built-in player calls the 3D audio decoder to decode the 3D audio stream, obtaining the audio data and the 3D audio metadata.
[0152] The rendering module 1003 is used for:
[0153] The audio data and the three-dimensional acoustic data are written into the audio track object through the system's built-in player.
[0154] Optionally, the decoding module 1002 is used for:
[0155] In response to the decoding request of the application, the three-dimensional sound decoder is invoked to decode the three-dimensional sound audio stream to obtain the audio data and the three-dimensional sound data, wherein the decoded audio data and the three-dimensional sound data are sent back to the application;
[0156] The rendering module 1003 is used for:
[0157] The audio data and the three-dimensional audio data sent by the application are written into the audio track object.
[0158] Optionally, the three-dimensional acoustic metadata is cached in the metadata buffer queue of the audio track object, and the audio data is cached in the audio data buffer of the audio track object.
[0159] Optionally, the rendering module 1003 is used for:
[0160] Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object;
[0161] The audio data is obtained from the audio buffer of the audio track object;
[0162] The audio data and the effective 3D audio data are rendered using the 3D audio renderer to obtain the 3D audio signal.
[0163] Optionally, during the process of retrieving effective 3D audio metadata from the metadata buffer queue of the audio track object, the rendering module 1003 is configured to:
[0164] Based on the object timestamp maintained by the audio track object and the effective timestamp of the three-dimensional audio data in the metadata buffer queue, the effective three-dimensional audio data is obtained from the metadata buffer queue of the audio track object. The object timestamp is a timestamp relative to the audio playback start time, and the effective timestamp is the timestamp when the three-dimensional audio data begins to take effect.
[0165] Optionally, the device further includes:
[0166] The generation module is used to generate 3D acoustic metadata in response to metadata generation operations;
[0167] The rendering module 1003 is also used to write the audio data and the generated three-dimensional audio data into the audio track object.
[0168] Optionally, the generation module is used for:
[0169] Based on the object timestamp maintained by the audio track object, the effective timestamp of the generated three-dimensional audio data is determined. The object timestamp is a timestamp relative to the start time of audio playback, and the effective timestamp is the timestamp when the three-dimensional audio data begins to take effect.
[0170] Optionally, the system further includes a three-dimensional acoustic encoder, and the device further includes:
[0171] The encoding module is used to encode the audio data and the generated three-dimensional acoustic data through the three-dimensional acoustic encoder in the audio recording state to obtain a three-dimensional acoustic encoded bitstream.
[0172] Optionally, the device further includes:
[0173] The preprocessing module is used to preprocess the decoded three-dimensional acoustic data, wherein the preprocessing method includes at least one of metadata completion and metadata conversion.
[0174] Optionally, the rendering module 1003 is further configured to:
[0175] Based on the audio data and the three-dimensional sound data, the sound is rendered by the three-dimensional sound renderer that matches the device type of the audio playback device to obtain the three-dimensional sound audio signal.
[0176] In summary, in this embodiment of the application, to achieve system-level 3D audio decoding and rendering, a 3D audio decoder and a 3D audio renderer are configured in the terminal system. The system decodes the 3D audio stream sent by the application using the 3D audio decoder, and renders the decoded audio data and 3D audio metadata using the 3D audio renderer to obtain a 3D audio signal. Since both decoding and rendering are completed on the system side, the system can provide 3D audio playback functionality even if the application does not natively support 3D audio playback, thus expanding the application scenarios for 3D audio playback. Furthermore, compared to the application submitting the rendered channel data to the system for rendering and playback, the solution provided in this embodiment of the application allows the system to obtain more information from the original 3D audio metadata and audio data compared to the rendered channel data, thereby improving the rendering effect.
[0177] It should be noted that the apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.
[0178] See Figure 11 , Figure 11 This is a schematic diagram of the structure of a terminal provided in an exemplary embodiment of this application. The terminal may also include one or more of the following components: a processor 1110 and a memory 1120.
[0179] Optionally, the processor 1110 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1120, and by calling data stored in the memory 1120. Optionally, the processor 1110 can be implemented in at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA).
[0180] The processor 1110 can integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), and a baseband chip. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content displayed on the touchscreen; the NPU implements artificial intelligence (AI) functions; and the baseband chip handles wireless communication. It is understood that the baseband chip can also be implemented as a separate chip without being integrated into the processor 1110.
[0181] The memory 1120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1120 may include a non-transitory computer-readable storage medium. The memory 1120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created according to the use of the electronic device, etc.
[0182] In addition, those skilled in the art will understand that the structure of the terminal shown in the above figures does not constitute a limitation on the terminal. The terminal may include more (e.g., power supply components, display components, sensor components) or fewer components than shown, or combine certain components, or have different component arrangements.
[0183] This application provides a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the three-dimensional audio-visual processing method as described in the above embodiments.
[0184] On the other hand, embodiments of this application provide a computer program product, the computer program product including computer instructions stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the three-dimensional audio-visual processing method as described in the above embodiments.
[0185] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0186] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A three-dimensional audio-visual processing method, characterized in that, The method is executed by a terminal system, the system being equipped with a 3D sound decoder and a 3D sound renderer, and the method includes: Obtain the 3D audio and video streams issued by the application; The three-dimensional audio stream is decoded by the three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio data. Based on the audio data and the three-dimensional sound data, the sound is rendered using the three-dimensional sound renderer to obtain a three-dimensional sound audio signal.
2. The method according to claim 1, characterized in that, The process of rendering sound using the 3D sound renderer based on the audio data and the 3D sound data to obtain a 3D sound audio signal includes: Write the audio data and the three-dimensional acoustic data into the audio track object; The audio track object is rendered using the 3D audio renderer to obtain the 3D audio signal.
3. The method according to claim 2, characterized in that, The step of decoding the three-dimensional audio stream using the three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio metadata includes: In response to the application's playback request, the system's built-in player calls the 3D audio decoder to decode the 3D audio stream, obtaining the audio data and the 3D audio metadata. The step of writing the audio data and the three-dimensional acoustic data into the audio track object includes: The audio data and the three-dimensional acoustic data are written into the audio track object through the system's built-in player.
4. The method according to claim 2, characterized in that, The step of decoding the three-dimensional audio stream using the three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio metadata includes: In response to the decoding request of the application, the three-dimensional sound decoder is invoked to decode the three-dimensional sound audio stream to obtain the audio data and the three-dimensional sound data, wherein the decoded audio data and the three-dimensional sound data are sent back to the application; The step of writing the audio data and the three-dimensional acoustic data into the audio track object includes: The audio data and the three-dimensional audio data sent by the application are written into the audio track object.
5. The method according to claim 2, characterized in that, The three-dimensional audio metadata is cached in the metadata buffer queue of the audio track object, and the audio data is cached in the audio data buffer of the audio track object.
6. The method according to claim 5, characterized in that, The step of rendering the audio track object using the 3D audio renderer to obtain the 3D audio signal includes: Retrieve the effective 3D audio metadata from the metadata buffer queue of the audio track object; The audio data is obtained from the audio buffer of the audio track object; The audio data and the effective 3D audio data are rendered using the 3D audio renderer to obtain the 3D audio signal.
7. The method according to claim 6, characterized in that, The step of retrieving effective 3D audio metadata from the metadata buffer queue of the audio track object includes: Based on the object timestamp maintained by the audio track object and the effective timestamp of the three-dimensional audio data in the metadata buffer queue, the effective three-dimensional audio data is obtained from the metadata buffer queue of the audio track object. The object timestamp is a timestamp relative to the audio playback start time, and the effective timestamp is the timestamp when the three-dimensional audio data begins to take effect.
8. The method according to claim 2, characterized in that, The method further includes: In response to the metadata generation operation, generate 3D acoustic metadata; The audio data and the generated 3D acoustic data are written into the audio track object.
9. The method according to claim 8, characterized in that, The generation of three-dimensional acoustic metadata includes: Based on the object timestamp maintained by the audio track object, the effective timestamp of the generated three-dimensional audio data is determined. The object timestamp is a timestamp relative to the start time of audio playback, and the effective timestamp is the timestamp when the three-dimensional audio data begins to take effect.
10. The method according to claim 8, characterized in that, The system is also equipped with a three-dimensional acoustic encoder, and the method further includes: In audio recording mode, the audio data and the generated three-dimensional acoustic data are encoded by the three-dimensional acoustic encoder to obtain a three-dimensional acoustic encoded bitstream.
11. The method according to any one of claims 1 to 10, characterized in that, Before the step of rendering sound using the 3D sound renderer based on the audio data and the 3D sound data to obtain the 3D sound audio signal, the method further includes: The decoded 3D acoustic data is preprocessed, wherein the preprocessing method includes at least one of metadata completion and metadata conversion.
12. The method according to any one of claims 1 to 10, characterized in that, The process of rendering sound using the 3D sound renderer based on the audio data and the 3D sound data to obtain a 3D sound audio signal includes: Based on the audio data and the three-dimensional sound data, the sound is rendered by the three-dimensional sound renderer that matches the device type of the audio playback device to obtain the three-dimensional sound audio signal.
13. A three-dimensional audio-visual processing device, characterized in that, The device includes: The acquisition module is used to acquire the 3D audio and video bitstreams issued by the application; The decoding module is used to decode the three-dimensional audio stream using a three-dimensional audio decoder to obtain audio data and the corresponding three-dimensional audio data. The rendering module is used to render sound using a 3D sound renderer based on the audio data and the 3D sound data to obtain a 3D sound audio signal, wherein the 3D sound decoder and the 3D sound renderer are set in the system.
14. A terminal, characterized in that, The terminal includes a processor and a memory, the memory storing at least one computer instruction, which is loaded and executed by the processor to implement the three-dimensional audio-visual processing method as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer instruction, which is executed by a processor to implement the three-dimensional audio-visual processing method as described in any one of claims 1 to 12.
16. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the three-dimensional audio-visual processing method as described in any one of claims 1 to 12.