Audio rendering methods, devices, media and electronic equipment
By configuring the receiver field in the metadata, audio data is rendered differently based on information such as the receiver's location, device, and rendering type. This solves the problem that existing technologies cannot meet diverse audio playback needs and achieves personalized audio rendering effects.
Patent Information
- Application Number
- CN202211297361.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing audio rendering technologies cannot meet the diverse audio playback needs in complex scenarios, and users cannot obtain personalized auditory effects when they are in different positions, postures, or using different devices.
By configuring multiple receiver fields in the metadata corresponding to multiple receivers, audio data can be rendered differently based on the receiver's location information, device information, rendering type, and occlusion object information to meet the needs of different receivers.
It enables multiple receivers to obtain different auditory effects from the same audio data, meeting the needs of diverse audio rendering scenarios and providing personalized audio rendering solutions.
Smart Images

Figure CN115696137B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio and video technology, specifically relating to an audio rendering method, an audio rendering device, a computer-readable medium, an electronic device, and a computer program product. Background Technology
[0002] When playing audio to users in different scenarios, different audio playback effects are usually produced. For example, for the same audio source, users should be able to obtain different auditory effects when they are in different positions, using different postures, or using different playback devices. However, related audio rendering technologies cannot meet the diverse audio playback needs in complex scenarios. Summary of the Invention
[0003] This application provides an audio rendering method, an audio rendering apparatus, a computer-readable medium, an electronic device, and a computer program product, with the aim of improving the diverse rendering effects of audio.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0006] Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0007] When there are multiple receivers, multiple receiver fields corresponding to the multiple receivers are configured in the metadata according to the number information. The receiver fields include feature fields related to the location information of the receivers.
[0008] The metadata is sent to an audio renderer, which renders the audio data differently for each of the multiple recipient fields.
[0009] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0010] The first acquisition module is configured to acquire the number of receivers, the number of receivers being used to indicate the number of receivers receiving audio data.
[0011] The first assignment module is configured to, when there are multiple receivers, configure multiple receiver fields in the metadata according to the quantity information, each corresponding to one of the multiple receivers. The receiver fields include feature fields related to the location information of the receivers.
[0012] The first sending module is configured to send the metadata to an audio renderer, which is used to render the audio data differently for multiple recipients based on multiple recipient fields.
[0013] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0014] Extract the receiver field corresponding to the receiver receiving the audio data from the metadata. The receiver field includes feature fields related to the location information of the receiver.
[0015] If there are multiple receivers, the audio data is rendered differently for each receiver based on the multiple receiver fields.
[0016] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0017] The first extraction module is configured to extract a receiver field from metadata corresponding to the receiver receiving the audio data. The receiver field includes feature fields related to the location information of the receiver.
[0018] The first rendering module is configured to render the audio data differently for each of the multiple receivers based on the multiple receiver fields if there are multiple receivers.
[0019] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0020] Obtain the receiver's device information, which indicates the audio playback device used by the receiver when receiving audio data;
[0021] The device information field in the metadata is assigned a value based on the device information, and the device information field is used to indicate the characteristics of the audio playback device that affect the audio data rendering effect;
[0022] The corresponding device identifier field in the metadata is assigned a value based on the device information field. The corresponding device identifier field is a sub-element of the rendering information field, which is an element in the metadata used to indicate the rendering effect of the audio data.
[0023] The metadata is sent to the audio renderer, which renders the audio data for the audio playback device based on the device information field.
[0024] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0025] The second acquisition module is configured to acquire the device information of the receiver, wherein the device information is used to indicate the audio playback device used by the receiver when receiving audio data;
[0026] The second assignment module is configured to assign a value to the device information field in the metadata based on the device information, wherein the device information field is used to indicate the characteristics of the audio playback device that affect the audio data rendering effect; and to assign a value to the corresponding device identifier field in the metadata based on the device information field, wherein the corresponding device identifier field is a sub-element of the rendering information field, and the rendering information field is an element in the metadata used to indicate the audio data rendering effect;
[0027] The second sending module is configured to send the metadata to the audio renderer, which is used to render the audio data for the audio playback device based on the device information field.
[0028] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0029] Extract the corresponding device identifier field from the rendering information field of the metadata, where the rendering information field is an element in the metadata used to indicate the rendering effect of the audio data;
[0030] The device information field is extracted based on the corresponding device identifier field. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect.
[0031] The audio data is rendered for the audio playback device based on the device information field.
[0032] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0033] The second extraction module is configured to extract a corresponding device identifier field from the rendering information field of the metadata, wherein the rendering information field is an element in the metadata used to indicate the audio data rendering effect; and to extract a device information field based on the corresponding device identifier field, wherein the device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect.
[0034] The second rendering module is configured to render the audio data for the audio playback device based on the device information field.
[0035] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0036] Obtain rendering type information of the audio data, wherein the rendering type information is used to indicate the rendering type corresponding to each audio component in the audio data;
[0037] The audio component information field in the metadata is assigned a value according to the rendering type information. The audio component information field is used to indicate the audio component with the corresponding rendering type that is processed by the audio renderer.
[0038] The metadata is sent to the audio renderer, which renders the audio component according to the rendering type.
[0039] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0040] The third acquisition module is configured to acquire rendering type information of audio data, wherein the rendering type information is used to indicate the rendering type corresponding to each audio component in the audio data;
[0041] The third assignment module is configured to assign a value to the audio component information field in the metadata according to the rendering type information. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer.
[0042] The third sending module is configured to send the metadata to the audio renderer, which is used to render the audio component according to the rendering type.
[0043] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0044] Extract the audio component information field from the metadata. The audio component information field is used to indicate the audio component with the corresponding rendering type that is processed by the audio renderer.
[0045] The rendering type corresponding to each audio component in the audio data is determined based on the audio component information field.
[0046] The audio component is rendered according to the rendering type.
[0047] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0048] The third extraction module is configured to extract an audio component information field from metadata, the audio component information field being used to indicate an audio component with a corresponding rendering type processed by the audio renderer;
[0049] The third rendering module is configured to determine the rendering type corresponding to each audio component in the audio data based on the audio component information field; and to render the audio component according to the rendering type.
[0050] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0051] Obtain the attribute information of occluded objects in the rendering scene of the audio data;
[0052] The occlusion object information field in the metadata is assigned a value based on the attribute information of the occluding object. The occlusion object information field is used to indicate the characteristics of the occluding object that produces the audio occlusion effect.
[0053] The metadata is sent to the audio renderer, which renders audio data with the audio occlusion effect based on the occlusion object information field.
[0054] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0055] The fourth acquisition module is configured to acquire attribute information of occluded objects in the rendering scene of the audio data;
[0056] The fourth assignment module is configured to assign a value to the occlusion object information field in the metadata based on the attribute information of the occluding object. The occlusion object information field is used to indicate the characteristics of the occluding object that produces the audio occlusion effect.
[0057] The fourth sending module is configured to send the metadata to the audio renderer, which is used to render audio data with the audio occlusion effect based on the occlusion object information field.
[0058] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0059] Extract the occlusion object information field from the metadata, which is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect;
[0060] Render audio data with the audio occlusion effect based on the occlusion object information field.
[0061] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0062] The fourth extraction module is configured to extract the occlusion object information field from the metadata, the occlusion object information field being used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0063] The fourth rendering module is configured to render audio data with the audio occlusion effect based on the occlusion object information field.
[0064] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0065] Obtain the scene configuration information of the receiver, wherein the scene configuration information is the configuration information that affects the rendering effect of audio data in the rendering scene;
[0066] The metadata feature fields are assigned values according to the scene configuration information. The feature fields include at least one of the following: receiver field, device information field, audio component information field, or occlusion object information field. The receiver field includes a field related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0067] The metadata is sent to an audio renderer, which renders the audio data based on the feature fields.
[0068] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0069] The fifth acquisition module is configured to acquire the scene configuration information of the receiver, which is the configuration information that affects the rendering effect of audio data in the rendering scene;
[0070] The fifth assignment module is configured to assign values to the feature fields of the metadata according to the scene configuration information. The feature fields include at least one of the following: receiver field, device information field, audio component information field, or occlusion object information field. The receiver field includes a field related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0071] The fifth sending module is configured to send the metadata to an audio renderer, which renders the audio data based on the feature fields.
[0072] According to one aspect of the embodiments of this application, an audio rendering method is provided, including:
[0073] Feature fields are extracted from metadata. The feature fields include at least one of the following: receiver field, device information field, audio component information field, or occlusion object information field. The receiver field includes fields related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0074] Render audio data based on the aforementioned feature fields.
[0075] According to one aspect of the embodiments of this application, an audio rendering apparatus is provided, comprising:
[0076] The fifth extraction module is configured to extract feature fields from metadata. The feature fields include at least one of a receiver field, a device information field, an audio component information field, or an occlusion object information field. The receiver field includes a field related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with a corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0077] The fifth rendering module is configured to render audio data based on the feature fields.
[0078] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the audio rendering method as described in the above technical solutions.
[0079] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform an audio rendering method as described above by executing the executable instructions.
[0080] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program, which, when executed by a processor, performs the audio rendering method as described in the above technical solutions.
[0081] In the technical solution provided in this application embodiment, by configuring multiple receiver fields corresponding to multiple receivers in the metadata, personalized audio rendering can be performed for multiple receivers according to the multiple receiver fields, so that multiple receivers can obtain different auditory effects from the same audio data, thus meeting the diverse audio rendering scenario requirements.
[0082] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0083] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0084] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0085] Figure 2 This schematically illustrates the placement of audio / video encoding and decoding devices in a streaming environment.
[0086] Figure 3 This schematically illustrates a system framework for realizing audio content expression in virtual reality application scenarios.
[0087] Figure 4 A flowchart illustrating the steps of an audio acquisition terminal performing audio rendering based on the number of receivers in one embodiment of this application is shown.
[0088] Figure 5 A flowchart illustrating the steps of audio rendering based on receiver degrees of freedom in one embodiment of this application is shown.
[0089] Figure 6 A flowchart illustrating the steps of audio rendering based on the receiver identifier in one embodiment of this application is shown.
[0090] Figure 7 A flowchart illustrating the steps of audio rendering based on coordinate system type in one embodiment of this application is shown.
[0091] Figure 8 A flowchart illustrating the steps of audio rendering based on a receiver at the same location in one embodiment of this application is shown.
[0092] Figure 9A flowchart illustrating the steps of a method for audio rendering based on the number of receivers performed by an audio receiver in one embodiment of this application is shown.
[0093] Figure 10 A flowchart illustrating the steps of a method for audio rendering based on device information performed by an audio acquisition terminal in one embodiment of this application is shown.
[0094] Figure 11 A flowchart illustrating the steps of a method for audio rendering based on device information performed by an audio receiver in one embodiment of this application is shown.
[0095] Figure 12 The flowchart illustrates the steps of a method for audio rendering based on rendering type performed by the audio acquisition terminal in one embodiment of this application.
[0096] Figure 13 The flowchart illustrates the steps of a method for audio rendering based on rendering type performed by an audio receiver in one embodiment of this application.
[0097] Figure 14 The flowchart illustrates the steps of a method for audio rendering based on occlusion object information executed by the audio acquisition terminal in one embodiment of this application.
[0098] Figure 15 The flowchart illustrates the steps of a method for audio rendering based on occlusion parameters performed by an audio acquisition terminal in one embodiment of this application.
[0099] Figure 16 A flowchart illustrating the steps of a method for audio rendering based on occlusion object information performed by an audio receiver in one embodiment of this application is shown.
[0100] Figure 17 The flowchart illustrates the steps of a method for audio rendering based on scene configuration information executed by the audio acquisition terminal in one embodiment of this application.
[0101] Figure 18 A flowchart illustrating the steps of a method for audio rendering based on scene configuration information performed by an audio receiver in one embodiment of this application is shown.
[0102] Figure 19 A schematic diagram of a computer system architecture suitable for implementing the embodiments of this application is shown. Detailed Implementation
[0103] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0104] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0105] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0106] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0107] In the specific embodiments of this application, data related to user location information, posture information, device information, etc. are involved. When the various embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0108] The relevant terms or abbreviations used in the embodiments of this application are explained as follows.
[0109] Panoramic audio: 720-degree omnidirectional audio in three-dimensional space, providing an immersive auditory experience.
[0110] Immersive audio: Audio in a three-dimensional space that creates an immersive experience. It is often used in immersive communication and as a way to express panoramic or three-dimensional audio in immersive systems.
[0111] Virtual reality audio: A virtual representation of audio in the real world. It uses rendering technology to play audio signals in a virtual form through binaural or speaker playback to achieve the feeling of audio in real space. It is often used in virtual reality devices and systems to describe three-dimensional audio.
[0112] Metadata: Describes characteristic information related to the rendering scene and audio content.
[0113] ADM (Audio Definition Model): An audio metadata standard used to describe the components of an audio file.
[0114] HRTF (Head Related Transfer Function): The frequency domain acoustic transfer function from a sound source to both ears in a free field.
[0115] HRIR (Head-related impulse response): The impulse response from the sound source to both ears in a free field, which is the time-domain equivalent of HRTF.
[0116] HOA (Higher Order Ambisonic): A higher-order spherical harmonic function signal.
[0117] DoF (Degree of Freedom): refers to the degree of freedom a user has to move and interact with content while watching immersive media.
[0118] 3DoF: refers to three degrees of freedom, which means the user's head rotates around the x, y, and z axes.
[0119] 3DoF+: In addition to the three degrees of freedom, users also have finite degrees of freedom to move along the x, y, and z axes.
[0120] 6DoF: In addition to the three degrees of freedom, users also have the freedom to move freely along the x, y, and z axes.
[0121] AVS: Audio Video Coding Standard, China's national video coding standard.
[0122] The following is passed Figures 1 to 3 This paper introduces the system architectures of the technical solutions applied in this application in different application scenarios. Figure 1 This refers to a system architecture that enables audio and video interaction and transmission between multiple terminal devices, such as scenarios involving audio and video conferencing and audio and video calls. Figure 2 This refers to the system architecture that transmits audio and video streams from the acquisition end to the receiving end, such as in scenarios involving live web broadcasting, live television broadcasting, and radio broadcasting. Figure 3 This refers to a system architecture that decouples audio signals from metadata during transmission, such as in virtual reality applications.
[0123] Figure 1A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0124] like Figure 1 As shown, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. Figure 1 In one embodiment, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0125] For example, the first terminal device 110 can encode audio and video data (e.g., audio and video data streams collected by the terminal device 110) to transmit to the second terminal device 120 via the network 150. The encoded audio and video data is transmitted in the form of one or more encoded audio and video streams. The second terminal device 120 can receive the encoded audio and video data from the network 150, decode the encoded audio and video data to recover the audio and video data, and play or display the content based on the recovered audio and video data.
[0126] In one embodiment of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded audio and video data, such as during an audio-visual conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode audio and video data (e.g., an audio and video data stream acquired by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded audio and video data transmitted by the other terminal device, decode the encoded audio and video data to recover the audio and video data, and play or display content based on the recovered audio and video data.
[0127] exist Figure 1In the embodiments disclosed herein, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited to these. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated audio and video conferencing equipment. Network 150 refers to any number of networks that transmit encoded audio and video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of this application.
[0128] In one embodiment of this application, Figure 2 This illustration schematically shows the placement of audio / video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other audio / video supported applications, including, for example, audio / video conferencing, digital television (television), and storing compressed audio / video on digital media including CDs, DVDs, memory sticks, etc.
[0129] The streaming system may include an acquisition subsystem 213, which may include audio / video sources 201 such as microphones and cameras, which create an uncompressed audio / video data stream 202. The audio / video data stream 202 is depicted as a thick line to emphasize its high data volume compared to encoded audio / video data 204 (or encoded audio / video bitstream 204). The audio / video data stream 202 may be processed by an electronic device 220, which includes an audio / video encoding device 203 coupled to the audio / video source 201. The audio / video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the audio / video data stream 202, the encoded audio / video data 204 (or encoded audio / video bitstream 204) is depicted as thin lines to emphasize the lower data volume of the encoded audio / video data 204 (or encoded audio / video bitstream 204), which can be stored on the streaming server 205 for future use. One or more streaming client subsystems, such as... Figure 2Client subsystems 206 and 208 can access streaming server 205 to retrieve copies 207 and 209 of encoded audio and video data 204. Client subsystem 206 may include, for example, an audio / video decoding device 210 in electronic device 230. Audio / video decoding device 210 decodes the incoming copy 207 of the encoded audio and video data and produces an output audio / video data stream 211 that can be presented on output device 212 (e.g., a speaker, a display) or another presentation device. In some streaming systems, the encoded audio / video data 204, audio / video data 207, and audio / video data 209 (e.g., audio / video streams) may be encoded according to certain audio / video encoding / compression standards.
[0130] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include an audio / video decoding device, and electronic device 230 may also include an audio / video encoding device.
[0131] Figure 3 This schematically illustrates a system framework for realizing audio content expression in virtual reality application scenarios.
[0132] Virtual reality audio content expression broadly involves metadata, renderers, audio encoders, and audio decoders. In this application's embodiments, metadata, renderers, encoders, and decoders can be logically separated. When used for local storage and production, only the renderer needs to parse the metadata, without involving the audio encoding / decoding process; when used for transmission (e.g., live streaming or two-way communication), the transmission format of the metadata and audio stream needs to be defined.
[0133] like Figure 3 As shown, in the virtual reality audio content expression framework, the acquisition end uses input audio signals including channel, object, hoa, or a mixture thereof. Metadata information is generated based on the metadata definition. Dynamic metadata can be encoded and transmitted along with the audio stream; the specific encapsulation format is defined according to the transmission protocol type adopted by the system layer. At the playback end, the renderer renders and outputs the decoded audio file based on the decoded metadata. Logically, metadata and audio encoding / decoding are independent, and the decoder and renderer are decoupled. The renderer adopts a registration system (ID1: renderer based on binaural output, ID2: renderer based on speaker output, ID3: other methods, ID4: other methods), and each registered renderer supports the same set of metadata definitions.
[0134] by Figure 3 The scenario shown is an example. Performing the audio rendering method proposed in the embodiments of this application may include the following steps.
[0135] 1. The server generates corresponding audio metadata files based on information such as user feedback regarding location, device information, and audio content.
[0136] 2. The server sends the audio metadata file to the audio renderer.
[0137] a) If the renderer is a cloud renderer, then the renderer is located on the server side.
[0138] b) If the renderer is a user-local renderer, then the renderer is located on the client side.
[0139] 3. The audio renderer extracts location and device information from the metadata file as rendering input parameters to calculate the rendering effect for the corresponding user and device. It also extracts object-related location information and occlusion algorithm parameters to render the corresponding sound occlusion effect.
[0140] 4. Present the audio on the corresponding user's corresponding device.
[0141] a) If the renderer is a cloud renderer, the server sends the rendered audio to the client, and the client then renders it.
[0142] b) If the renderer is the user's local renderer, the client renders it directly.
[0143] The following combination Figures 4 to 10 This application provides a detailed description of the audio rendering method for multi-user scenarios in its embodiments. Figures 4 to 8 This application's embodiment describes an audio rendering method for multi-user scenarios executed at an audio acquisition terminal. The audio acquisition terminal may be, for example, a... Figure 1 The terminal device shown Figure 2 The acquisition subsystem shown or Figure 3 The data acquisition terminal shown. Figure 9 This application's embodiment describes an audio rendering method for multi-user scenarios executed at an audio receiving end. The audio receiving end may be, for example, a... Figure 1 The terminal device shown Figure 2 The client subsystem shown or Figure 3 The registered renderer is shown.
[0144] Figure 4 This document illustrates a flowchart of the method steps for audio rendering based on the number of receivers performed by the audio acquisition terminal in one embodiment of this application. Figure 4 As shown, the audio rendering method involving a multi-user scenario executed at the audio acquisition end includes the following steps S410 to S430.
[0145] S410: Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0146] S420: When there are multiple receivers, configure multiple receiver fields in the metadata according to the number information, each corresponding to a different receiver. The receiver fields include feature fields related to the location information of the receivers.
[0147] S430: Sends metadata to the audio renderer, which renders audio data differently for multiple recipients based on multiple recipient fields.
[0148] This application embodiment configures multiple receiver fields corresponding to multiple receivers in the metadata, and can perform personalized audio rendering for multiple receivers according to the multiple receiver fields, so that multiple receivers can obtain different auditory effects from the same audio data, meeting the needs of diverse audio rendering scenarios.
[0149] In one embodiment of this application, after obtaining the number of recipients, a value can also be assigned to the multi-recipient flag field in the metadata based on the number information. The multi-recipient flag field is used to indicate that the number of recipients is one or more.
[0150] In one embodiment of this application, the multi-listener flag field can be an element in the root metadata. As shown in Table 1, in the VR extended root metadata implemented using the embodiments of this application, there exists a multi-listener flag, multi_listener. Based on different values of this field, differentiated audio rendering can be determined according to the number of receivers. For example, when the multi_listener flag is 1, it indicates that the current renderer needs to perform personalized audio rendering for multiple receivers; when the multi_listener flag is 0, it indicates that the current renderer does not need to perform personalized audio rendering for multiple receivers.
[0151] Table 1. Relevant Specifications for VR Extended Root Metadata Including Multi-Receiver Flag Field
[0152]
[0153] In one embodiment of this application, the metadata includes a presentation information field, which is an element related to the content of the audio data, and a multi-receiver flag field, which is an attribute or sub-element of the presentation information field. As shown in Table 1, the metadata includes a presentation information field called presenceInfo, which is extended metadata related to the content of the audio data.
[0154] In one embodiment of this application, the multi-listener flag field is an attribute of the presentation information field. As shown in Table 2, the presentation information field `presenceInfo` contains a multi-listener flag `multi_listener`. Based on different values of this field, differentiated audio rendering can be determined according to the number of listeners. For example, when the multi-listener flag `multi_listener` is 1, it indicates that the current renderer needs to perform personalized audio rendering for multiple listeners; when the multi-listener flag `multi_listener` is 0, it indicates that the current renderer does not need to perform personalized audio rendering for multiple listeners.
[0155] Table 2. Specifications for attributes in the presenceInfo field.
[0156]
[0157]
[0158] In one embodiment of this application, the multi-receiver flag field is a sub-element of the presentation information field. As shown in Table 3, the presentation information field `presenceInfo` contains a multi-receiver flag `multi_listener`. Based on different values of this field, differentiated audio rendering can be determined according to the number of receivers. For example, when the multi-receiver flag `multi_listener` is 1, it indicates that the current renderer needs to perform personalized audio rendering for multiple receivers; when the multi-receiver flag `multi_listener` is 0, it indicates that the current renderer does not need to perform personalized audio rendering for multiple receivers.
[0159] Table 3. Specifications for sub-elements in the presenceInfo field.
[0160]
[0161] In one embodiment of this application, before sending the metadata to the audio renderer, a value can be assigned to the receiver's degrees of freedom field in the metadata, thereby determining the degrees of freedom of the corresponding receiver to support movement and generate content interaction when watching immersive media based on the receiver's degrees of freedom field.
[0162] Figure 5 A flowchart illustrating the steps of audio rendering based on receiver degrees of freedom in one embodiment of this application is shown. Figure 5 As shown, the method for audio rendering based on metadata carrying the receiver's degrees of freedom field includes the following steps S510 to S550.
[0163] S510: Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0164] S520: When there are multiple receivers, configure multiple receiver fields in the metadata according to the number information, each corresponding to a different receiver. The receiver fields include feature fields related to the location information of the receivers.
[0165] S530: Obtain the receiver's degree of freedom information, which is used to indicate the receiver's degree of freedom when receiving audio;
[0166] S540: Assign a value to the receiver's degrees of freedom field in the metadata based on the degrees of freedom information. The receiver's degrees of freedom field is used to indicate whether the receiver has three or six degrees of freedom.
[0167] S550: Sends metadata to the audio renderer, which renders audio data differently for multiple receivers based on multiple receiver fields and receiver degrees of freedom fields.
[0168] This application embodiment configures the receiver's degree of freedom field in the metadata, which can be used to identify the degree of freedom of different receivers, thereby performing personalized audio rendering based on the degree of freedom of different receivers and obtaining an auditory effect that meets their degree of freedom requirements.
[0169] In one embodiment of this application, the receiver degree of freedom field can be a sub-element of the receiver field. That is, in the receiver field corresponding to each receiver, the receiver degree of freedom field can indicate the degree of freedom of that receiver, satisfying the requirement of rendering audio data corresponding to different degrees of freedom for different receivers. In addition, the receiver degree of freedom field can also be an element in the root metadata, indicating that multiple receivers have the same degree of freedom.
[0170] In one embodiment of this application, a receiver identifier field in the metadata can be assigned a value before the metadata is sent to the audio renderer, thereby distinguishing different receivers of the audio data based on the receiver identifier field.
[0171] Figure 6 A flowchart illustrating the steps of audio rendering based on a receiver identifier in one embodiment of this application is shown. Figure 6 As shown, the method for audio rendering based on metadata carrying a receiver identifier field includes the following steps S610 to S640.
[0172] S610: Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0173] S620: When there are multiple receivers, configure multiple receiver fields in the metadata according to the number information, each corresponding to a different receiver. The receiver fields include feature fields related to the location information of the receivers.
[0174] S630: When there are multiple recipients, assign a value to the recipient identifier field in the metadata corresponding to each recipient. The recipient identifier field is used to distinguish different recipients.
[0175] S640: Sends metadata to the audio renderer, which renders audio data differently for multiple receivers based on multiple receiver fields and a receiver identifier field.
[0176] This application embodiment configures a receiver identifier field in the metadata, which can be used to identify different receivers, thereby enabling personalized audio rendering for different receivers and meeting the diverse auditory effects required by different receivers in their respective scenarios.
[0177] In one embodiment of this application, the receiver identifier field is a sub-element of the receiver field, that is, in the receiver field corresponding to each receiver, the receiver identifier field can indicate the identifier that distinguishes the receiver from other receivers.
[0178] In one embodiment of this application, before sending the metadata to the audio renderer, a coordinate system type field in the metadata can be assigned a value, thereby determining the accurate location of different receivers receiving the audio data based on the coordinate system type field.
[0179] Figure 7 A flowchart illustrating the steps of audio rendering based on coordinate system type in one embodiment of this application is shown. Figure 7 As shown, the method for audio rendering based on metadata carrying a coordinate system type field includes the following steps S710 to S750.
[0180] S710: Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0181] S720: When there are multiple receivers, configure multiple receiver fields in the metadata according to the number information, each corresponding to a different receiver. The receiver fields include feature fields related to the location information of the receivers.
[0182] S730: When there are multiple receivers, the location information of each receiver is obtained. The location information is used to indicate the position coordinates of the receiver in the coordinate system used.
[0183] S740: Assign values to the coordinate system type field in the metadata corresponding to each receiver based on the location information. The coordinate system type field is used to indicate whether the coordinate system is a local coordinate system or a world coordinate system.
[0184] S750: Sends metadata to the audio renderer, which renders audio data differently for multiple receivers based on multiple receiver fields and coordinate system type fields.
[0185] This application embodiment configures a coordinate system type field in the metadata, which can be used to identify the coordinate system type used by different receivers, thereby determining the accurate position of each receiver and performing personalized audio rendering for different receivers based on the accurate position, so as to meet the diverse auditory effects required by different receivers in their respective scenarios.
[0186] In one embodiment of this application, the coordinate system type field can be a sub-element of the receiver field. That is, in the receiver field corresponding to each receiver, the coordinate system type field can indicate the coordinate system type used by that receiver, thus enabling the rendering of audio data corresponding to different coordinate system types for different receivers. Alternatively, the coordinate system type field can also be an element in the root metadata, indicating that multiple receivers use the same coordinate system type.
[0187] In one embodiment of this application, when the coordinate system is the local coordinate system, the origin offset field in the metadata is assigned a value according to the location information. The origin offset field is used to indicate the offset of the origin of the local coordinate system in the world coordinate system.
[0188] In one embodiment of this application, the origin offset field includes offset component fields corresponding to each coordinate axis, which are used to indicate the components of the offset on each coordinate axis.
[0189] In one embodiment of this application, after obtaining the location information of each receiver, the same location receiver identifier field in the metadata can be assigned a value based on whether the locations of multiple receivers are the same. Based on this field, the same location information can be configured for receivers in the same location, reducing the duplication of field data and reducing the consumption of storage resources, computing resources and network resources.
[0190] Figure 8 A flowchart illustrating the steps of audio rendering based on a receiver at the same location in one embodiment of this application is shown. Figure 8 As shown, the method for audio rendering based on metadata carrying a receiver identifier field at the same location includes the following steps S810 to S860.
[0191] S810: Obtain the number of receivers, which indicates the number of receivers receiving audio data;
[0192] S820: When there are multiple receivers, configure multiple receiver fields in the metadata according to the number information, each corresponding to a different receiver. The receiver fields include feature fields related to the location information of the receivers.
[0193] S830: When there are multiple receivers, the location information of each receiver is obtained. The location information is used to indicate the position coordinates of the receiver in the coordinate system used.
[0194] S840: Determine whether each receiver is at the same location based on the location information;
[0195] S850: If the first receiver and the second receiver are at the same position, then the position field corresponding to the first receiver is assigned a value according to the position information, and the same position receiver identifier field corresponding to the second receiver is assigned a value. The position field is used to indicate the position coordinates of the first receiver in the coordinate system, and the same position receiver identifier field is used to indicate the first receiver at the same position as the second receiver.
[0196] S860: Sends metadata to the audio renderer, which renders audio data differently for multiple receivers based on multiple receiver fields, a location field, and a same-location receiver identifier field.
[0197] In one embodiment of this application, the fields related to the receiver's location information can be sub-elements of the receiver field `listener` in the metadata. As shown in Table 4, the sub-elements of the receiver field `listener` can include a receiver degrees of freedom field `listener_DoF`. Based on different values of this field, the corresponding receiver's degrees of freedom can be determined. For example, when this field is 0, it indicates that the receiver has only three degrees of freedom; when this field is 1, it indicates that the receiver has six degrees of freedom.
[0198] The listener field can also include a listener identifier field, listener_id, as a child element. This field can be used to distinguish different listeners when the renderer needs to process user data from multiple listeners.
[0199] The listener field's sub-elements can also include a field called equal_listener_id, which identifies receivers at the same location. This field indicates the identifier of another receiver that is equivalent to receiver A, when receiver A is considered to have the same location information as receiver B.
[0200] The listener field's sub-elements can also include a local_coordinates field, which determines the coordinate system used by the receiver. For example, a value of 0 indicates that multiple receivers use their own local coordinate systems, while a value of 1 indicates that multiple receivers use a common world coordinate system.
[0201] The listener field can also include an origin offset field, which can in turn include offset component fields corresponding to each coordinate axis.
[0202] For example, the offset component field origin_offset_x corresponding to the x-axis represents the x-axis component of the local coordinate system origin offset; when multiple receivers use their own local coordinate systems, this field indicates the offset value of the local coordinate system origin of the respective receiver on the x-axis of the world coordinate system.
[0203] The field origin_offset_y, corresponding to the y-axis offset component, represents the y-axis component of the offset from the origin of the local coordinate system. When multiple receivers use their own local coordinate systems, this field indicates the offset value of the origin of the corresponding receiver's local coordinate system on the y-axis of the world coordinate system.
[0204] The field origin_offset_z, corresponding to the z-axis offset component, represents the z-axis component of the local coordinate system origin offset. When multiple receivers use their own local coordinate systems, this field indicates the offset value of the local coordinate system origin of the respective receiver on the z-axis of the world coordinate system.
[0205] Table 4. Specifications for child elements in the listener field
[0206]
[0207]
[0208] Figure 9 This document illustrates a flowchart of the method steps for audio rendering based on the number of receivers, performed by the audio receiver in one embodiment of this application. Figure 9 As shown, the audio rendering method involving a multi-user scenario executed at the audio receiver includes the following steps S910 to S920.
[0209] S910: Extract the receiver field corresponding to the receiver of the received audio data from the metadata. The receiver field includes feature fields related to the receiver's location information.
[0210] S920: If there are multiple receivers, the audio data is rendered differently for each receiver based on the multiple receiver fields.
[0211] Details of the audio rendering method executed at the audio receiving end and Figures 4 to 8 The audio rendering method executed on the audio acquisition end is the same, so it will not be described again here.
[0212] In one embodiment of this application, the metadata used to indicate the rendering of audio data also includes fields other than those involved in the above embodiments. The specifications of each field are illustrated in Tables 5 to 20 below. Here, "instance" corresponds to "audioChannelFormat," the purpose of which is to allow modification of the metadata content without modifying the ADM.
[0213] Table 5. Specifications for attributes in the instance field
[0214] property describe specification id Unique identifier for this instance INS_0001_0001 type Corresponding audio type code {0001,0003,0004} typeLabel Corresponding audio type name {DirectSpeaker,Objects,HOA}
[0215] Table 6. Specifications for child elements in the instance field.
[0216] child elements describe specification audioChannelFormatRefID ID matching with ADM element AC_00010001 unitInfo There can be multiple corresponding audioBlockFormats.
[0217] Table 7. Specifications for attributes in the unitInfo field
[0218] property describe specification id Unique identifier for this unit UNI_00010001 start The start time of this unit, valid for Objects type. 00:00:00.00000 duration The duration of this unit, valid for Objects type. 00:00:00.00000
[0219] Table 8 specifies the relevant specifications for child elements in the unitInfo field (typeLabel == DirectSpeakers).
[0220] child elements describe specification speakerLabel Speaker layout Refer to speaker layout specifications gain Linear gain 0-16
[0221] Table 9 specifies the relevant specifications for child elements in the unitInfo field (typeLabel == Objects).
[0222] child elements describe specification azimuth Horizontal angle, in degrees -180-180 elevation Elevation angle, in degrees -90~90 distance Distance, in meters 0-50 gain Linear gain 0-16
[0223] Table 10 specifies the relevant specifications for child elements in the unitInfo field (typeLabel == HOA).
[0224] child elements describe specification order order of the corresponding channel 0-7 degree corresponding channel angle -7~+7 normalization Normalization method {0,1…} gain Linear gain 0-16
[0225] Table 11 Specifications for child elements in the staticControl field
[0226] child elements describe specification ambisonicOrder spherical harmonic coding order 1-7 acousticEnv Acoustic environment related rendererInfo Post-processing related to rendering
[0227] Table 12 Specifications for sub-elements in the acousticEnv field
[0228] child elements describe specification type Environmental acoustic types {0,1,2} typeLabel Environmental acoustics type tags {Physical / Artificial / Sample} earlyReflectionGain Early reflection gain [0.0-1.0] lateReverbGain Post-reverberation gain [0.0-1.0] lowFreqProFlag Low-frequency separation processing 0 / 1, low frequencies can be reverberated or not. convolutionReverbType Sampled reverberation type {0,1,2…} surface Reflecting surface in geometric space Supports spatial models composed of an unlimited number of reflective surfaces
[0229] Table 13 Specifications for child elements in the surface field
[0230]
[0231] Table 14 Specifications for child elements in the audioEffect field
[0232] child elements describe specification EQ EQ Post-processing DRC DRC post-processing Gain Gain post-processing
[0233] Table 15 Specifications for Attributes in the EQ Field
[0234] property describe specification index Indicates the order of sound effect links {0,1,2…}
[0235] Table 16 Attribute Specifications for Multi-segment Item Sub-elements in the EQ Field
[0236] item attributes describe specification type Filter type {0,1,2…} typeLabel Filter type label {lowpass, highpass, bandpass...} frequency Cutoff frequency 20-16000Hz gain Gain [-40,-40]dB Q quality factor 0.1-12
[0237] Table 17 EQ Types
[0238] type typeLabel 0 LowPass 1 HighPass 2 BandPass 3 BandReject 4 AllPass 5 LowShelving 6 HighShelving
[0239]
[0240] Table 18 Specifications for Attributes in DRC Fields
[0241] property describe specification index Indicates the order of sound effect links {0,1,2…} attackTime Start time [0-100]ms releaseTime Release time [50-300]ms threshold Threshold [-80,10]dB preGain Preamplifier Gain [-10,10]dB postGain Post-gain [0,20]dB ratio Compression ratio 1-100
[0242] Table 19 Specifications for attributes in the Gain field
[0243] property describe specification index Indicates the order of sound effect links {0,1,2…} gain Gain -20~20dB
[0244] Table 20 Specifications for Sub-elements in the dynamicControl Field
[0245]
[0246] The following combination Figures 10 to 11 This application provides a detailed description of the audio rendering methods for various audio playback devices described in its embodiments. Figure 10 This application describes an audio rendering method implemented at an audio acquisition terminal, which involves various audio playback devices. The audio acquisition terminal may be, for example, a... Figure 1 The terminal device shown Figure 2 The acquisition subsystem shown or Figure 3 The data acquisition terminal shown. Figure 11 This application describes an audio rendering method implemented at an audio receiving end, which involves various audio playback devices. The audio receiving end may be, for example, a... Figure 1 The terminal device shown Figure 2 The client subsystem shown or Figure 3 The registered renderer is shown.
[0247] Figure 10 This document illustrates a flowchart of the method steps for audio rendering based on device information performed by the audio acquisition terminal in one embodiment of this application. Figure 10 As shown, the audio rendering method involving diverse audio playback devices executed at the audio acquisition end includes the following steps S1010 to S1040.
[0248] S1010: Obtain the receiver's device information, which indicates the audio playback device used by the receiver when receiving audio data;
[0249] S1020: Assign a value to the device information field in the metadata based on the device information. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect.
[0250] S1030: Assign a value to the corresponding device identifier field in the metadata based on the device information field. The corresponding device identifier field is a sub-element of the rendering information field, which is an element in the metadata used to indicate the rendering effect of the audio data.
[0251] S1040: Send metadata to the audio renderer, which is used to render audio data for the audio playback device based on the device information field.
[0252] This application embodiment configures corresponding device identifier fields and device information fields in the metadata. The relevant fields can be used to identify the audio playback device of the receiver receiving the audio data, so that the audio data can be rendered to form a personalized audio rendering effect adapted to different audio playback devices under the indication of the relevant fields, thus meeting the needs of diverse audio rendering scenarios.
[0253] In one embodiment of this application, the corresponding device identifier field is a sub-element of the rendering information field, and the corresponding device identifier field can directly reference the sub-element of the device information field for assignment.
[0254] As shown in Table 21, the metadata implemented using the embodiments of this application includes a rendering information field `rendererInfo`, and within the sub-elements of the `rendererInfo` field is a corresponding device identifier field `refDeviceID`. The `refDeviceID` field is a sub-element used to describe the device identifier or device group identifier corresponding to the current rendering effect, and its value is either the device identifier or the device group identifier. As shown in Table 22, the metadata implemented using the embodiments of this application includes a device information field `deviceInfo`, which provides several sub-elements related to the audio playback device.
[0255] Table 21 specifies the relevant specifications for child elements in the rendererInfo field of the rendering information.
[0256]
[0257] Table 22 Specifications for Sub-elements in the DeviceInfo Field of Device Information
[0258]
[0259]
[0260] In one embodiment of this application, the device information field includes a device identifier field or a device group identifier field. The device identifier field is used to indicate an identifier that distinguishes different audio playback devices, and the device group identifier field is used to indicate the device group to which the audio playback device belongs. Audio playback devices belonging to the same device group have the same audio data rendering effect. For example, the device identifier field deviceID shown in Table 22 is used to indicate the identifier of the current device, and the device group identifier field deviceGroupID is used to indicate the identifier of the device group to which the current device belongs. Devices belonging to the same device group have the same rendering effect. The device identifier and the device group identifier are not duplicated.
[0261] In one embodiment of this application, the device information field includes a device name field or a device type field. The device name field indicates a readable name of the audio playback device, and the device type field indicates the device type of the audio playback device. For example, the device name field deviceName shown in Table 22 represents a readable device name. The device type field deviceType indicates the type of the current device; a value of 0 indicates headphones, and a value of 1 indicates a speaker.
[0262] In one embodiment of this application, the device information field includes a device location field, which indicates the distribution location of the audio playback device relative to the receiver. For example, the device location field devicePos is shown in Table 22.
[0263] In one embodiment of this application, when the audio playback device is a headphone, the distribution position includes the left ear or the right ear; when the audio playback device is a speaker, the distribution position includes the center, front left, front right, rear left, or rear right. For example, when the audio playback device is a headphone, a devicePos value of 0 indicates the left earphone; a value of 1 indicates the right earphone. When the audio playback device is a speaker, a devicePos value of 0 indicates the center speaker; a value of 1 indicates the front left speaker; a value of 2 indicates the front right speaker; a value of 3 indicates the rear left speaker; and a value of 4 indicates the rear right speaker.
[0264] In one embodiment of this application, the device information field includes a corresponding receiver identifier field, which indicates the identifier of the receiver corresponding to the audio playback device. For example, the RefListenerID shown in Table 22 is used to indicate the identifier of the receiver to which the current device belongs.
[0265] Figure 11 This document illustrates a flowchart of the method steps for audio rendering based on device information performed by an audio receiver in one embodiment of this application. Figure 11 As shown, the audio rendering method involving diverse audio playback devices executed at the audio receiving end includes the following steps S1110 to S1130.
[0266] S1110: Extract the corresponding device identifier field from the rendering information field of the metadata. The rendering information field is an element in the metadata used to indicate the rendering effect of the audio data.
[0267] S1120: Extract the device information field based on the corresponding device identifier field. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect.
[0268] S1130: Render audio data for the audio playback device based on the device information field.
[0269] Details of the audio rendering method executed at the audio receiving end and Figure 10 The audio rendering method executed on the audio acquisition end is the same, so it will not be described again here.
[0270] The following combination Figures 12 to 13 This application provides a detailed description of audio rendering methods involving diverse rendering types in its embodiments. Among them, Figure 12This application's embodiment describes an audio rendering method involving diverse rendering types executed at an audio acquisition terminal. The audio acquisition terminal could be, for example, […]. Figure 1 The terminal device shown Figure 2 The acquisition subsystem shown or Figure 3 The data acquisition terminal shown. Figure 13 This application's embodiment describes an audio rendering method involving diverse rendering types executed at an audio receiving end. The audio receiving end may be, for example, […]. Figure 1 The terminal device shown Figure 2 The client subsystem shown or Figure 3 The registered renderer is shown.
[0271] Figure 12 This document illustrates a flowchart of the method steps for audio rendering based on rendering type, executed by the audio acquisition terminal in one embodiment of this application. Figure 12 As shown, the audio rendering method involving diverse rendering types executed at the audio acquisition end includes the following steps S1210 to S1230.
[0272] S1210: Obtain the rendering type information of the audio data. The rendering type information is used to indicate the rendering type of each audio component in the audio data.
[0273] S1220: Assign a value to the audio component information field in the metadata according to the rendering type information. The audio component information field is used to indicate the audio component with the corresponding rendering type that is processed by the audio renderer.
[0274] S1230: Send metadata to the audio renderer, which is used to render audio components according to the rendering type.
[0275] This application embodiment configures an audio component information field in the metadata, which can identify the rendering type of the audio component. This allows the audio data to be rendered under the guidance of the relevant fields to form a personalized audio rendering effect that adapts to different rendering types, thus meeting the needs of diverse audio rendering scenarios.
[0276] In one embodiment of this application, the audio component information field can be an element in the root metadata. As shown in Table 23, in the VR extended root metadata implemented using the embodiments of this application, there exists an audio component information field called componentsInfo, which can be used to indicate the audio component processed by the audio renderer.
[0277] Table 23 includes the relevant specifications for VR extended root metadata, including the audio component information field.
[0278] <vrext> < / vrext> describe specification version Version number of extended metadata 0.0.1 name Naming extended metadata vrRExt level Extended metadata priority 1 presenceInfo Extended metadata content related sections componentsInfo Audio components processed by the renderer staticControl Static content of extended metadata dynamicControl Dynamic content of extended metadata
[0279] In one embodiment of this application, the audio component information field includes a component type field, which indicates whether the rendering type corresponding to the audio component is immersive audio rendering or non-immersive audio rendering. As shown in Table 24, the sub-element of the audio component information field componentsInfo includes a component type field componentType. The type of audio component participating in the rendering can be determined according to the different values of this field. For example, a value of 0 indicates that the audio component participates in immersive audio rendering; a value of 1 indicates that the audio component participates in non-immersive audio rendering, i.e., normal audio rendering.
[0280] Table 24 Specifications for Sub-elements in the Audio Component Information Field (componentsInfo)
[0281]
[0282] In one embodiment of this application, the audio component information field includes a component identifier field or a component group identifier field. The component identifier field is used to indicate an identifier that distinguishes different audio components; the component group identifier field is used to indicate the component group to which the audio component belongs. Audio components belonging to the same component group are used as input parameters of the audio renderer. As shown in Table 24, the sub-elements of the audio component information field componentsInfo contain a component identifier field componentID and a component group identifier field componentGroupID. Audio components within the same group are used as input to the renderer algorithm.
[0283] In one embodiment of this application, the audio component information field includes a corresponding audio format identifier field, which is used to indicate the identifier of the audio format corresponding to the audio component in the audio definition model. As shown in Table 24, the audio component information field componentsInfo contains a corresponding audio format identifier field refAudioFormatID, which is used to describe the identifier in the ADM corresponding to the audio component.
[0284] Figure 13 This document illustrates a flowchart of the method steps for audio rendering based on rendering type performed by an audio receiver in one embodiment of this application. Figure 13 As shown, the audio rendering method involving diverse rendering types executed at the audio receiving end includes the following steps S1310 to S1320.
[0285] S1310: Extract the audio component information field from the metadata. The audio component information field is used to indicate the audio component with the corresponding rendering type that is processed by the audio renderer.
[0286] S1320: Determine the rendering type corresponding to each audio component in the audio data based on the audio component information field.
[0287] S1330: Render the audio component according to the rendering type.
[0288] Details of the audio rendering method executed at the audio receiving end and Figure 12 The audio rendering method executed on the audio acquisition end is the same, so it will not be described again here.
[0289] The following combination Figures 14 to 16 This application provides a detailed description of audio rendering methods involving diverse audio occlusion effects in its embodiments. Among them, Figures 14 to 15 This application's embodiment describes an audio rendering method involving diverse audio occlusion effects executed at an audio acquisition terminal. The audio acquisition terminal could be, for example, […]. Figure 1 The terminal device shown Figure 2 The acquisition subsystem shown or Figure 3 The data acquisition terminal shown. Figure 16 This application's embodiment describes an audio rendering method involving diverse audio occlusion effects executed at an audio receiving end. The audio receiving end may be, for example, [the following is a list of possible methods]. Figure 1 The terminal device shown Figure 2 The client subsystem shown or Figure 3 The registered renderer is shown.
[0290] Figure 14 This document illustrates a flowchart of the method steps for audio rendering based on occlusion object information performed by the audio acquisition terminal in one embodiment of this application. Figure 14 As shown, the audio rendering method involving diverse audio occlusion effects executed at the audio acquisition end includes the following steps S1410 to S1430.
[0291] S1410: Obtain the attribute information of occluded objects in the rendering scene of the audio data;
[0292] S1420: Assign a value to the occlusion object information field in the metadata based on the attribute information of the occluding object. The occlusion object information field is used to indicate the characteristics of the occluding object that produces the audio occlusion effect.
[0293] S1430: Send metadata to the audio renderer, which renders audio data with audio occlusion effect based on the occlusion object information field.
[0294] This application embodiment configures an occlusion object information field in the metadata, which can be used to identify occlusion objects in the audio rendering scene. This allows audio data to be rendered under the guidance of the relevant fields to form personalized audio rendering effects with different audio occlusion effects, thus meeting the diverse needs of audio rendering scenarios.
[0295] As shown in Table 25, the occlusionObjectInfo field can include multiple sub-elements that describe the characteristics of the occluded object.
[0296] Table 25 Specifications for Child Elements in the OcclusionObjectInfo Field Occlusion Object Information Table 25: Specifications for Child Elements in the OcclusionObjectInfo Field
[0297] child elements describe specification objectID Identifiers of occluded objects within the scene pos_x x-coordinate of the occluded object any valid position in the three-dimensional coordinate system pos_y y-coordinate of the occluded object any valid position in the three-dimensional coordinate system pos_z z-coordinate of the occluded object any valid position in the three-dimensional coordinate system boundingBox_x x-coordinate of the occluded object's size any valid position in the three-dimensional coordinate system boundingBox_y y-coordinate of the occluded object's size any valid position in the three-dimensional coordinate system boundingBox_z z-coordinate of the occluded object any valid position in the three-dimensional coordinate system
[0298] In one embodiment of this application, the occlusion object information field includes an occlusion position field, which indicates the position coordinates of the occluding object in the coordinate system. As shown in Table 25, the occlusion object information field occlusionObjectInfo contains occlusion position fields pos_x, pos_y, and pos_z. Specifically, pos_x represents the x-coordinate of the occluding object's position, pos_y represents the y-coordinate, and pos_z represents the z-coordinate.
[0299] In one embodiment of this application, the occlusion object information field includes an occlusion size field, which indicates the size coordinates of the occluded object in the coordinate system. As shown in Table 25, the occlusion object information field `occlusionObjectInfo` contains occlusion size fields `boundingBox_x`, `boundingBox_y`, and `boundingBox_z`. Specifically, `boundingBox_x` represents the x-coordinate of the occluded object's size, `boundingBox_y` represents the y-coordinate, and `boundingBox_z` represents the z-coordinate.
[0300] In another embodiment of this application, the occlusion object information field includes a vertex position field, which indicates the maximum and minimum position coordinates of vertices on the occlusion object. That is, the position and size fields of the occlusion object can be determined by the minimum and maximum values of the coordinates of each vertex of the object in the coordinate system. As shown in Table 26, the sub-elements of the occlusion object information field occlusionObjectInfo contain vertex position fields pos_x_min, pos_y_min, pos_z_min, pos_x_max, pos_y_max, and pos_z_max. Here, pos_x_min represents the minimum x-coordinate of a vertex position on the occlusion object, pos_y_min represents the minimum y-coordinate of a vertex position on the occlusion object, pos_z_min represents the minimum z-coordinate of a vertex position on the occlusion object, pos_x_max represents the maximum x-coordinate of a vertex position on the occlusion object, pos_y_max represents the maximum y-coordinate of a vertex position on the occlusion object, and pos_z_max represents the maximum z-coordinate of a vertex position on the occlusion object.
[0301] Table 26 Specifications for child elements in the occlusionObjectInfo field of Table 26
[0302] child elements describe specification objectID Identifiers of occluded objects within the scene pos_x_min Minimum x-coordinate of the vertex position of the occluded object any valid position in the three-dimensional coordinate system pos_y_min Minimum y-coordinate of the vertex position of the occluded object any valid position in the three-dimensional coordinate system pos_z_min Minimum z-coordinate of the vertex position of the occluded object any valid position in the three-dimensional coordinate system pos_x_max The maximum x-coordinate of the vertex position of the occluded object. any valid position in the three-dimensional coordinate system pos_y_max The maximum y-coordinate of the vertex position of the occluded object. any valid position in the three-dimensional coordinate system pos_z_max The maximum z-coordinate of the vertex position of the occluded object any valid position in the three-dimensional coordinate system
[0303] In one embodiment of this application, the occlusion object information field includes an object identifier field, which is used to indicate the identifier that distinguishes different occlusion objects. As shown in Table 25 or Table 26, the object identifier field objectID exists in the sub-elements of the occlusion object information field occlusionObjectInfo. The object identifier field objectID represents the identifier of the occlusion object located in the rendering scene.
[0304] Figure 15 This document illustrates a flowchart of the method steps for audio rendering based on occlusion parameters performed by the audio acquisition terminal in one embodiment of this application. Figure 15 As shown, the method for audio rendering based on diverse occlusion parameters executed at the audio acquisition end may include the following steps S1510 to S1540.
[0305] S1510: Obtain the attribute information of occluded objects in the rendering scene of the audio data;
[0306] S1520: Assign a value to the occlusion object information field in the metadata based on the attribute information of the occluding object. The occlusion object information field is used to indicate the characteristics of the occluding object that produces the audio occlusion effect.
[0307] S1530: Assign values to the occlusion parameter field in the metadata based on the attribute information of the occluded object. The occlusion parameter field is used to indicate the parameters that affect the audio occlusion effect.
[0308] S1540: Send metadata to the audio renderer, which renders audio data with audio occlusion effect based on the occlusion object information field and the occlusion parameter field.
[0309] This application embodiment configures occlusion object information fields and occlusion parameter fields in the metadata. By using these fields, occlusion objects and occlusion parameters in the audio rendering scene can be identified. This allows audio data to be rendered under the guidance of these fields to form personalized audio rendering effects with different audio occlusion effects, thus meeting the diverse needs of audio rendering scenarios.
[0310] As shown in Table 27, the occlusion parameter field can include multiple sub-elements that describe the occlusion parameters.
[0311] Table 27 Specifications for Sub-elements in the OcclusionParameter Field of Table 27
[0312]
[0313] In one embodiment of this application, the occlusion parameter field includes an occlusion type field, which indicates the type of algorithm used when rendering the audio occlusion effect.
[0314] In one embodiment of this application, the algorithm type includes achieving audio occlusion by adjusting the volume or by filtering high-frequency signals.
[0315] As shown in Table 27, the occlusion parameter field occlusion parameter has an occlusion type field occlusionType in its sub-elements. The different values of this field can describe the type of algorithm used when rendering the occlusion effect. For example, a value of 0 indicates that the occlusion effect is achieved by adjusting the volume; a value of 1 indicates that the occlusion effect is achieved by filtering high-frequency signals.
[0316] In one embodiment of this application, the occlusion parameter field includes an occlusion intensity field, which indicates the intensity of the rendered audio occlusion effect. As shown in Table 26, the occlusion parameter field occlusionParameter contains an occlusion intensity field occlusionIntensity, and different values of this field can describe the intensity of the occlusion effect.
[0317] In one embodiment of this application, the occlusion parameter field includes a maximum occlusion intensity field, which indicates the maximum intensity of the audio occlusion effect on the audio data. This maximum intensity can represent the intensity required to achieve complete audio occlusion, or it can be the intensity required to achieve a specified degree of audio occlusion (e.g., the intensity corresponding to a 90% occlusion effect). As shown in Table 27, the sub-elements of the occlusion parameter field occlusionParameter contain a maximum occlusion intensity field maxOcclusionIntensity. When the occlusion effect reaches the maximum occlusion intensity indicated by this field, the sound can be considered completely occluded. In terms of distance, if the maximum occlusion intensity is 10 and the occlusion intensity corresponding to a certain occluding object is 5, then the occlusion effect produced after the audio passes through the occluding object can be considered to be a 50% attenuation.
[0318] In one embodiment of this application, the occlusion parameter field includes a corresponding object identifier field, which is used to indicate the identifiers of one or more occluding objects referenced when rendering the audio occlusion effect. As shown in Table 27, the occlusion parameter field occlusionParameter contains a corresponding object identifier field refObjectID, which can describe the identifiers of the occluding objects referenced when calculating the occlusion effect and can directly index one or more occluding objects.
[0319] Figure 16 This document illustrates a flowchart of the method steps for audio rendering based on occlusion object information performed by an audio receiver in one embodiment of this application. Figure 16 As shown, the audio rendering method involving diverse occlusion parameters executed at the audio receiver includes the following steps S1610 to S1620.
[0320] S1610: Extract the occlusion information field from the metadata. The occlusion information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0321] S1620: Render audio data with audio occlusion effect based on the occlusion object information field.
[0322] Details of the audio rendering method executed at the audio receiving end and Figure 14 The audio rendering method executed on the audio acquisition end is the same, so it will not be described again here.
[0323] The following combination Figures 17 to 18 This application provides a detailed description of audio rendering methods involving diverse scene configurations in its embodiments. Among them, Figure 17 This application's embodiment describes an audio rendering method involving diverse scene configurations executed at an audio acquisition terminal. The audio acquisition terminal could be, for example, […]. Figure 1The terminal device shown Figure 2 The acquisition subsystem shown or Figure 3 The data acquisition terminal shown. Figure 18 This application's embodiment describes an audio rendering method involving diverse scene configurations executed at an audio receiving end. The audio receiving end may be, for example, [the following is a list of possible methods]. Figure 1 The terminal device shown Figure 2 The client subsystem shown or Figure 3 The registered renderer is shown.
[0324] Figure 17 This document illustrates a flowchart of the method steps for audio rendering based on scene configuration information executed by the audio acquisition terminal in one embodiment of this application. Figure 17 As shown, the audio rendering method involving diverse scene configurations executed at the audio acquisition end includes the following steps S1710 to S1730.
[0325] S1710: Obtain the scene configuration information of the receiver. The scene configuration information is the configuration information that affects the rendering effect of audio data in the rendering scene.
[0326] S1720: Assign values to the feature fields of the metadata according to the scene configuration information. The feature fields include at least one of the following: receiver field, device information field, audio component information field, or occlusion object information field. The receiver field includes fields related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0327] S1730: Send metadata to the audio renderer, which is used to render audio data based on feature fields.
[0328] This application embodiment configures feature fields related to scene configuration information in the metadata. By using these fields, various scene configurations that affect the audio rendering effect in the audio rendering scene can be identified, so that the audio data can obtain different audio rendering effects under the indication of the relevant fields, thus meeting the diverse needs of audio rendering scenes.
[0329] In some optional implementations, any of the aforementioned feature fields can be configured in the metadata based on the receiver's scene configuration information to achieve audio rendering effects adapted to a single scene configuration type. In other optional implementations, combinations of two or more of the aforementioned feature fields can also be configured in the metadata to achieve audio rendering effects adapted to multiple scene configuration types.
[0330] For example, the receiver field and device information field can be configured in the metadata at the same time to meet the audio rendering scenario requirements that are adapted to the receiver's location information and the device characteristics of the audio playback device.
[0331] For example, device information fields and occlusion object information fields can be configured simultaneously in the metadata to meet the audio rendering scenario requirements that are adapted to the device characteristics of the audio playback device and the characteristics of the occluding object.
[0332] Figure 18 This document illustrates a flowchart of the method steps for audio rendering based on scene configuration information performed by an audio receiver in one embodiment of this application. Figure 18 As shown, the audio rendering method involving diverse scene configurations executed at the audio receiver includes the following steps S1810 to S1820.
[0333] S1810: Extract feature fields from metadata. Feature fields include at least one of receiver fields, device information fields, audio component information fields, or occlusion object information fields. Receiver fields include fields related to the location information of the receiver. Device information fields are used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. Audio component information fields are used to indicate the audio component with the corresponding rendering type processed by the audio renderer. Occlusion object information fields are used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0334] S1820: Render audio data based on feature fields.
[0335] Details of the audio rendering method executed at the audio receiving end and Figure 17 The audio rendering method executed on the audio acquisition end is the same, so it will not be described again here.
[0336] In the audio rendering method involving diverse scene configurations in the embodiments of this application, the implementation details related to the receiver field can be found in the corresponding above. Figures 4 to 9 The embodiments described herein will not be repeated here.
[0337] For details regarding the implementation methods related to device information fields in the audio rendering methods with diverse scene configurations described in this application, please refer to the corresponding methods described above. Figures 10 to 11 The embodiments described herein will not be repeated here.
[0338] For details regarding the implementation methods related to the audio component information field in the audio rendering methods with diverse scene configurations described in this application, please refer to the corresponding methods described above. Figures 12 to 13 The embodiments described herein will not be repeated here.
[0339] In the audio rendering method involving diverse scene configurations in this application, the implementation details related to the occlusion object information field can be found in the corresponding description above. Figures 14 to 16 The embodiments described herein will not be repeated here.
[0340] The following examples of two specific application scenarios illustrate the audio rendering method provided in this application.
[0341] In one application scenario of this application, a method for audio rendering based on user location, device, content and other related information may include the following steps.
[0342] 1. The server generates corresponding audio metadata files based on user-provided location, device information, etc. Among them:
[0343]
[0344]
[0345] 2. The server sends the audio metadata file to the audio renderer.
[0346] a) If the renderer is a cloud renderer, then the renderer is located on the server side.
[0347] b) If the renderer is a user-local renderer, then the renderer is located on the client side.
[0348] 3. The audio renderer extracts location and device information from the metadata file as input parameters for rendering, and calculates the rendering effect for the corresponding user and device.
[0349] For receiver 1, it only has three degrees of freedom, so it only has rotation-related information and not position information.
[0350] Both receivers 2 and 3 have six degrees of freedom, thus possessing both rotation-related information and position-related information.
[0351] When the renderer renders based on the coordinates of receivers 1, 2, and 3, no coordinate transformation is needed because the coordinates are in the global coordinate system. Furthermore, since receivers 2 and 3 are in the same position, the calculation of the 3D spatial rendering effect only needs to be based on the coordinates of receiver 2.
[0352] Regarding the rendering information, assuming the calculated rendering effect corresponds to device ID 100, the device information list shows that it corresponds to a device group. This indicates that all devices within device group 100 are eligible for this rendering effect, namely device1 and device2.
[0353] For each component in the audio file, component 1 corresponds to audio channel 1 in the file, and component 2 corresponds to audio channel 2. Component 1 participates in immersive rendering, while component 2 does not. Therefore, when rendering audio, the renderer needs to obtain the data for audio channel 1 and audio channel 2 based on this metadata information and process them separately.
[0354] 4. Present the audio on the corresponding user's corresponding device.
[0355] a) If the renderer is a cloud renderer, the server sends the rendered audio to the client, and the client then renders it.
[0356] b) If the renderer is the user's local renderer, the client renders it directly.
[0357] In another application scenario of this application embodiment, the method for audio rendering based on occlusion effect may include the following steps.
[0358] 1. The server generates corresponding audio metadata files based on user-provided location, device information, etc. Among them:
[0359]
[0360]
[0361] 2. The server sends the audio metadata file to the audio renderer.
[0362] a) If the renderer is a cloud renderer, then the renderer is located on the server side.
[0363] b) If the renderer is a user-local renderer, then the renderer is located on the client side.
[0364] 3. The audio renderer extracts receiver location information, occlusion information, and occlusion algorithm information from the metadata file as input parameters related to occlusion effect rendering.
[0365] By using the location of listener1, the location information of the obstruction, and the sound source location information already supported in the current standard, the degree to which the sound source is obstructed when it propagates to the receiver in the scene can be calculated.
[0366] Then, based on the occlusion intensity (occlusionIntensity) / maxOcclusionIntensity, the proportion of sound blocked is finally calculated. For example, if the occlusion intensity (occlusionIntensity) is 2 and the maximum occlusion intensity (maxOcclusionIntensity) is 10, then the proportion of sound blocked is 20%.
[0367] Finally, based on the occlusion effect algorithm type (occlusionType=0), the occlusion effect is ultimately achieved by adjusting the volume of the sound source.
[0368] 4. Present the audio on the corresponding user's corresponding device.
[0369] a) If the renderer is a cloud renderer, the server sends the rendered audio to the client, and the client then renders it.
[0370] b) If the renderer is the user's local renderer, the client renders it directly.
[0371] As can be seen from the above application scenario examples, the audio rendering method provided by this application embodiment can perform differentiated audio rendering based on the location information feedback of multiple users. In addition, it can distinguish between 3DoF and 6DoF audio rendering scenarios, more accurately define user devices, distinguish audio components in audio content that participate in immersive rendering and normal rendering, and provide diverse audio occlusion rendering effects, thereby supporting richer immersive audio application scenarios.
[0372] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0373] The following describes an apparatus embodiment of this application, which can be used to execute the audio rendering method in the above embodiments of this application.
[0374] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on the number of receivers at an audio acquisition end. The apparatus includes:
[0375] The first acquisition module is configured to acquire the number of receivers, the number of receivers being used to indicate the number of receivers receiving audio data.
[0376] The first assignment module is configured to, when there are multiple receivers, configure multiple receiver fields in the metadata according to the quantity information, each corresponding to one of the multiple receivers. The receiver fields include feature fields related to the location information of the receivers.
[0377] The first sending module is configured to send the metadata to an audio renderer, which is used to render the audio data differently for multiple recipients based on multiple recipient fields.
[0378] In one embodiment of this application, the first assignment module is further configured to assign a value to the multi-receiver flag field in the metadata according to the quantity information, wherein the multi-receiver flag field is used to indicate that the number of the receivers is one or more.
[0379] In one embodiment of this application, the metadata includes a presentation information field, which is an element related to the content of the audio data, and the multi-receiver flag field is an attribute or sub-element of the presentation information field.
[0380] In one embodiment of this application, the first acquisition module is further configured to acquire the degree-of-freedom information of the receiver, the degree-of-freedom information being used to indicate the degree of freedom of the receiver when receiving the audio; the first assignment module is further configured to assign a value to the receiver degree-of-freedom field in the metadata according to the degree-of-freedom information, the receiver degree-of-freedom field being used to indicate that the receiver has three degrees of freedom or six degrees of freedom.
[0381] In one embodiment of this application, the first assignment module is further configured to assign a value to the receiver identifier field in the metadata corresponding to each receiver when there are multiple receivers, wherein the receiver identifier field is used to distinguish different receivers.
[0382] In one embodiment of this application, the first acquisition module is further configured to acquire the location information of each of the receivers when there are multiple receivers, wherein the location information is used to indicate the location coordinates of the receiver in the coordinate system it is using; the first assignment module is further configured to assign a value to the coordinate system type field in the metadata corresponding to each receiver according to the location information, wherein the coordinate system type field is used to indicate whether the type of the coordinate system is a local coordinate system or a world coordinate system.
[0383] In one embodiment of this application, the first assignment module is further configured to assign a value to the origin offset field in the metadata according to the location information when the coordinate system is a local coordinate system. The origin offset field is used to indicate the offset of the origin of the local coordinate system in the world coordinate system.
[0384] In one embodiment of this application, the origin offset field includes an offset component field corresponding to each coordinate axis, the offset component field being used to indicate the component of the offset on each of the coordinate axes.
[0385] In one embodiment of this application, the first assignment module is further configured to determine whether each of the receivers is in the same position based on the location information; if the first receiver and the second receiver are in the same position, then a value is assigned to the position field corresponding to the first receiver based on the location information, and a value is assigned to the same-position receiver identifier field corresponding to the second receiver. The position field is used to indicate the position coordinates of the first receiver in the coordinate system, and the same-position receiver identifier field is used to indicate the first receiver that is in the same position as the second receiver.
[0386] In one embodiment of this application, an audio rendering apparatus located at an audio receiver that performs audio rendering based on the number of receivers is provided. The apparatus includes:
[0387] The first extraction module is configured to extract a receiver field from metadata corresponding to the receiver receiving the audio data. The receiver field includes feature fields related to the location information of the receiver.
[0388] The first rendering module is configured to render the audio data differently for each of the multiple receivers based on the multiple receiver fields if there are multiple receivers.
[0389] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on device information at an audio acquisition end. The apparatus includes:
[0390] The second acquisition module is configured to acquire the device information of the receiver, wherein the device information is used to indicate the audio playback device used by the receiver when receiving audio data;
[0391] The second assignment module is configured to assign a value to the device information field in the metadata based on the device information, wherein the device information field is used to indicate the characteristics of the audio playback device that affect the audio data rendering effect; and to assign a value to the corresponding device identifier field in the metadata based on the device information field, wherein the corresponding device identifier field is a sub-element of the rendering information field, and the rendering information field is an element in the metadata used to indicate the audio data rendering effect;
[0392] The second sending module is configured to send the metadata to the audio renderer, which is used to render the audio data for the audio playback device based on the device information field.
[0393] In one embodiment of this application, the device information field includes a device identifier field or a device group identifier field. The device identifier field is used to indicate an identifier that distinguishes different audio playback devices, and the device group identifier field is used to indicate the device group to which the audio playback device belongs. Audio playback devices belonging to the same device group have the same audio data rendering effect.
[0394] In one embodiment of this application, the device information field includes a device name field or a device type field, wherein the device name field is used to indicate the readable name of the audio playback device, and the device type field is used to indicate the device type of the audio playback device.
[0395] In one embodiment of this application, the device information field includes a device location field, which is used to indicate the distribution location of the audio playback device relative to the receiver.
[0396] In one embodiment of this application, when the audio playback device is an earphone, the distribution position includes the left ear or the right ear; when the audio playback device is a speaker, the distribution position includes the middle, front left, front right, rear left, or rear right.
[0397] In one embodiment of this application, the device information field includes a corresponding receiver identifier field, which is used to indicate the identifier of the receiver corresponding to the audio playback device.
[0398] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on device information at an audio receiving end. The apparatus includes:
[0399] The second extraction module is configured to extract a corresponding device identifier field from the rendering information field of the metadata, wherein the rendering information field is an element in the metadata used to indicate the audio data rendering effect; and to extract a device information field based on the corresponding device identifier field, wherein the device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect.
[0400] The second rendering module is configured to render the audio data for the audio playback device based on the device information field.
[0401] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on rendering type at an audio acquisition end. The apparatus includes:
[0402] The third acquisition module is configured to acquire rendering type information of audio data, wherein the rendering type information is used to indicate the rendering type corresponding to each audio component in the audio data;
[0403] The third assignment module is configured to assign a value to the audio component information field in the metadata according to the rendering type information. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer.
[0404] The third sending module is configured to send the metadata to the audio renderer, which is used to render the audio component according to the rendering type.
[0405] In one embodiment of this application, the audio component information field includes a component type field, which is used to indicate whether the rendering type corresponding to the audio component is immersive audio rendering or non-immersive audio rendering.
[0406] In one embodiment of this application, the audio component information field includes a component identifier field or a component group identifier field. The component identifier field is used to indicate an identifier that distinguishes different audio components; the component group identifier field is used to indicate the component group to which the audio component belongs, and audio components belonging to the same component group are used together as input parameters of the audio renderer.
[0407] In one embodiment of this application, the audio component information field includes a corresponding audio format identifier field, which is used to indicate the identifier of the audio format corresponding to the audio component in the audio definition model.
[0408] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on rendering type at an audio receiving end. The apparatus includes:
[0409] The third extraction module is configured to extract an audio component information field from metadata, the audio component information field being used to indicate an audio component with a corresponding rendering type processed by the audio renderer;
[0410] The third rendering module is configured to determine the rendering type corresponding to each audio component in the audio data based on the audio component information field; and to render the audio component according to the rendering type.
[0411] In one embodiment of this application, an audio rendering apparatus is provided that performs audio rendering based on occlusion objects at an audio acquisition end. The apparatus includes:
[0412] The fourth acquisition module is configured to acquire attribute information of occluded objects in the rendering scene of the audio data;
[0413] The fourth assignment module is configured to assign a value to the occlusion object information field in the metadata based on the attribute information of the occluding object. The occlusion object information field is used to indicate the characteristics of the occluding object that produces the audio occlusion effect.
[0414] The fourth sending module is configured to send the metadata to the audio renderer, which is used to render audio data with the audio occlusion effect based on the occlusion object information field.
[0415] In one embodiment of this application, the occlusion object information field includes an occlusion position field, which is used to indicate the position coordinates of the occlusion object in a coordinate system.
[0416] In one embodiment of this application, the occlusion object information field includes an occlusion size field, which is used to indicate the size coordinates of the occlusion object in a coordinate system.
[0417] In one embodiment of this application, the occlusion object information field includes a vertex position field, which is used to indicate the maximum and minimum position coordinates of vertices on the occlusion object.
[0418] In one embodiment of this application, the occlusion object information field includes an object identifier field, which is used to indicate an identifier that distinguishes different occlusion objects.
[0419] In one embodiment of this application, the fourth assignment module is further configured to assign a value to the occlusion parameter field in the metadata based on the attribute information of the occluding object, wherein the occlusion parameter field is used to indicate parameters that affect the audio occlusion effect.
[0420] In one embodiment of this application, the occlusion parameter field includes an occlusion type field, which indicates the type of algorithm used when rendering the audio occlusion effect.
[0421] In one embodiment of this application, the algorithm type includes achieving the audio occlusion effect by adjusting the volume or by filtering high-frequency signals.
[0422] In one embodiment of this application, the occlusion parameter field includes an occlusion intensity field, which is used to indicate the intensity of the audio occlusion effect being rendered.
[0423] In one embodiment of this application, the occlusion parameter field includes a maximum occlusion intensity field, which indicates the maximum intensity of the audio occlusion effect on the audio data.
[0424] In one embodiment of this application, the occlusion parameter field includes a corresponding object identifier field, which is used to indicate the identifier of one or more occlusion objects referenced when rendering the audio occlusion effect.
[0425] In one embodiment of this application, an audio rendering apparatus is provided for performing audio rendering based on occlusion objects at an audio receiving end. The apparatus includes:
[0426] The fourth extraction module is configured to extract the occlusion object information field from the metadata, the occlusion object information field being used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0427] The fourth rendering module is configured to render audio data with the audio occlusion effect based on the occlusion object information field.
[0428] In one embodiment of this application, an audio rendering device is provided that performs audio rendering based on diverse scene configurations at an audio acquisition end. The device includes:
[0429] The fifth acquisition module is configured to acquire the scene configuration information of the receiver, which is the configuration information that affects the rendering effect of audio data in the rendering scene;
[0430] The fifth assignment module is configured to assign values to the feature fields of the metadata according to the scene configuration information. The feature fields include at least one of the following: receiver field, device information field, audio component information field, or occlusion object information field. The receiver field includes a field related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with the corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0431] The fifth sending module is configured to send the metadata to an audio renderer, which renders the audio data based on the feature fields.
[0432] In one embodiment of this application, an audio rendering apparatus is provided for performing audio rendering based on diverse scene configurations at an audio receiving end. The apparatus includes:
[0433] The fifth extraction module is configured to extract feature fields from metadata. The feature fields include at least one of a receiver field, a device information field, an audio component information field, or an occlusion object information field. The receiver field includes a field related to the location information of the receiver. The device information field is used to indicate the characteristics of the audio playback device that affects the audio data rendering effect. The audio component information field is used to indicate the audio component with a corresponding rendering type processed by the audio renderer. The occlusion object information field is used to indicate the characteristics of the occlusion object that produces the audio occlusion effect.
[0434] The fifth rendering module is configured to render audio data based on the feature fields.
[0435] The specific details of the audio rendering apparatus provided in the various embodiments of this application have been described in detail in the corresponding method embodiments, and will not be repeated here.
[0436] Figure 19 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.
[0437] It should be noted that, Figure 19 The computer system 1900 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0438] like Figure 19As shown, the computer system 1900 includes a central processing unit (CPU) 1901, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1902 or programs loaded from storage section 1908 into random access memory (RAM) 1903. The RAM 1903 also stores various programs and data required for system operation. The CPU 1901, ROM 1902, and RAM 1903 are interconnected via a bus 1904. An input / output interface 1905 (I / O interface) is also connected to the bus 1904.
[0439] The following components are connected to the input / output interface 1905: an input section 1906 including a keyboard, mouse, etc.; an output section 1907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1908 including a hard disk, etc.; and a communication section 1909 including a network interface card such as a local area network card, modem, etc. The communication section 1909 performs communication processing via a network such as the Internet. A drive 1910 is also connected to the input / output interface 1905 as needed. Removable media 1911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on the drive 1910 as needed so that computer programs read from them can be installed into the storage section 1908 as needed.
[0440] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1909, and / or installed from removable medium 1911. When the computer program is executed by central processing unit 1901, it performs various functions defined in the system of this application.
[0441] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0442] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0443] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0444] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0445] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0446] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An audio rendering method, characterized by, The method comprises: obtaining quantity information of the receivers, the quantity information being used to indicate a number of receivers receiving the audio data; when the number of the receivers is multiple, configuring multiple receiver fields respectively corresponding to the multiple receivers in the metadata according to the quantity information, the receiver fields comprising feature fields related to position information of the receivers; sending the metadata to an audio renderer, the audio renderer being used to render the audio data differently for the multiple receivers according to the multiple receiver fields respectively.
2. The audio rendering method of claim 1, wherein, After obtaining the quantity information of the receivers, the method further comprises: assigning a multi-receiver flag field in the metadata according to the quantity information, the multi-receiver flag field being used to indicate that the number of the receivers is one or multiple.
3. The audio rendering method of claim 2, wherein, The metadata comprises a presentation information field, the presentation information field being an element related to content of the audio data, and the multi-receiver flag field being an attribute or a sub-element of the presentation information field.
4. The audio rendering method of any of claims 1-3, wherein, Before sending the metadata to the audio renderer, the method further comprises: obtaining freedom degree information of the receivers, the freedom degree information being used to indicate freedom degrees of the receivers when receiving the audio; assigning a receiver freedom degree field in the metadata according to the freedom degree information, the receiver freedom degree field being used to indicate that the receivers have three freedom degrees or six freedom degrees.
5. The audio rendering method of any of claims 1-3, wherein, Before sending the metadata to the audio renderer, the method further comprises: when the number of the receivers is multiple, assigning receiver identifier fields corresponding to the receivers in the metadata, the receiver identifier fields being used to distinguish the different receivers.
6. The audio rendering method of any one of claims 1-3, wherein, Before sending the metadata to the audio renderer, the method further comprises: when the number of the receivers is multiple, obtaining position information of the receivers respectively, the position information being used to indicate position coordinates of the receivers in a coordinate system used; assigning coordinate system type fields corresponding to the receivers in the metadata according to the position information, the coordinate system type fields being used to indicate that the coordinate system is a local coordinate system or a world coordinate system.
7. The audio rendering method of claim 6, wherein, After obtaining the position information of the receivers respectively, the method further comprises: when the coordinate system is the local coordinate system, assigning an origin offset field in the metadata according to the position information, the origin offset field being used to indicate an offset amount of a coordinate origin of the local coordinate system in the world coordinate system.
8. The audio rendering method of claim 7, wherein, The origin offset field comprises offset component fields corresponding to coordinate axes, the offset component fields being used to indicate components of the offset amount on the coordinate axes.
9. The audio rendering method of claim 6, wherein, After obtaining the position information of the receivers respectively, the method further comprises: determining whether the receivers are at the same position according to the position information; If the first receiver and the second receiver are in the same position, a position field corresponding to the first receiver is assigned a value according to the position information, and a same-position receiver identifier field corresponding to the second receiver is assigned a value, the position field being used to indicate position coordinates of the first receiver in a coordinate system, and the same-position receiver identifier field being used to indicate the first receiver in the same position as the second receiver.
10. The audio rendering method of claim 1, wherein, The method further comprises: obtaining device information of a receiver, the device information being used to indicate an audio playback device used by the receiver when receiving audio data; assigning a device information field in metadata a value according to the device information, the device information field being used to indicate a feature of the audio playback device that affects rendering effect of the audio data; assigning a corresponding device identifier field in the metadata a value according to the device information field, the corresponding device identifier field being a sub-element of a rendering information field, the rendering information field being an element in the metadata used to indicate the rendering effect of the audio data; sending the metadata to an audio renderer, the audio renderer being used to render the audio data for the audio playback device according to the device information field.
11. The audio rendering method of claim 10, wherein, The device information field comprises a device identifier field or a device group identifier field, the device identifier field being used to indicate an identifier that distinguishes different audio playback devices, and the device group identifier field being used to indicate a device group to which the audio playback device belongs, audio playback devices belonging to the same device group having the same audio data rendering effect.
12. The audio rendering method of claim 10, wherein, The device information field comprises a device name field or a device type field, the device name field being used to indicate a readable name of the audio playback device, and the device type field being used to indicate a device type of the audio playback device.
13. The audio rendering method of claim 12, wherein, The device information field comprises a device position field, the device position field being used to indicate a distribution position of the audio playback device relative to the receiver.
14. The audio rendering method of claim 13, wherein, When the device type of the audio playback device is earphones, the distribution position comprises a left ear or a right ear; when the device type of the audio playback device is a loudspeaker, the distribution position comprises a center, a front left, a front right, a back left, or a back right.
15. The audio rendering method of claim 10, wherein, The device information field comprises a corresponding receiver identifier field, the corresponding receiver identifier field being used to indicate an identifier of the receiver corresponding to the audio playback device.
16. The audio rendering method of claim 1, wherein, The method further comprises: obtaining rendering type information of audio data, the rendering type information being used to indicate a rendering type corresponding to each audio component in the audio data; assigning an audio component information field in metadata a value according to the rendering type information, the audio component information field being used to indicate an audio component with the corresponding rendering type processed by an audio renderer; sending the metadata to the audio renderer, the audio renderer being used to render the audio component according to the rendering type.
17. The audio rendering method of claim 16, wherein, The audio component information field comprises a component type field, the component type field being used to indicate that the rendering type corresponding to the audio component is immersive audio rendering or non-immersive audio rendering.
18. The audio rendering method of claim 16, wherein, The audio component information field comprises a component identifier field or a component group identifier field, the component identifier field being used to indicate an identifier distinguishing different audio components; the component group identifier field being used to indicate a component group to which the audio component belongs, audio components belonging to the same component group being used as input parameters of the audio renderer together.
19. The audio rendering method of claim 16, wherein, The audio component information field comprises a corresponding audio format identifier field, the corresponding audio format identifier field being used to indicate an identifier of an audio format corresponding to the audio component in an audio definition model.
20. The audio rendering method of claim 1, wherein, The method further comprises: obtaining attribute information of an occlusion object located in a rendering scene of audio data; assigning a value to an occlusion object information field in the metadata according to the attribute information of the occlusion object, the occlusion object information field being used to indicate a feature of the occlusion object generating an audio occlusion effect; sending the metadata to an audio renderer, the audio renderer being used to render audio data with the audio occlusion effect according to the occlusion object information field.
21. The audio rendering method of claim 20, wherein, The occlusion object information field comprises an occlusion position field, the occlusion position field being used to indicate position coordinates of the occlusion object in a coordinate system.
22. The audio rendering method of claim 20, wherein, The occlusion object information field comprises an occlusion size field, the occlusion size field being used to indicate size coordinates of the occlusion object in the coordinate system.
23. The audio rendering method of claim 20, wherein, The occlusion object information field comprises a vertex position field, the vertex position field being used to indicate maximum and minimum position coordinates of a vertex on the occlusion object.
24. The audio rendering method of claim 20, wherein, The occlusion object information field comprises an object identifier field, the object identifier field being used to indicate an identifier distinguishing different occlusion objects.
25. The audio rendering method of claim 20, wherein, After obtaining the attribute information of the occlusion object located in the rendering scene of the audio data, the method further comprises: assigning a value to an occlusion parameter field in the metadata according to the attribute information of the occlusion object, the occlusion parameter field being used to indicate a parameter affecting the audio occlusion effect.
26. The audio rendering method of claim 25, wherein, The occlusion parameter field comprises an occlusion type field, the occlusion type field being used to indicate a type of algorithm used when rendering the audio occlusion effect.
27. The audio rendering method of claim 26, wherein, The type of algorithm comprises adjusting a volume to achieve the audio occlusion effect, or filtering a high-frequency signal to achieve the audio occlusion effect.
28. The audio rendering method of claim 25, wherein, The occlusion parameter field comprises an occlusion intensity field, the occlusion intensity field being used to indicate an intensity of rendering the audio occlusion effect.
29. The audio rendering method of claim 25, wherein, The occlusion parameter field comprises a maximum occlusion intensity field, the maximum occlusion intensity field being used to indicate a maximum intensity of the audio occlusion effect on the audio data.
30. The audio rendering method of claim 25, wherein, The occlusion parameter field comprises a corresponding object identifier field, the corresponding object identifier field being used to indicate identifiers of one or more occlusion objects referenced when rendering the audio occlusion effect.
31. A method of audio rendering, the method comprising: The method further comprises: extracting a receiver field corresponding to a receiver of the received audio data from the metadata, the receiver field comprising a feature field related to position information of the receiver; if the number of receivers is multiple, rendering the audio data for the multiple receivers differently according to the multiple receiver fields respectively.
32. The audio rendering method of claim 31, wherein, The method further comprises: extract a corresponding device identification field from a rendering information field of the metadata, the rendering information field being an element in the metadata for indicating a rendering effect of the audio data; extract a device information field according to the corresponding device identification field, the device information field being for indicating a feature of an audio playback device affecting the rendering effect of the audio data; render the audio data for the audio playback device according to the device information field.
33. The audio rendering method of claim 31, wherein, The method further comprises: extract an audio component information field from the metadata, the audio component information field being for indicating an audio component with a corresponding rendering type processed by an audio renderer; determine the rendering type corresponding to each audio component in the audio data according to the audio component information field; render the audio component according to the rendering type.
34. The audio rendering method of claim 31, wherein, The method further comprises: extract an occlusion object information field from the metadata, the occlusion object information field being for indicating a feature of an occlusion object producing an audio occlusion effect; render the audio data with the audio occlusion effect according to the occlusion object information field.
35. An audio rendering apparatus, characterized by comprises: a first obtaining module configured to obtain quantity information of receivers, the quantity information being for indicating a quantity of receivers receiving audio data; a first assigning module configured to, when the quantity of receivers is multiple, configure multiple receiver fields respectively corresponding to multiple receivers in metadata according to the quantity information, the receiver field including a feature field related to position information of the receiver; a first sending module configured to send the metadata to an audio renderer, the audio renderer being for rendering the audio data for multiple receivers differently according to multiple receiver fields respectively.
36. The audio rendering apparatus of claim 35, wherein, The apparatus further comprises: a second obtaining module configured to obtain device information of receivers, the device information being for indicating an audio playback device used by the receiver when receiving audio data; a second assigning module configured to assign a value to a device information field in metadata according to the device information, the device information field being for indicating a feature of the audio playback device affecting a rendering effect of the audio data; and assign a value to a corresponding device identification field in the metadata according to the device information field, the corresponding device identification field being a sub-element of a rendering information field, the rendering information field being an element in the metadata for indicating the rendering effect of the audio data; a second sending module configured to send the metadata to an audio renderer, the audio renderer being for rendering the audio data for the audio playback device according to the device information field.
37. The audio rendering apparatus of claim 35, wherein, The apparatus further comprises: a third obtaining module configured to obtain rendering type information of audio data, the rendering type information being for indicating a rendering type corresponding to each audio component in the audio data; a third assigning module configured to assign a value to an audio component information field in metadata according to the rendering type information, the audio component information field being for indicating an audio component with a corresponding rendering type processed by an audio renderer; The third sending module is configured to send the metadata to the audio renderer, and the audio renderer is used to render the audio component according to the rendering type.
38. The audio rendering apparatus of claim 35, wherein, The device further comprises: The fourth obtaining module is configured to obtain attribute information of an occlusion object located in a rendering scene of audio data. The fourth assignment module is configured to assign a value to an occlusion object information field in the metadata according to the attribute information of the occlusion object, and the occlusion object information field is used to indicate a feature of the occlusion object that produces an audio occlusion effect. The fourth sending module is configured to send the metadata to the audio renderer, and the audio renderer is used to render audio data with the audio occlusion effect according to the occlusion object information field.
39. An audio rendering apparatus, comprising: Comprise: The first extraction module is configured to extract a receiver field corresponding to a receiver of audio data from metadata, and the receiver field comprises a feature field related to position information of the receiver. The first rendering module is configured to, if the number of receivers is multiple, render the audio data for the multiple receivers differently according to the multiple receiver fields.
40. The audio rendering apparatus of claim 39, wherein, The device further comprises: The second extraction module is configured to extract a corresponding device identifier field from a rendering information field of the metadata, and the rendering information field is an element in the metadata used to indicate an audio data rendering effect; and extract a device information field according to the corresponding device identifier field, and the device information field is used to indicate a feature of an audio playback device that affects the audio data rendering effect. The second rendering module is configured to render the audio data for the audio playback device according to the device information field.
41. The audio rendering apparatus of claim 39, wherein, The device further comprises: The third extraction module is configured to extract an audio component information field from the metadata, and the audio component information field is used to indicate an audio component with a corresponding rendering type processed by an audio renderer. The third rendering module is configured to determine a rendering type corresponding to each audio component in the audio data according to the audio component information field; and render the audio component according to the rendering type.
42. The audio rendering apparatus of claim 39, wherein, The device further comprises: The fourth extraction module is configured to extract an occlusion object information field from the metadata, and the occlusion object information field is used to indicate a feature of an occlusion object that produces an audio occlusion effect. The fourth rendering module is configured to render audio data with the audio occlusion effect according to the occlusion object information field.
43. A computer readable medium characterized by The computer program is stored on the computer readable medium and is executed by the processor to implement the audio rendering method in any one of claims 1 to 34.
44. An electronic device, comprising: Comprise: A processor; And A memory for storing executable instructions of the processor; Wherein the processor is configured to execute the executable instructions to enable the electronic device to perform the audio rendering method in any one of claims 1 to 34.
45. A computer program product comprising a computer program, characterised in that, The computer program is executed by the processor to implement the audio rendering method in any one of claims 1 to 34. The computer program is executed by the processor to implement the audio rendering method in any one of claims 1 to 34.
Citation Information
Patent Citations
An audio processor and a method for providing loudspeaker signals
CN113016197A
Audio processing method and device, electronic equipment and computer readable storage medium
CN114339297A