Audio signal rendering method and apparatus

By acquiring control information and adaptively selecting the rendering method, the problem of insufficient audio signal rendering effect in virtual reality systems is solved, and flexible format conversion and efficient rendering of audio signals are realized, thereby enhancing the audio immersion.

CN114067810BActive Publication Date: 2025-12-12HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010763577.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-12-12
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of improvement in the rendering effect of audio signals, especially in virtual reality systems where the rendering effect of audio signals in audiovisual fusion is insufficient.

Method used

By decoding the received bitstream to obtain control information, and utilizing content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, and attitude information, the rendering method is adaptively selected. This includes techniques such as pre-processing, signal format conversion, local reverb processing, group processing, dynamic range compression, binaural rendering, and speaker rendering to improve the rendering effect of the audio signal.

Benefits of technology

It achieves adaptive audio rendering based on different application scenarios and terminal device performance, improving the rendering effect and immersiveness of audio signals, and is suitable for various signal formats and terminal devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067810B_ABST
    Figure CN114067810B_ABST
Patent Text Reader

Abstract

The application provides an audio signal rendering method and device. The audio signal rendering method can include: an audio signal rendering method, characterized by comprising: obtaining a to-be-rendered audio signal by decoding a received code stream; obtaining control information, the control information being used for indicating at least one of content description metadata, rendering format flag information, loudspeaker configuration information, application scenario information, tracking information, posture information or position information; and rendering the to-be-rendered audio signal according to the control information to obtain a rendered audio signal. The embodiments of the application can improve the rendering effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, and in particular to an audio signal rendering method and device. BACKGROUND

[0002] With the continuous development of multimedia technology, audio has been widely applied in the fields of multimedia communication, consumer electronics, virtual reality, human-computer interaction, etc. Users have higher and higher requirements for audio quality. Three-dimensional audio (3D audio) has close-to-real spatial sense and can provide users with better immersive experience, and has become a new trend of multimedia technology.

[0003] Taking virtual reality (VR) as an example, a VR system with immersion not only needs a shocking visual effect, but also needs a realistic auditory effect. The fusion of vision and hearing can greatly improve the experience of virtual reality, and the core of the audio of virtual reality is three-dimensional audio technology. Based on sound channel, based on object, and based on scene are three common formats in three-dimensional audio technology. By rendering the decoded audio signals based on sound channel, object and scene, the audio signal playback can be realized to achieve realistic and immersive auditory experience.

[0004] Among them, how to improve the rendering effect of the audio signal has become a technical problem to be solved. SUMMARY

[0005] The present application provides an audio signal rendering method and device, which is beneficial to improve the rendering effect of the audio signal.

[0006] In a first aspect, an audio signal rendering method is provided, which can include: obtaining a to-be-rendered audio signal by decoding a received code stream. Control information is obtained, the control information being used to indicate one or more of content description metadata, rendering format flag information, loudspeaker configuration information, application scenario information, tracking information, posture information or position information. The to-be-rendered audio signal is rendered according to the control information to obtain a rendered audio signal.

[0007] The content description metadata is used to indicate a signal format of the audio signal to be rendered. The signal format comprises at least one of a channel-based signal format, a scene-based signal format, or an object-based signal format. The rendering format flag information is used to indicate an audio signal rendering format. The audio signal rendering format comprises a loudspeaker rendering or a binaural rendering. The loudspeaker configuration information is used to indicate a layout of loudspeakers. The application scenario information is used to indicate renderer scene description information. The tracking information is used to indicate whether the rendered audio signal changes with a head rotation of a listener. The pose information is used to indicate an orientation and an amplitude of the head rotation. The position information is used to indicate an orientation and an amplitude of a body movement of the listener.

[0008] According to the adaptive selection of the rendering mode based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the pose information, or the position information, the audio rendering effect can be improved.

[0009] In a possible design, the rendering of the audio signal to be rendered according to the control information comprises at least one of the following: rendering pre-processing of the audio signal to be rendered according to the control information; or signal format conversion of the audio signal to be rendered according to the control information; or local reverberation processing of the audio signal to be rendered according to the control information; or group processing of the audio signal to be rendered according to the control information; or dynamic range compression of the audio signal to be rendered according to the control information; or binaural rendering of the audio signal to be rendered according to the control information; or loudspeaker rendering of the audio signal to be rendered according to the control information.

[0010] According to the adaptive selection of the rendering mode based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the pose information, or the position information, the audio rendering effect can be improved.

[0011] In a possible design, the to-be-rendered audio signal includes at least one of a channel-based audio signal, an object-based audio signal, or a scene-based audio signal, and when the to-be-rendered audio signal is rendered according to the control information, including pre-processing of the to-be-rendered audio signal according to the control information, the method further includes: obtaining first reverberation information from the code stream, the first reverberation information including at least one of first reverberation output loudness information, first time difference information of direct sound and early reflection sound, first reverberation duration information, first room shape and size information, or first sound scattering degree information. Correspondingly, the pre-processing of the to-be-rendered audio signal according to the control information to obtain the rendered audio signal can include: performing control processing on the to-be-rendered audio signal according to the control information to obtain a control-processed audio signal, the control processing including at least one of performing initial three-degree-of-freedom (3DoF) processing on the channel-based audio signal, performing transformation processing on the object-based audio signal, or performing initial 3DoF processing on the scene-based audio signal, and performing reverberation processing on the control-processed audio signal according to the first reverberation information to obtain a first audio signal. The first audio signal is binaurally rendered or loudspeaker rendered to obtain the rendered audio signal.

[0012] In a possible design, when the to-be-rendered audio signal is rendered according to the control information, the signal format conversion further includes: performing signal format conversion on the first audio signal according to the control information to obtain a second audio signal. The second audio signal is binaurally rendered or loudspeaker rendered to obtain the rendered audio signal.

[0013] The signal format conversion includes at least one of the following: converting the channel-based audio signal in the first audio signal into a scene-based or object-based audio signal; or converting the scene-based audio signal in the first audio signal into a channel-based or object-based audio signal; or converting the object-based audio signal in the first audio signal into a channel-based or scene-based audio signal.

[0014] In the implementation, the signal format conversion on the to-be-rendered audio signal according to the control information can achieve flexible conversion of the signal format, so that the audio signal rendering method of the embodiments of the present application is applicable to any signal format, and the audio rendering effect can be improved by rendering the audio signal in a suitable signal format.

[0015] In a possible design, the signal format conversion on the first audio signal according to the control information can include: performing signal format conversion on the first audio signal according to the control information, a signal format of the first audio signal, and a processing capability of the terminal device.

[0016] In the implementation manner, the first audio signal is converted in a signal format according to the processing capability of the terminal device, so that the signal format is matched with the processing capability of the terminal device, rendering is performed, and the audio rendering effect is optimized.

[0017] In a possible design, when the to-be-rendered audio signal is rendered according to the control information, the local reverberation processing on the to-be-rendered audio signal can further include: performing binaural rendering or loudspeaker rendering on the second audio signal to obtain the rendered audio signal, which can include: obtaining second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information including at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information; performing local reverberation processing on the second audio signal according to the control information and the second reverberation information to obtain a third audio signal; and performing binaural rendering or loudspeaker rendering on the third audio signal to obtain the rendered audio signal.

[0018] In the implementation manner, the second reverberation information corresponding to the real-time input application scene information can be generated, and the second reverberation information can be used for rendering processing, so that the audio rendering effect is improved, and real-time reverberation corresponding to a scene can be provided for an AR application scene.

[0019] In a possible design, the local reverberation processing on the second audio signal according to the control information and the second reverberation information to obtain a third audio signal can include: performing clustering processing on different signal formats of the audio signal in the second audio signal according to the control information to obtain at least one of a channel-based group signal, a scene-based group signal, or an object-based group signal; and performing local reverberation processing on at least one of the channel-based group signal, the scene-based group signal, or the object-based group signal according to the second reverberation information to obtain the third audio signal.

[0020] In a possible design, when the to-be-rendered audio signal is rendered according to the control information, the binaural rendering or the loudspeaker rendering is further performed on the third audio signal according to the control information to obtain the rendered audio signal, which can include: performing real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on a group signal of each signal format in the third audio signal according to the control information to obtain a fourth audio signal; and performing binaural rendering or loudspeaker rendering on the fourth audio signal to obtain the rendered audio signal.

[0021] According to the implementation, uniform processing is performed on audio signals of each format, and processing complexity can be reduced on the basis of ensuring processing performance.

[0022] In a possible design, when the to-be-rendered audio signal is rendered according to the control information, the binaural rendering or the loudspeaker rendering is further performed on the fourth audio signal according to the control information to obtain the rendered audio signal, which can include: performing dynamic range compression on the fourth audio signal according to the control information to obtain a fifth audio signal; and performing binaural rendering or loudspeaker rendering on the fifth audio signal to obtain the rendered audio signal.

[0023] According to the implementation, dynamic range compression is performed on an audio signal according to control information, to improve the playback quality of a rendered audio signal.

[0024] In a possible design, the to-be-rendered audio signal is rendered according to the control information to obtain a rendered audio signal, which can include: performing signal format conversion on the to-be-rendered audio signal according to the control information to obtain a sixth audio signal; and performing binaural rendering or loudspeaker rendering on the sixth audio signal to obtain the rendered audio signal.

[0025] The signal format conversion includes at least one of the following: converting a channel-based audio signal in the to-be-rendered audio signal into a scene-based or object-based audio signal; or converting a scene-based audio signal in the to-be-rendered audio signal into a channel-based or object-based audio signal; or converting an object-based audio signal in the to-be-rendered audio signal into a channel-based or scene-based audio signal.

[0026] In a possible design, the signal format conversion on the to-be-rendered audio signal according to the control information can include: performing signal format conversion on the to-be-rendered audio signal according to the control information, a signal format of the to-be-rendered audio signal, and processing performance of a terminal device.

[0027] The terminal device can be a device that executes the audio signal rendering method of the first aspect of the embodiments of the present application. The implementation manner can convert the signal format of the to-be-rendered audio signal in combination with the processing performance of the terminal device, so that the audio signal rendering is suitable for terminal devices with different performance.

[0028] For example, the signal format conversion can be performed in combination with the processing performance of the terminal device from two dimensions of algorithm complexity of the audio signal rendering method and rendering effect. For example, when the processing performance of the terminal device is good, the to-be-rendered audio signal can be converted into a signal format with better rendering effect, even if the algorithm complexity corresponding to the signal format with better rendering effect is high. When the processing performance of the terminal device is poor, the to-be-rendered audio signal can be converted into a signal format with lower algorithm complexity, so as to ensure the rendering output efficiency. The processing performance of the terminal device can be the processor performance of the terminal device. For example, when the frequency of the processor of the terminal device is greater than a certain threshold and the bit number is greater than a certain threshold, the processing performance of the terminal device is good. The specific implementation manner of the signal format conversion in combination with the processing performance of the terminal device can also be other manners, for example, based on a preset corresponding relationship and the model of the processor of the terminal device, a processing performance parameter value of the terminal device is obtained, and when the processing performance parameter value is greater than a certain threshold, the to-be-rendered audio signal is converted into a signal format with better rendering effect. The embodiments of the present application do not enumerate all the examples. The signal format with better rendering effect can be determined based on the control information.

[0029] In a possible design, the rendering of the to-be-rendered audio signal according to the control information to obtain the rendered audio signal can include: obtaining second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information including at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information. The to-be-rendered audio signal is subjected to local reverberation processing according to the control information and the second reverberation information to obtain a seventh audio signal. The seventh audio signal is subjected to binaural rendering or loudspeaker rendering to obtain the rendered audio signal.

[0030] In a possible design, the rendering of the to-be-rendered audio signal according to the control information to obtain the rendered audio signal can include: performing real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom 6DoF processing on the audio signal of each signal format in the to-be-rendered audio signal according to the control information to obtain an eighth audio signal. The eighth audio signal is subjected to binaural rendering or loudspeaker rendering to obtain the rendered audio signal.

[0031] In a possible design, the rendering of the to-be-rendered audio signal according to the control information to obtain a rendered audio signal can include: performing dynamic range compression on the to-be-rendered audio signal according to the control information to obtain a ninth audio signal. The ninth audio signal is subjected to binaural rendering or loudspeaker rendering to obtain the rendered audio signal.

[0032] In a second aspect, an embodiment of the present application provides an audio signal rendering apparatus. The audio signal rendering apparatus can be an audio renderer, a chip or a system on chip of an audio decoding device, or a functional module in the audio renderer for implementing the method in the first aspect or any possible design of the first aspect. The audio signal rendering apparatus can implement the functions performed in the first aspect or any possible design of the first aspect. The functions can be implemented by hardware or software. The hardware or software includes one or more modules corresponding to the functions. For example, in a possible design, the audio signal rendering apparatus can include: an obtaining module configured to obtain a to-be-rendered audio signal by decoding a received bitstream; a control information generation module configured to obtain control information, the control information being used to indicate one or more of content description metadata, rendering format flag information, loudspeaker configuration information, application scenario information, tracking information, pose information, or position information; and a rendering module configured to render the to-be-rendered audio signal according to the control information to obtain a rendered audio signal.

[0033] The content description metadata is used to indicate a signal format of the to-be-rendered audio signal. The signal format includes at least one of channel-based, scene-based, or object-based. The rendering format flag information is used to indicate an audio signal rendering format. The audio signal rendering format includes loudspeaker rendering or binaural rendering. The loudspeaker configuration information is used to indicate a layout of loudspeakers. The application scenario information is used to indicate renderer scene description information. The tracking information is used to indicate whether the rendered audio signal changes with a head rotation of a listener. The pose information is used to indicate a direction and an amplitude of the head rotation. The position information is used to indicate a direction and an amplitude of a body movement of the listener.

[0034] In a possible design, the rendering module is configured to perform at least one of the following: pre-processing of the to-be-rendered audio signal according to the control information; or signal format conversion of the to-be-rendered audio signal according to the control information; or local reverberation processing of the to-be-rendered audio signal according to the control information; or group processing of the to-be-rendered audio signal according to the control information; or dynamic range compression of the to-be-rendered audio signal according to the control information; or binaural rendering of the to-be-rendered audio signal according to the control information; or loudspeaker rendering of the to-be-rendered audio signal according to the control information.

[0035] In a possible design, the to-be-rendered audio signal comprises at least one of a channel-based audio signal, an object-based audio signal, or a scene-based audio signal, and the obtaining module is further configured to: obtain, by decoding the code stream, first reverberation information, the first reverberation information comprising at least one of first reverberation output loudness information, first time difference information of direct sound and early reflection sound, first reverberation duration information, first room shape and size information, or first sound scattering degree information. Correspondingly, the rendering module is configured to: perform control processing on the to-be-rendered audio signal according to the control information, to obtain a to-be-controlled audio signal, the control processing comprising at least one of performing initial three-degree-of-freedom (3DoF) processing on the channel-based audio signal, performing transformation processing on the object-based audio signal, or performing initial 3DoF processing on the scene-based audio signal, and perform reverberation processing on the to-be-controlled audio signal according to the first reverberation information, to obtain a first audio signal. Perform binaural rendering or loudspeaker rendering on the first audio signal, to obtain the rendered audio signal.

[0036] In a possible design, the rendering module is configured to: perform signal format conversion on the first audio signal according to the control information, to obtain a second audio signal. Perform binaural rendering or loudspeaker rendering on the second audio signal, to obtain the rendered audio signal.

[0037] In a possible design, the signal format conversion comprises at least one of: converting channel-based audio signal in the first audio signal into scene-based or object-based audio signal; or converting scene-based audio signal in the first audio signal into channel-based or object-based audio signal; or converting object-based audio signal in the first audio signal into channel-based or scene-based audio signal.

[0038] In a possible design, the rendering module is configured to: perform signal format conversion on the first audio signal according to the control information, a signal format of the first audio signal, and a processing capability of the terminal device.

[0039] In a possible design, the rendering module is configured to: obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information comprising at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information. Perform local reverberation processing on the second audio signal according to the control information and the second reverberation information, to obtain a third audio signal. Perform binaural rendering or loudspeaker rendering on the third audio signal, to obtain the rendered audio signal.

[0040] In a possible design, the rendering module is configured to: perform clustering processing on audio signals of different signal formats in the second audio signals respectively according to the control information, to obtain at least one of a channel-based group signal, a scene-based group signal, or an object-based group signal; and perform local reverberation processing on the at least one of the channel-based group signal, the scene-based group signal, or the object-based group signal respectively according to the second reverberation information, to obtain the third audio signals.

[0041] In a possible design, the rendering module is configured to: perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on the group signals of each signal format in the third audio signals according to the control information, to obtain fourth audio signals; and perform binaural rendering or loudspeaker rendering on the fourth audio signals, to obtain the rendered audio signals.

[0042] In a possible design, the rendering module is configured to: perform dynamic range compression on the fourth audio signals according to the control information, to obtain fifth audio signals; and perform binaural rendering or loudspeaker rendering on the fifth audio signals, to obtain the rendered audio signals.

[0043] In a possible design, the rendering module is configured to: perform signal format conversion on the to-be-rendered audio signals according to the control information, to obtain sixth audio signals; and perform binaural rendering or loudspeaker rendering on the sixth audio signals, to obtain the rendered audio signals.

[0044] The signal format conversion includes at least one of the following: converting channel-based audio signals in the to-be-rendered audio signals into scene-based or object-based audio signals; or converting scene-based audio signals in the to-be-rendered audio signals into channel-based or object-based audio signals; or converting object-based audio signals in the to-be-rendered audio signals into channel-based or scene-based audio signals.

[0045] In a possible design, the rendering module is configured to: perform signal format conversion on the to-be-rendered audio signals according to the control information, a signal format of the to-be-rendered audio signals, and a processing capability of the terminal device.

[0046] In a possible design, the rendering module is configured to: obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the to-be-rendered audio signal is located, the second reverberation information including at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information; perform local reverberation processing on the to-be-rendered audio signal according to the control information and the second reverberation information, to obtain a seventh audio signal; and perform binaural rendering or loudspeaker rendering on the seventh audio signal, to obtain the rendered audio signal.

[0047] In a possible design, the rendering module is configured to: perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on an audio signal of each signal format in the to-be-rendered audio signal according to the control information, to obtain an eighth audio signal; and perform binaural rendering or loudspeaker rendering on the eighth audio signal, to obtain the rendered audio signal.

[0048] In a possible design, the rendering module is configured to: perform dynamic range compression on the to-be-rendered audio signal according to the control information, to obtain a ninth audio signal; and perform binaural rendering or loudspeaker rendering on the ninth audio signal, to obtain the rendered audio signal.

[0049] In a third aspect, an embodiment of the present application provides an audio signal rendering apparatus, which is characterized by comprising: a non-volatile memory and a processor coupled with each other, and the processor is configured to invoke program code stored in the memory to execute the method in the first aspect or any possible design of the first aspect.

[0050] In a fourth aspect, an embodiment of the present application provides an audio signal decoding device, which is characterized by comprising: a renderer configured to execute the method in the first aspect or any possible design of the first aspect.

[0051] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, which includes a computer program, and the computer program is configured to cause a computer to execute the method in any one of the first aspect when the computer program is executed on the computer.

[0052] In a sixth aspect, the present application provides a computer program product, which includes a computer program, and the computer program is configured to execute the method in any one of the first aspect when the computer program is executed on a computer.

[0053] In a seventh aspect, the present application provides a chip, which includes a processor and a memory, the memory is configured to store a computer program, and the processor is configured to invoke and run the computer program stored in the memory, to execute the method in any one of the first aspect.

[0054] The audio signal rendering method and device of the embodiments of the present application can obtain the to-be-rendered audio signal by decoding the received code stream, obtain the control information, the control information being used for indicating at least one of content description metadata, rendering format flag information, loudspeaker configuration information, application scenario information, tracking information, posture information or position information, render the to-be-rendered audio signal according to the control information to obtain the rendered audio signal, and can realize adaptive selection of the rendering mode based on at least one of the input information of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information, thereby improving the audio rendering effect. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 A schematic diagram of an audio encoding and decoding system instance in the embodiments of the present application;

[0056] Figure 2 A schematic diagram of an audio signal rendering application in the embodiments of the present application;

[0057] Figure 3 A flowchart of an audio signal rendering method in the embodiments of the present application;

[0058] Figure 4 A layout schematic diagram of a loudspeaker in the embodiments of the present application;

[0059] Figure 5 A schematic diagram of generation of control information in the embodiments of the present application;

[0060] Figure 6A A flowchart of another audio signal rendering method in the embodiments of the present application;

[0061] Figure 6B A schematic diagram of pre-processing of rendering in the embodiments of the present application;

[0062] Figure 7 A schematic diagram of loudspeaker rendering provided by the embodiments of the present application;

[0063] Figure 8 A schematic diagram of binaural rendering provided by the embodiments of the present application;

[0064] Figure 9A A flowchart of another audio signal rendering method in the embodiments of the present application;

[0065] Figure 9B A schematic diagram of signal format conversion in the embodiments of the present application;

[0066] Figure 10A A flowchart of another audio signal rendering method in the embodiments of the present application;

[0067] Figure 10B a schematic diagram of a local reverberation processing according to an embodiment of the application;

[0068] Figure 11A a flowchart of another audio signal rendering method according to an embodiment of the application;

[0069] Figure 11B a schematic diagram of a grouped source Transformations according to an embodiment of the application;

[0070] Figure 12A a flowchart of another audio signal rendering method according to an embodiment of the application;

[0071] Figure 12B a schematic diagram of a dynamic range compression according to an embodiment of the application;

[0072] Figure 13A a schematic diagram of an architecture of an audio signal rendering apparatus according to an embodiment of the application;

[0073] Figure 13B a schematic diagram of a detailed architecture of an audio signal rendering apparatus according to an embodiment of the application;

[0074] Figure 14 a schematic diagram of an architecture of an audio signal rendering apparatus according to an embodiment of the application;

[0075] Figure 15 a schematic diagram of an architecture of an audio signal rendering apparatus according to an embodiment of the application. DETAILED DESCRIPTION

[0076] The terms "first", "second", and the like, as used herein do not imply or mean relative importance or order. Furthermore, the terms "comprises", "comprising", "includes", "including", and the like, are inclusive and are intended to cover non-exclusive inclusion, i.e., the process, method, product, or apparatus that comprises a list of steps or units are not necessarily limited to those steps or units that are clearly listed, but can include other steps or units that are not clearly listed or are inherent to such processes, methods, products, or apparatuses.

[0077] It should be understood that in this application, "at least one" means one or more, "multiple" means two or more. "And / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including single or multiple items in any combination. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple, respectively, or some single and some multiple.

[0078] The system architecture to which embodiments of the present application apply is described below. Referring to Figure 1 , Figure 1 An exemplary block diagram of an audio encoding and decoding system 10 to which embodiments of the present application apply is shown. As Figure 1 indicated, the audio encoding and decoding system 10 can include a source device 12 that generates encoded audio data and, thus, the source device 12 can be referred to as an audio encoding apparatus. A destination device 14 can decode the encoded audio data generated by the source device 12 and, thus, the destination device 14 can be referred to as an audio decoding apparatus. Various implementations of the source device 12, the destination device 14, or both, can include one or more processors and a memory coupled to the one or more processors. The memory can include, but is not limited to, RAM, ROM, EEPROM, flash memory or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. The source device 12 and the destination device 14 can comprise various apparatuses, including a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a so-called "smart" phone, a television, an audio speaker, a digital media player, a video game console, an in-vehicle computer, a wireless communication device, any wearable device (e.g., a smart watch, smart glasses), or the like.

[0079] Although Figure 1Source device 12 and destination device 14 are illustrated as separate devices, but device embodiments can also include both source device 12 and destination device 14 or functionality of both, i.e., source device 12 or corresponding functionality and destination device 14 or corresponding functionality, simultaneously. In such embodiments, source device 12 or corresponding functionality and destination device 14 or corresponding functionality can be implemented using the same hardware and / or software, or separate hardware and / or software, or any combination thereof.

[0080] Source device 12 and destination device 14 can be communicatively coupled via a link 13 over which destination device 14 can receive encoded audio data from source device 12. Link 13 can comprise one or more media or devices capable of conveying encoded audio data from source device 12 to destination device 14. In one example, link 13 can comprise one or more communication media that enable source device 12 to transmit encoded audio data directly to destination device 14 in real-time. In such an example, source device 12 can modulate encoded audio data according to a communication standard, such as a wireless communication protocol, and can transmit the modulated audio data to destination device 14. The one or more communication media can include wireless and / or wired communication media, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The one or more communication media can include routers, switches, base stations, or other equipment that facilitate communication from source device 12 to destination device 14.

[0081] Source device 12 includes an encoder 20, and can optionally also include an audio source 16, a preprocessor 18, and a communication interface 22. In implementations, the encoder 20, audio source 16, preprocessor 18, and communication interface 22 can be hardware components of source device 12, or software programs of source device 12. Each is described separately as follows:

[0082] The audio source 16 can comprise or be any kind of sound capturing device, e.g. for capturing real world sound, and / or any kind of audio generating device. The audio source 16 can be a microphone for capturing sound or a memory for storing audio data, and can further comprise any kind of (internal or external) interface for storing previously captured or generated audio data and / or for obtaining or receiving audio data. When the audio source 16 is a microphone, the audio source 16 can e.g. be a local or integrated microphone integrated in the source device; when the audio source 16 is a memory, the audio source 16 can be a local or integrated memory e.g. integrated in the source device. When the audio source 16 comprises an interface, the interface can e.g. be an external interface for receiving audio data from an external audio source, e.g. an external sound capturing device such as a microphone, an external memory or an external audio generating device. The interface can be any kind of interface according to any proprietary or standardized interface protocol, e.g. a wired or wireless interface, an optical interface.

[0083] In the embodiments of the present application, the audio data transmitted by the audio source 16 to the pre-processor 18 can also be referred to as raw audio data 17.

[0084] The pre-processor 18 is configured to receive the raw audio data 17 and to perform pre-processing on the raw audio data 17 to obtain pre-processed audio 19 or pre-processed audio data 19. The pre-processing performed by the pre-processor 18 can comprise filtering, or de-noising, etc.

[0085] The encoder 20 (or audio encoder 20) is configured to receive the pre-processed audio data 19 and to process the pre-processed audio data 19 to provide encoded audio data 21.

[0086] The communication interface 22 can be configured to receive the encoded audio data 21 and to transmit the encoded audio data 21 to the destination device 14 or any other device (e.g. a memory) over the link 13 for storage or direct reconstruction, which can be any kind of device for decoding or storage. The communication interface 22 can e.g. be configured to encapsulate the encoded audio data 21 into a suitable format, e.g. a packet, for transmission over the link 13.

[0087] The destination device 14 comprises a decoder 30, and can additionally comprise a communication interface 28, an audio post-processor 32 and a rendering device 34. These are described separately as follows:

[0088] The communication interface 28 can be configured to receive the encoded audio data 21 from the source device 12 or any other source, such as a storage device, e.g. an encoded audio data storage device. The communication interface 28 can be configured to transmit or receive the encoded audio data 21 via a link 13 between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any kind of network, such as a wired or wireless network or any combination thereof, or any kind of private and public network or any combination thereof. The communication interface 28 can be configured to, for example, de-encapsulate data packets transmitted by the communication interface 22 to obtain the encoded audio data 21.

[0089] The communication interface 28 and the communication interface 22 can both be configured as unidirectional or bidirectional communication interfaces and can be configured to, for example, send and receive messages to establish a connection, to acknowledge and exchange any other information related to the communication link and / or the transmission of data, such as the transmission of encoded audio data.

[0090] The decoder 30 (or decoding unit 30) is configured to receive the encoded audio data 21 and to provide decoded audio data 31 or decoded audio 31.

[0091] The audio post-processor 32 is configured to perform post-processing on the decoded audio data 31 (also referred to as reconstructed audio data) to obtain post-processed audio data 33. The post-processing performed by the audio post-processor 32 can include, for example, rendering or any other processing and can be configured to transmit the post-processed audio data 33 to a rendering device 34. The audio post-processor can be configured to perform the various embodiments described in the following to enable the application of the audio signal rendering methods described in the present application.

[0092] The rendering device 34 is configured to receive the post-processed audio data 33 to play audio to, for example, a user or a viewer. The rendering device 34 can be or can include any kind of playback device for rendering reconstructed sound. The rendering device can include a loudspeaker or a headphone.

[0093] Although, Figure 1 While the source device 12 and the destination device 14 are depicted as separate devices, a device embodiment can also comprise both the source device 12 and the destination device 14 or the functionality of both, i.e. the source device 12 or the corresponding functionality and the destination device 14 or the corresponding functionality. In such embodiments, the source device 12 or the corresponding functionality and the destination device 14 or the corresponding functionality can be implemented using the same hardware and / or software or using separate hardware and / or software or any combination thereof.

[0094] It is obvious to a person skilled in the art on the basis of the description that the functionality of different units or Figure 1The presence and (precise) division of the functionalities of the illustrated source device 12 and / or destination device 14 can differ depending on the actual device and application. The source device 12 and the destination device 14 can comprise any of a variety of devices, including handheld or stationary devices of any category, e.g., a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a video camera, a desktop computer, a set-top box, a television, a camera, a car device, a sound system, a digital media player, an audio game console, an audio streaming device (e.g., a content service server or a content distribution server), a broadcast receiver device, a broadcast transmitter device, smart glasses, a smart watch, etc., and can not use or use any category of operating system.

[0095] The encoder 20 and the decoder 30 can each be implemented as any of a variety of suitable circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combinations thereof. If the techniques are implemented partially in software, a device can store instructions for the software in any suitable non-transitory computer-readable storage medium and can execute the instructions using one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be considered to be one or more processors.

[0096] In some cases, Figure 1 The audio encoding and decoding system 10 illustrated in FIG. 1 is merely an example, and the techniques of this disclosure can be applied to an audio encoding setup (e.g., audio encoding or audio decoding) that does not necessarily include any data communication between the encoding and decoding devices. In other examples, data can be retrieved from local memory, streamed over a network, etc. An audio encoding device can encode data and store the data to memory, and / or an audio decoding device can retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to memory and / or retrieve data from memory and decode the data.

[0097] The encoder described above can be a multi-channel encoder, such as a stereo encoder, a 5.1 channel encoder, or a 7.1 channel encoder, etc. It is of course understood that the encoder described above can also be a mono-channel encoder. The audio post-processor described above can be used to perform the audio signal rendering method described below in the embodiments of the present disclosure to improve the audio playback effect.

[0098] The audio data can also be referred to as an audio signal, the decoded audio data can also be referred to as a to-be-rendered audio signal, and the post-processed audio data can also be referred to as a rendered audio signal. The audio signal in the embodiment of the present application refers to an input signal of the audio rendering device, and the audio signal can include multiple frames. For example, the current frame can specifically refer to a frame in the audio signal, and the rendering of the current frame of the audio signal is exemplarily described in the embodiment of the present application. The embodiment of the present application is used to implement the rendering of the audio signal.

[0099] Figure 2 is a simplified block diagram of an apparatus 200 according to an exemplary embodiment. The apparatus 200 can implement the techniques of the present application. In other words, Figure 2 is a schematic block diagram of an implementation of an encoding or decoding device (referred to as a coding device 200) of the present application. The apparatus 200 can include a processor 210, a memory 230, and a bus system 250. The processor and the memory are connected via the bus system, the memory is configured to store instructions, and the processor is configured to execute the instructions stored in the memory. The memory of the coding device stores program code, and the processor can invoke the program code stored in the memory to execute the methods described in the present application. To avoid repetition, no longer detailed description is given here.

[0100] In the present application, the processor 210 can be a central processing unit (CPU), and the processor 210 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), ready programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0101] The memory 230 can include read-only memory (ROM) devices or random access memory (RAM) devices. Any other suitable type of storage device can also be used as the memory 230. The memory 230 can include code and data 231 accessed by the processor 210 using the bus 250. The memory 230 can further include an operating system 233 and an application program 235.

[0102] The bus system 250 can include not only a data bus, but also a power bus, a control bus, and a status signal bus, etc. However, for the purpose of clear illustration, all the buses are marked as the bus system 250 in the figure.

[0103] Optionally, the decoding device 200 can also include one or more output devices, such as a speaker 270. In one example, the speaker 270 can be headphones or external speakers. The speaker 270 can be connected to the processor 210 via the bus 250.

[0104] The audio signal rendering method of the embodiments of the present application is applicable to audio rendering in voice communication of any communication system, which can be an LTE system, or a 5G system, or a future evolved PLMN system, etc. The audio signal rendering method of the embodiments of the present application is also applicable to audio rendering in VR or augmented reality (AR) or audio playback applications. Of course, it can also be other application scenarios of audio signal rendering, which are not enumerated one by one in the embodiments of the present application.

[0105] Taking VR as an example, at the encoding end, the audio signal A is preprocessed (Audio Preprocessing) after being acquired by the acquisition module, the preprocessing operation includes filtering out the low frequency part in the signal, usually taking 20Hz or 50Hz as the dividing point, extracting the azimuth information in the audio signal, and then performing encoding processing (Audio encoding) and packaging (File / Segment encapsulation), and then sending (Delivery) to the decoding end. The decoding end first performs unpacking (File / Segment decapsulation), then decoding (Audio decoding), rendering (Audio rendering) the decoded signal, and mapping the rendered signal to the headphones or loudspeakers of the listener. The headphones can be independent headphones, or headphones on glasses devices or other wearable devices. The decoded signal can be rendered (Audio rendering) by the audio signal rendering method described in the following embodiments.

[0106] The audio signal rendering of the embodiments of the present application is to convert the audio signal to be rendered into an audio signal of a specific playback format, i.e. the rendered audio signal, so that the rendered audio signal is adapted to at least one of the playback environment or the playback device, thereby improving the user's auditory experience. The playback device can be the rendering device 34 described above, and can include headphones or loudspeakers. The playback environment can be the environment where the playback device is located. The specific processing method used for audio signal rendering can be referred to the explanation and description of the following embodiments.

[0107] The audio signal rendering apparatus can execute the audio signal rendering method of the embodiments of the present application to adaptively select a rendering processing mode and improve the rendering effect of the audio signal. The audio signal rendering apparatus can be an audio post-processor in the destination device described above, and the destination device can be any terminal device, such as a mobile phone, a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, or the like. The specific implementation of the embodiments can be seen from the following detailed description of the embodiments of the present application. Figure 3 The destination device can also be referred to as a playback end or a rendering end or a decoding and rendering end, and the like.

[0108] Figure 3 For the flowchart of the audio signal rendering method of the embodiments of the present application, the execution subject of the embodiments of the present application can be the audio signal rendering apparatus described above, such as Figure 3 The method of the embodiments can include the following steps.

[0109] In step 401, the received bitstream is decoded to obtain the audio signal to be rendered.

[0110] The received bitstream is decoded to obtain the audio signal to be rendered. The signal format of the audio signal to be rendered can include one signal format or a mixture of multiple signal formats, and the signal format can include a channel-based format, a scene-based format, or an object-based format, and the like.

[0111] Among the three different signal formats, the channel-based signal format is the most traditional audio signal format, which is easy to store and transmit, and can be directly played back by loudspeakers without much additional processing, i.e., the channel-based audio signal is for some standard loudspeaker arrangement, such as a 5.1 channel loudspeaker arrangement, a 7.1.4 channel loudspeaker arrangement, etc. One channel signal corresponds to one loudspeaker device. In actual applications, if the loudspeaker configuration format is different from the loudspeaker configuration format required by the audio signal to be rendered, upmixing or downmixing processing needs to be performed to adapt to the loudspeaker configuration format of the current application, and the downmixing processing will reduce the accuracy of the sound image in the playback sound field to some extent. For example, the channel-based signal format is in line with the 7.1.4 channel loudspeaker arrangement, but the loudspeaker configuration format of the current application is a 5.1 channel loudspeaker, and then the 7.1.4 channel signal needs to be downmixed to obtain a 5.1 channel signal, so as to be able to be played back by using a 5.1 channel loudspeaker. If it is necessary to use a headphone for playback, the loudspeaker signal can be further processed by a head-related transfer function (HRTF) / BRIR convolution to obtain a binaural rendering signal for binaural playback by a headphone or the like. The channel-based audio signal can be a mono audio signal, or can also be a multi-channel signal, for example, a stereo signal.

[0112] The object-based signal format is used to describe object audio, which includes a series of sound objects and corresponding metadata. The sound objects include independent sound sources, and the metadata includes static metadata such as language and starting time, and dynamic metadata such as the position, orientation, and sound pressure of the sound source. Therefore, the greatest advantage of the object-oriented signal format is that it can be used for any loudspeaker playback system for selective playback, while increasing the interactivity, such as adjusting the language, increasing the volume of some sound sources, and adjusting the position of the sound source object according to the movement of the listener, etc.

[0113] The scene-based signal format, which expands the actual physical sound signal or the sound signal collected by a microphone by using an orthogonal basis function, stores not the direct loudspeaker signal but the corresponding basis function expansion coefficient, and at the playback end, uses a corresponding sound field synthesis algorithm for binaural rendering playback. It can also be played back by using a variety of loudspeaker configurations, and the loudspeaker placement has great flexibility. The scene-based audio signal can include a first-order Ambisonics (FOA) signal, or a high-order Ambisonics (HOA) signal, etc.

[0114] The signal format is a signal format obtained by a collection end. For example, in a remote telephone conference application scenario in which multiple parties participate, some terminal devices send a stereo signal, that is, a sound channel-based audio signal, some terminal devices send an object-based audio signal of a remote participant, and some terminal devices send a high-order Ambisonics (HOA) signal, that is, a scene-based audio signal. A playback end decodes the received code stream to obtain a to-be-rendered audio signal, which is a mixed signal of the three signal formats. The audio signal rendering apparatus of the embodiment of the present application can support flexible rendering of an audio signal mixed with one or more signal formats.

[0115] The decoding of the received code stream can also obtain content description metadata. The content description metadata is used to indicate the signal format of the to-be-rendered audio signal. For example, in the above-mentioned remote telephone conference application scenario in which multiple parties participate, the playback end can obtain the content description metadata by decoding, and the content description metadata is used to indicate that the signal format of the to-be-rendered audio signal includes three signal formats of sound channel-based, object-based, and scene-based.

[0116] Step 402, obtaining control information, the control information being used to indicate at least one of content description metadata, rendering format flag information, loudspeaker configuration information, application scenario information, tracking information, posture information, or position information.

[0117] The content description metadata is used to indicate the signal format of the to-be-rendered audio signal, and the signal format includes at least one of sound channel-based, scene-based, or object-based.

[0118] The rendering format flag information is used to indicate an audio signal rendering format. The audio signal rendering format can include loudspeaker rendering or binaural rendering. In other words, the rendering format flag information is used to indicate that the audio rendering apparatus outputs a loudspeaker rendering signal or a binaural rendering signal. The rendering format flag information can be obtained from the decoding of the received code stream, or can be determined according to the hardware setting of the playback end, or is obtained according to the configuration information of the playback end.

[0119] The loudspeaker configuration information is used to indicate the layout of the loudspeakers. The layout of the loudspeakers can include the positions and quantities of the loudspeakers. The layout of the loudspeakers enables the audio rendering apparatus to generate a loudspeaker rendering signal of a corresponding layout. Figure 4 A layout of loudspeakers of an embodiment of the present application is shown in FIG. 8. As shown in FIG. 8, eight loudspeakers on a horizontal plane form a 7.1 layout configuration, in which a solid loudspeaker represents a subwoofer, and four loudspeakers on a plane above the horizontal plane form a 4.0 layout configuration. Figure 4 Figure 4 ​The four loudspeakers on the dotted line box together form a 7.1.4 loudspeaker layout. The loudspeaker configuration information can be determined according to the layout of the loudspeakers of the playback end, or can be obtained from the configuration information of the playback end.

[0120] The application scenario information is used to indicate renderer scene description information. The renderer scene description information can indicate a scene in which the rendered audio signal is output, i.e., a rendering sound field environment. The scene can be at least one of an indoor conference room, an indoor classroom, an outdoor grassland, or a concert venue. The application scenario information can be determined according to information obtained by a sensor of the playback end. For example, the environment data of the playback end is collected by one or more sensors such as an ambient light sensor, an infrared sensor, and the like, and the application scenario information is determined according to the environment data. For another example, the application scenario information can be determined according to an access point (AP) connected to the playback end. For example, the access point (AP) is a home wifi, and when the playback end is connected to the home wifi, the application scenario information can be determined as a home indoor. For another example, the application scenario information can be obtained from the configuration information of the playback end.

[0121] The tracking information is used to indicate whether the rendered audio signal changes with the head rotation of the listener. The tracking information can be obtained from the configuration information of the playback end. The pose information is used to indicate the orientation and amplitude of the head rotation. The pose information can be three degree of freedom (3DoF) data. The 3DoF data is used to represent the rotation information of the head of the listener. The 3DoF data can include three rotation angles of the head. The pose information can be 3DoF+ data, which represents the motion information of the upper body of the listener in the premise that the body is not moving on the seat. The 3DoF+ data can include three rotation angles of the head and the amplitude of the forward and backward motion and the amplitude of the left and right motion of the upper body. Alternatively, the 3DoF+ data can include three rotation angles of the head and the amplitude of the forward and backward motion of the upper body. Alternatively, the 3DoF+ data can include three rotation angles of the head and the amplitude of the left and right motion of the upper body. The position information is used to indicate the orientation and amplitude of the body movement of the listener. The pose information and the position information can be six degree of freedom (6DoF) data, which represents the information of the unconstrained free motion of the listener. The 6DoF data can include three rotation angles of the head and the amplitude of the forward and backward motion, the amplitude of the left and right motion, and the amplitude of the up and down motion of the body movement.

[0122] The control information can be obtained by the aforementioned audio signal rendering device generating control information based on at least one of the following: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, attitude information, or position information. Alternatively, control information can be obtained by receiving control information from other devices; the specific implementation of this method is not limited in the embodiments of this application.

[0123] For example, before rendering the audio signal to be rendered, embodiments of this application can generate control information based on at least one of the following: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, posture information, or position information. (Refer to...) Figure 5 As shown, the input information includes at least one of the following: metadata describing the above content, rendering format flag information, speaker configuration information, application scenario information, tracking information, posture information, or position information. The input information is analyzed to generate control information. This control information can be applied to the rendering process, enabling adaptive selection of the rendering processing method to improve the rendering effect of the audio signal. This control information may include the rendering format of the output signal (i.e., the rendered audio signal), application scenario information, the rendering processing method used, and the database used for rendering, etc.

[0124] Step 403: Render the audio signal to be rendered according to the control information to obtain the rendered audio signal.

[0125] Since the control information is generated based on at least one of the above-mentioned metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, posture information, or position information, the corresponding rendering method is used based on the control information to achieve adaptive selection of the rendering method based on the input information, thereby improving the audio rendering effect.

[0126] In some embodiments, the step 403 can include at least one of: rendering pre-processing of the to-be-rendered audio signals according to the control information; or format conversion of the to-be-rendered audio signals according to the control information; or local reverberation processing of the to-be-rendered audio signals according to the control information; or grouped source Transformations of the to-be-rendered audio signals according to the control information; or Dynamic Range Compression of the to-be-rendered audio signals according to the control information; or Binaural rendering of the to-be-rendered audio signals according to the control information; or Loudspeaker rendering of the to-be-rendered audio signals according to the control information.

[0127] The rendering pre-processing is used to perform static initialization processing on the to-be-rendered audio signals by using the relevant information of the sending end, which can include the reverberation information of the sending end. The rendering pre-processing can provide a basis for one or more dynamic rendering processing methods such as signal format conversion, local reverberation processing, grouped source Transformations, Dynamic Range Compression, Binaural rendering, or Loudspeaker rendering, so that the rendered audio signals are matched with at least one of the playback device or the playback environment, thereby providing better auditory effect. The specific implementation of the rendering pre-processing can refer to the explanation of the embodiment shown in 6A.

[0128] The grouped source Transformations is used to perform real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the audio signals of each signal format in the to-be-rendered audio signals, i.e., the same processing is performed on the audio signals of the same signal format, so as to reduce the processing complexity. The specific implementation of the grouped source Transformations can refer to the explanation of the embodiment shown in 11A.

[0129] The Dynamic Range Compression is used to compress the dynamic range of the to-be-rendered audio signals, so as to improve the playback quality of the rendered audio signals. The dynamic range is the intensity difference between the strongest signal and the weakest signal in the to-be-rendered audio signals, and the unit is “db”. The specific implementation of the Dynamic Range Compression can refer to the explanation of the embodiment shown in 12A.

[0130] The Binaural rendering is used to convert the to-be-rendered audio signals into binaural signals, so as to be played back through earphones. The specific implementation of the Binaural rendering can refer to the explanation of the step 504 of the embodiment shown in 6A.

[0131] The loudspeaker rendering is used to convert the audio signal to be rendered into a signal matching the loudspeaker layout, so as to be played back through the loudspeakers. The implementation of the loudspeaker rendering can refer to the explanation of step 504 of the embodiment shown in FIG. 6A.

[0132] For example, the content description metadata, the rendering format flag information and the tracking information are indicated in the control information. The implementation of rendering the audio signal to be rendered according to the control information is explained. An example: the content description metadata indicates that the input signal format is a scene-based audio signal, the rendering format flag information indicates that the rendering is binaural rendering, and the tracking information indicates that the rendered audio signal does not change with the head rotation of the listener. Then, the audio signal to be rendered is rendered according to the control information, which can be: converting the scene-based audio signal into a channel-based audio signal, directly convolving the channel-based audio signal with HRTF / BRIR to generate a binaural rendering signal, and the binaural rendering signal is the rendered audio signal. Another example: the content description metadata indicates that the input signal format is a scene-based audio signal, the rendering format flag information indicates that the rendering is binaural rendering, and the tracking information indicates that the rendered audio signal changes with the head rotation of the listener. Then, the audio signal to be rendered is rendered according to the control information, which can be: performing spherical harmonic decomposition on the scene-based audio signal to generate a virtual loudspeaker signal, convolving the virtual loudspeaker signal with HRTF / BRIR to generate a binaural rendering signal, and the binaural rendering signal is the rendered audio signal. Another example: the content description metadata indicates that the input signal format is a channel-based audio signal, the rendering format flag information indicates that the rendering is binaural rendering, and the tracking information indicates that the rendered audio signal does not change with the head rotation of the listener. Then, the audio signal to be rendered is rendered according to the control information, which can be: directly convolving the channel-based audio signal with HRTF / BRIR to generate a binaural rendering signal, and the binaural rendering signal is the rendered audio signal. Another example: the content description metadata indicates that the input signal format is a channel-based audio signal, the rendering format flag information indicates that the rendering is binaural rendering, and the tracking information indicates that the rendered audio signal changes with the head rotation of the listener. Then, the audio signal to be rendered is rendered according to the control information, which can be: converting the channel-based audio signal into a scene-based audio signal, performing spherical harmonic decomposition on the scene-based audio signal to generate a virtual loudspeaker signal, and convolving the virtual loudspeaker signal with HRTF / BRIR to generate a binaural rendering signal, and the binaural rendering signal is the rendered audio signal. It should be noted that the above examples are only exemplary, and are not limited to the above examples in actual application. Therefore, the information indicated by the control information is used to adaptively select a suitable processing mode to render the input signal, so as to improve the rendering effect.

[0133] For example, taking a control message that indicates content description metadata, rendering format flags, application scenario information, tracking information, posture information, and position information, the specific implementation of rendering the audio signal to be rendered based on the control message can be as follows: Local reverb processing, grouping processing, and binaural rendering or speaker rendering can be performed on the audio signal to be rendered based on the content description metadata, rendering format flags, application scenario information, tracking information, posture information, and position information; or, signal format conversion, local reverb processing, grouping processing, and binaural rendering or speaker rendering can be performed on the audio signal to be rendered based on the content description metadata, rendering format flags, application scenario information, tracking information, posture information, and position information. Thus, by using the information indicated by the control message, an appropriate processing method is adaptively selected to render the input signal, thereby improving the rendering effect. It should be noted that the above examples are merely illustrative and are not intended to limit the application to only these examples.

[0134] In this embodiment, the audio signal to be rendered is obtained by decoding the received bitstream, and control information is obtained. This control information is used to indicate at least one of the following: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, posture information, or position information. The audio signal to be rendered is rendered according to the control information to obtain the rendered audio signal. This can realize adaptive selection of the rendering method based on at least one of the following input information: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, posture information, or position information, thereby improving the audio rendering effect.

[0135] Figure 6A This is a flowchart of another audio signal rendering method according to an embodiment of this application. Figure 6B This is a schematic diagram of a pre-rendering process according to an embodiment of this application. The execution entity of this embodiment can be the aforementioned audio signal rendering device. This embodiment is the aforementioned... Figure 3 One possible implementation of the illustrated embodiment is to specifically explain the rendering pre-processing of the audio signal rendering method of this application. Rendering pre-processing includes: setting the precision of rotation and translation for channel-based, object-based, or scene-based audio signals and performing three-degree-of-freedom (3DoF) processing, as well as reverberation processing, such as... Figure 6A As shown, the method in this embodiment may include:

[0136] Step 501: Obtain the audio signal to be rendered and the first reverberation information by decoding the received bitstream.

[0137] The audio signal to be rendered includes at least one of channel-based audio signal, object-based audio signal, or scene-based audio signal, and the first reverberation information includes at least one of the following: first reverberation output loudness information, time difference information between the first direct sound and the early reflected sound, first reverberation duration information, first room shape and size information, or first sound scattering information.

[0138] Step 502: Obtain control information, which is used to indicate at least one of the following: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, attitude information, or position information.

[0139] For an explanation of step 502, please refer to [link / reference needed]. Figure 3 The specific explanation of step 402 in the illustrated embodiment will not be repeated here.

[0140] Step 503: According to the control information, perform control processing on the audio signal to be rendered, obtain the audio signal after control processing, and perform reverberation processing on the audio signal after control processing according to the first reverberation information to obtain the first audio signal.

[0141] The aforementioned control processing includes at least one of the following: performing initial 3DoF processing on channel-based audio signals in the audio signal to be rendered, performing transformation processing on object-based audio signals in the audio signal to be rendered, or performing initial 3DoF processing on scene-based audio signals in the audio signal to be rendered.

[0142] This application embodiment can perform pre-processing on individual sound sources according to control information. An individual sound source can be a channel-based audio signal, an object-based audio signal, or a scene-based audio signal. Taking a pulse code modulation (PCM) signal 1 as an example, see [link to example]. Figure 6B As shown, the input signal for pre-rendering processing is PCM signal 1, and the output signal is PCM signal 2. If the control information indicates that the input signal format includes channel-based processing, then pre-rendering processing includes initial 3DoF processing and reverb processing of the channel-based audio signal. If the control information indicates that the input signal format includes object-based processing, then pre-rendering processing includes transformation and reverb processing of the object-based audio signal. If the control information indicates that the input signal format includes scene-based processing, then pre-rendering processing includes initial 3DoF processing and reverb processing of the scene-based audio signal. The output PCM signal 2 is obtained after pre-rendering processing.

[0143] For example, when the audio signal to be rendered includes a channel-based audio signal and a scene-based audio signal, the channel-based audio signal and the scene-based audio signal can be respectively pre-processed according to the control information. That is, the channel-based audio signal is initially processed by 3DoF according to the control information, and the channel-based audio signal is reverberation-processed according to the first reverberation information to obtain a pre-processed channel-based audio signal. The scene-based audio signal is initially processed by 3DoF according to the control information, and the scene-based audio signal is reverberation-processed according to the first reverberation information to obtain a pre-processed scene-based audio signal. The first audio signal includes the pre-processed channel-based audio signal and the pre-processed scene-based audio signal. When the audio signal to be rendered includes a channel-based audio signal, an object-based audio signal, and a scene-based audio signal, the processing process is similar to the foregoing example. The first audio signal obtained by pre-processing can include a pre-processed channel-based audio signal, a pre-processed object-based audio signal, and a pre-processed scene-based audio signal. The foregoing two examples are illustrative. When the audio signal to be rendered includes other single signal format audio signals or a combination of multiple signal format audio signals, the specific implementation is similar, that is, the precision of rotation and translation of the single signal format audio signal is set and the initial 3DoF processing and the reverberation processing are completed. Here, the examples are not repeated.

[0144] The pre-rendering processing of the embodiments of the present application can select a corresponding processing method to perform pre-rendering processing on individual sources according to the control information. For a scene-based audio signal, the initial 3DoF processing can include moving and rotating the scene-based audio signal according to the starting position (determined based on the initial 3DoF data), and then performing virtual loudspeaker mapping on the processed scene-based audio signal to obtain a virtual loudspeaker signal corresponding to the scene-based audio signal. For a channel-based audio signal, the channel-based audio signal includes one or more channel signals, and the initial 3DoF processing can include calculating the initial position of the listener (determined based on the initial 3DoF data) and the relative position of each channel signal to select initial HRTF / BRIR data to obtain a corresponding channel signal and an initial HRTF / BRIR data index. For an object-based audio signal, the object-based audio signal includes one or more object signals, and the transformation processing can include calculating the initial position of the listener (determined based on the initial 3DoF data) and the relative position of each object signal to select initial HRTF / BRIR data to obtain a corresponding object signal and an initial HRTF / BRIR data index.

[0145] The above reverberation processing generates first reverberation information according to the output parameters of the decoder. The parameters required by the reverberation processing include but are not limited to one or more of the output loudness information of the reverberation, the time difference information of the direct sound and the early reflection sound, the duration information of the reverberation, the room shape and size information, or the scattering degree information of the sound, etc. The audio signals of the three signal formats are processed by the first reverberation information generated according to the three signal formats, to obtain output signals with the reverberation information of the sending end, i.e., the first audio signals.

[0146] Step 504, binaural rendering or loudspeaker rendering is performed on the first audio signal to obtain the rendered audio signal.

[0147] The rendered audio signal can be played through a loudspeaker or through a headphone.

[0148] In an implementation, the first audio signal can be speaker rendered according to the control information. For example, the input signal (i.e., the first audio signal herein) can be processed according to the speaker configuration information in the control information and the rendering format flag information in the control information. In this case, one part of the first audio signal can be processed by one speaker rendering manner, and another part of the first audio signal can be processed by another speaker rendering manner. The speaker rendering manner can include speaker rendering of a channel-based audio signal, speaker rendering of a scene-based audio signal, or speaker rendering of an object-based audio signal. The speaker processing of the channel-based audio signal can include upmixing or downmixing processing of the input channel-based audio signal to obtain a speaker signal corresponding to the channel-based audio signal. The speaker rendering of the object-based audio signal can include applying an amplitude panning processing method to the object-based audio signal to obtain a speaker signal corresponding to the object-based audio signal. The speaker rendering of the scene-based audio signal can include decoding processing of the scene-based audio signal to obtain a speaker signal corresponding to the scene-based audio signal. One or more of the speaker signal corresponding to the channel-based audio signal, the speaker signal corresponding to the object-based audio signal, and the speaker signal corresponding to the scene-based audio signal are fused to obtain a speaker signal. In some embodiments, the speaker signal can be further subjected to de-chorus processing and virtual height information through horizontal position speakers without height speakers.

[0149] Taking the first audio signal as the PCM signal 6 as an example, Figure 7 A schematic diagram of speaker rendering provided by an embodiment of the present application is shown in FIG. 6. Figure 7 As shown in FIG. 6, the input of the speaker rendering is the PCM signal 6, and the output is a speaker signal after the speaker rendering.

[0150] In another implementation, the first audio signal can be binaurally rendered according to the control information. For example, the input signal (i.e., the first audio signal herein) can be processed according to the rendering format flag information in the control information. In this case, the HRTF data corresponding to an initial HRTF data index obtained through pre-processing can be acquired from an HRTF database according to the HRTF data index. The head-centered HRTF data can be converted into binaural-centered HRTF data, and the HRTF data can be subjected to de-chorus processing, headphone equalization processing, and individualization processing. The input signal (i.e., the first audio signal herein) can be processed according to the HRTF data to obtain a binaural signal. The binaural signal processing can include processing of a channel-based audio signal and an object-based audio signal through a direct convolution method to obtain a binaural signal, and processing of a scene-based audio signal through a spherical harmonic decomposition convolution method to obtain a binaural signal.

[0151] Taking the first audio signal as the PCM signal 6 as an example, Figure 8 A schematic diagram of binaural rendering provided by the embodiment of the present application is shown in FIG. 6. Figure 8 As shown in FIG. 6, the input of binaural rendering is the PCM signal 6, and the output is a binaural signal after binaural rendering.

[0152] In the embodiment, the to-be-rendered audio signal and the first reverberation information are acquired by decoding the received code stream, at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information indicated by the control information is used to control process the to-be-rendered audio signal, the control-processed audio signal is acquired, the control processing includes at least one of initial 3DoF processing on the channel-based audio signal, transformation processing on the object-based audio signal or initial 3DoF processing on the scene-based audio signal, and the control-processed audio signal is processed with reverberation according to the first reverberation information, to acquire the first audio signal, the first audio signal is binaurally rendered or loudspeaker rendered to acquire the rendered audio signal, and the adaptive selection of the rendering mode based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information can be realized, so that the audio rendering effect is improved.

[0153] Figure 9A A flowchart of another audio signal rendering method of the embodiment of the present application is shown in FIG. 7. Figure 9B A schematic diagram of signal format conversion of the embodiment of the present application is shown in FIG. 8, and the execution subject of the embodiment of the present application can be the audio signal rendering device described above. The embodiment shown in FIG. 8 is a specific implementation of the audio signal rendering method of the embodiment of the present application. Figure 3 The signal format conversion (Format converter) of the audio signal rendering method of the embodiment of the present application is specifically explained and described. The signal format conversion (Format converter) can convert one signal format into another signal format to improve the rendering effect. As shown in FIG. 9, the method of the embodiment can include: Figure 9A

[0154] Step 601: acquiring a to-be-rendered audio signal by decoding a received code stream.

[0155] The explanation and description of step 601 can be referred to the specific explanation and description of step 401 of the embodiment shown in FIG. 4, which will not be repeated here. Figure 3

[0156] ​​In step 602, control information is acquired, the control information being used to indicate at least one of content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, pose information, or position information.

[0157] The explanation of step 602 can be referred to the explanation of step 602 in the embodiment shown in Figure 3 The specific explanation of step 402 in the embodiment shown in FIG. 4 will not be repeated here.

[0158] In step 603, signal format conversion is performed on the to-be-rendered audio signal according to the control information, to obtain a sixth audio signal.

[0159] The signal format conversion includes at least one of the following: converting a channel-based audio signal in the to-be-rendered audio signal into a scene-based or object-based audio signal; or converting a scene-based audio signal in the to-be-rendered audio signal into a channel-based or object-based audio signal; or converting an object-based audio signal in the to-be-rendered audio signal into a channel-based or scene-based audio signal.

[0160] Taking the to-be-rendered audio signal as the PCM signal 2 for example, as shown in Figure 9B The control information can be used to select corresponding signal format conversion, to convert the PCM signal 2 in one signal format into the PCM signal 3 in another signal format.

[0161] The embodiments of the present application can adaptively select signal format conversion according to the control information, and can implement conversion of a part of input signals (here, the to-be-rendered audio signal) by using one signal format conversion (for example, any of the above) and conversion of another part of input signals by using another signal format conversion.

[0162] For example, in the application scenario of binaural rendering, sometimes a part of input signals needs to be rendered by using direct convolution, and another part of input signals needs to be rendered by using HOA, and therefore the scene-based audio signal can be converted into the channel-based audio signal by signal format conversion, so as to perform direct convolution in the subsequent binaural rendering process, and the object-based audio signal can be converted into the scene-based audio signal, so as to be rendered by using the HOA mode. For another example, the pose information and the position information in the control information indicate that the listener needs to perform 6DoF rendering processing, and therefore the channel-based audio signal can be converted into the object-based audio signal by signal format conversion, and the scene-based audio signal can be converted into the object-based audio signal.

[0163] In the signal format conversion of the to-be-rendered audio signal, the processing performance of the terminal device can also be combined. The processing performance of the terminal device can be the processor performance of the terminal device, for example, the frequency, bit number, etc. of the processor. One implementation manner of the signal format conversion of the to-be-rendered audio signal according to the control information can include: performing the signal format conversion of the to-be-rendered audio signal according to the control information, the signal format of the to-be-rendered audio signal, and the processing performance of the terminal device. For example, the pose information and the position information in the control information indicate that the listener is to perform 6DoF rendering processing, and the processing performance of the terminal device is combined to determine whether to convert. For example, the processing performance of the terminal device is poor, and then the object-based audio signal or the channel-based audio signal can be converted into the scene-based audio signal. The processing performance of the terminal device is good, and then the scene-based audio signal or the channel-based audio signal can be converted into the object-based audio signal.

[0164] One implementation manner is to determine whether to convert and the signal format after conversion according to the pose information and the position information in the control information and the signal format of the to-be-rendered audio signal.

[0165] In the conversion of the scene-based audio signal into the object-based audio signal, the scene-based audio signal can be first converted into a virtual loudspeaker signal, and then each virtual loudspeaker signal and its corresponding position is an object-based audio signal, wherein the virtual loudspeaker signal is audio content, and the corresponding position is information in metadata.

[0166] Step 604: performing binaural rendering or loudspeaker rendering on the sixth audio signal to obtain a rendered audio signal.

[0167] The explanation of step 604 can be referred to the specific explanation of step 504 in Figure 6A , which will not be described here. That is, the first audio signal in step 504 in Figure 6A is replaced by the sixth audio signal.

[0168] In the embodiment, the to-be-rendered audio signal is acquired by decoding the received code stream, signal format conversion is performed on the to-be-rendered audio signal according to at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information indicated by the control information, the sixth audio signal is acquired, and binaural rendering or loudspeaker rendering is performed on the sixth audio signal to acquire the rendered audio signal, so that adaptive selection of a rendering mode based on at least one of the input information of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information is implemented, and the audio rendering effect is improved. The signal format conversion is performed on the to-be-rendered audio signal according to the control information, so that flexible conversion of the signal format is implemented, and the audio signal rendering method of the embodiment is applicable to any signal format. The audio signal of a suitable signal format is rendered, and the audio rendering effect is improved.

[0169] Figure 10A FIG. 6 is a flowchart of another audio signal rendering method according to an embodiment of the present application. Figure 10B FIG. 7 is a schematic diagram of local reverberation processing according to an embodiment of the present application. The execution subject of the embodiment can be the audio signal rendering apparatus, and the embodiment is a specific implementation of the local reverberation processing of the audio signal rendering method according to the embodiment shown in FIG. 5. Figure 3 The local reverberation processing can implement rendering based on the reverberation information of the playback end to improve the rendering effect, so that the audio signal rendering method can support the AR application scenario, and the like. Figure 10A The method can include the following steps.

[0170] In step 701, the to-be-rendered audio signal is acquired by decoding the received code stream.

[0171] The explanation of step 701 can be referred to the specific explanation of step 401 of the embodiment shown in FIG. 4, which will not be repeated here. Figure 3 The explanation of step 701 can be referred to the specific explanation of step 401 of the embodiment shown in FIG. 4, which will not be repeated here.

[0172] In step 702, the control information is acquired, and the control information is used to indicate at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information.

[0173] The explanation of step 702 can be referred to the specific explanation of step 402 of the embodiment shown in FIG. 4, which will not be repeated here. Figure 3The specific explanation of step 402 in the illustrated embodiment will not be repeated here.

[0174] Step 703: Obtain the second reverberation information, which is the reverberation information of the scene where the rendered audio signal is located. The second reverberation information includes at least one of the following: second reverberation output loudness information, time difference information between the second direct sound and the early reflected sound, second reverberation duration information, second room shape and size information, or second sound scattering information.

[0175] The second reverberation information is the reverberation information generated on the audio signal rendering device side. This second reverberation information can also be called local reverberation information.

[0176] In some embodiments, the second reverberation information can be generated based on the application scenario information of the audio signal rendering device. The application scenario information can be obtained through configuration information set by the listener or through sensors. This application scenario information may include location or environmental information, etc.

[0177] Step 704: Perform local reverb processing on the audio signal to be rendered based on the control information and the second reverb information to obtain the seventh audio signal.

[0178] The seventh audio signal is obtained by rendering based on the control information and the second reverberation information.

[0179] One possible approach is to cluster signals of different signal formats in the audio signal to be rendered based on control information to obtain at least one of channel-based group signals, scene-based group signals, or object-based group signals. Based on second reverberation information, local reverberation processing is then performed on at least one of the channel-based group signals, scene-based group signals, or object-based group signals to obtain a seventh audio signal.

[0180] In other words, the audio signal rendering device can generate reverberation information for three audio signal formats, enabling the audio signal rendering method of this embodiment to be applied to augmented reality scenarios to enhance the sense of presence. Because the environmental information of the real-time location of the playback end cannot be predicted in augmented reality scenarios, the reverberation information cannot be determined at the production end. This embodiment generates corresponding second reverberation information based on the real-time input application scenario information for rendering processing, thereby improving the rendering effect.

[0181] For example, such as Figure 10B As shown, for Figure 10B The PCM signal 3 shown is clustered into three types of signals: channel-based group signal, object-based group signal, and scene-based group signal. These three group signals are then subjected to reverberation processing to output the seventh audio signal, as shown below. Figure 10BThe PCM signal 4 is shown.

[0182] Step 705, binaural rendering or loudspeaker rendering is performed on the seventh audio signal to obtain a rendered audio signal.

[0183] The explanation of step 705 can refer to the specific explanation of step 504 in Figure 6A , which will not be repeated here. That is, the first audio signal in step 504 in Figure 6A is replaced by the seventh audio signal.

[0184] In this embodiment, the to-be-rendered audio signal is obtained by decoding the received code stream, and at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information, or the position information indicated by the control information, and the second reverberation information are used to perform local reverberation processing on the to-be-rendered audio signal to obtain the seventh audio signal. The binaural rendering or loudspeaker rendering is performed on the seventh audio signal to obtain a rendered audio signal. The adaptive selection of the rendering mode based on the at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information, or the position information input information can be realized, thereby improving the audio rendering effect. The second reverberation information corresponding to the real-time input application scenario information is used for rendering processing, which can improve the audio rendering effect and provide real-time reverberation corresponding to the scene for the AR application scenario.

[0185] Figure 11A FIG. 8 is a flowchart of another audio signal rendering method according to an embodiment of the present application, Figure 11B FIG. 9 is a schematic diagram of a grouped source transformation according to an embodiment of the present application. The execution subject of the present embodiment can be the audio signal rendering device described above. The present embodiment is a specific implementation of the audio signal rendering method according to the embodiment of the present application shown in Figure 3 . The grouped source transformation can reduce the complexity of the rendering processing. For example, as shown in Figure 11A , the method of the present embodiment can include:

[0186] Step 801, obtaining a to-be-rendered audio signal by decoding a received code stream.

[0187] The explanation of step 801 can refer to the specific explanation of step 401 in the embodiment shown in Figure 3 , which will not be repeated here.

[0188] Step 802, obtaining control information, the control information is used for indicating at least one of content description metadata, rendering format flag information, speaker configuration information, application scene information, tracking information, attitude information or position information.

[0189] The specific interpretation of step 802 in the embodiment shown can be referred to Figure 3 The specific interpretation of step 402 in the embodiment shown is not repeated here.

[0190] Step 803, performing real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the audio signal of each signal format in the to-be-rendered audio signal according to the control information, to obtain an eighth audio signal.

[0191] The embodiment can process the audio signals of the three signal formats according to the 3DoF, 3DoF+ and 6DoF information in the control information, that is, uniformly process the audio signals of each format, which can reduce the processing complexity on the basis of ensuring the processing performance.

[0192] The real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the channel-based audio signal is to calculate the relative orientation relationship between the listener and the channel-based audio signal in real time. The real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the object-based audio signal is to calculate the relative orientation and relative distance relationship between the listener and the object sound source signal in real time. The real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the scene-based audio signal is to calculate the position relationship between the listener and the center of the scene signal in real time.

[0193] In one implementation, the real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the channel-based audio signal is to obtain a processed HRTF / BRIR data index according to an initial HRTF / BRIR data index and 3DoF / 3DoF+ / 6DoF data of the listener at the current time. The processed HRTF / BRIR data index is used to reflect the orientation relationship between the listener and the channel signal.

[0194] In one implementation, the real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the object-based audio signal is to obtain a processed HRTF / BRIR data index according to an initial HRTF / BRIR data index and 3DoF / 3DoF+ / 6DoF data of the listener at the current time. The processed HRTF / BRIR data index is used to reflect the relative orientation and relative distance relationship between the listener and the object signal.

[0195] In an implementation, the real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the scene-based audio signal is to obtain processed HRTF / BRIR data index according to the virtual loudspeaker signal and the 3DoF / 3DoF+ / 6DoF data of the current time of the listener. The processed HRTF / BRIR data index is used to reflect the positional relationship between the listener and the virtual loudspeaker signal.

[0196] For example, referring to Figure 11B , different format types of signals in the PCM signal 4 as shown in Figure 11B are respectively subjected to real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing, and the PCM signal 5, i.e., the eighth audio signal, is output. The PCM signal 5 includes the PCM signal 4 and the processed HRTF / BRIR data index.

[0197] Step 804: binaural rendering or loudspeaker rendering is performed on the eighth audio signal to obtain a rendered audio signal.

[0198] The explanation of step 804 can refer to the specific explanation of step 504 in Figure 6A , which will not be repeated here. That is, the first audio signal in step 504 in Figure 6A is replaced by the eighth audio signal.

[0199] In the embodiment, the to-be-rendered audio signal is obtained by decoding the received code stream, real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing is performed on the audio signal of each signal format in the to-be-rendered audio signal according to at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the attitude information, or the position information indicated by the control information, the eighth audio signal is obtained, binaural rendering or loudspeaker rendering is performed on the eighth audio signal to obtain a rendered audio signal, and adaptive selection of a rendering mode based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the attitude information, or the position information input information is implemented, thereby improving the audio rendering effect. Uniform processing is performed on each format of audio signal, which can reduce the processing complexity on the basis of ensuring the processing performance.

[0200] Figure 12A FIG. 8 is a flowchart of another audio signal rendering method according to an embodiment of the present application, Figure 12B FIG. 9 is a schematic diagram of dynamic range compression according to an embodiment of the present application. The execution subject of the present embodiment can be the audio signal rendering device described above, and the present embodiment is the audio signal rendering device described above. Figure 3An implementation of the embodiment shown, i.e., the dynamic range compression of the audio signal rendering method of the embodiment of the present application, is specifically explained and described. As shown in Figure 12A The method of the embodiment shown can include, as shown in

[0201] Step 901: obtaining the audio signal to be rendered by decoding the received code stream.

[0202] The explanation and description of step 901 can be referred to the explanation and description of step 401 of the embodiment shown. Figure 3 The specific explanation and description of step 401 of the embodiment shown is not repeated here.

[0203] Step 902: obtaining control information, which is used to indicate at least one of content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, pose information or position information.

[0204] The explanation and description of step 902 can be referred to the explanation and description of step 402 of the embodiment shown. Figure 3 The specific explanation and description of step 402 of the embodiment shown is not repeated here.

[0205] Step 903: performing dynamic range compression on the audio signal to be rendered according to the control information to obtain a ninth audio signal.

[0206] The input signal (for example, the audio signal to be rendered here) can be dynamically range compressed according to the control information to output the ninth audio signal.

[0207] An implementation, based on the application scenario information and the rendering format flag in the control information, performs dynamic range compression on the audio signal to be rendered. For example, the home theater scenario and the headphone rendering scenario have different requirements for the amplitude of the frequency response. For another example, different channel program contents require similar sound loudness, and the same program content also needs to ensure appropriate dynamic range. For another example, a stage play needs to ensure that the dialogue content can be clearly heard when the light sound is spoken and the sound loudness is within a certain range when the music is played loudly, so that the overall effect will not have the feeling of high and low. For the above examples, the audio signal to be rendered can be dynamically range compressed according to the control information to ensure the quality of audio rendering.

[0208] For example, as shown in Figure 12B The PCM signal 5 shown in Figure 12B is dynamically range compressed to output a PCM signal 6, i.e., the ninth audio signal.

[0209] Step 904: performing binaural rendering or speaker rendering on the ninth audio signal to obtain a rendered audio signal.

[0210] The explanation of step 904 can be seen in the specific explanation of step 504 in Figure 6A , which will not be repeated here. The first audio signal in step 504 in Figure 6A is replaced by the ninth audio signal.

[0211] In this embodiment, the to-be-rendered audio signal is obtained by decoding the received code stream, at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information indicated by the control information is used for dynamic range compression of the to-be-rendered audio signal, and the ninth audio signal is obtained. The binaural rendering or loudspeaker rendering is performed on the ninth audio signal to obtain the rendered audio signal. The adaptive selection of the rendering mode based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the posture information or the position information can be realized, so as to improve the audio rendering effect.

[0212] The above uses Figures 6A to 12B , respectively, to explain the rendering pre-processing of the to-be-rendered audio signal according to the control information, the signal format conversion of the to-be-rendered audio signal according to the control information, the local reverberation processing of the to-be-rendered audio signal according to the control information, the grouped source transformation of the to-be-rendered audio signal according to the control information, the dynamic range compression of the to-be-rendered audio signal according to the control information, the binaural rendering of the to-be-rendered audio signal according to the control information, and the loudspeaker rendering of the to-be-rendered audio signal according to the control information. The control information can make the audio signal rendering device adaptively select the rendering processing mode, and improve the rendering effect of the audio signal.

[0213] In some embodiments, the above-mentioned various embodiments can also be implemented in combination, i.e., based on the control information, one or more of rendering pre-processing, signal format conversion, local reverberation processing, grouped source transformations, or dynamic range compression are selected to process the audio signal to be rendered to improve the rendering effect of the audio signal.

[0214] The following embodiment illustrates the audio signal rendering method of the present application by taking rendering pre-processing, signal format conversion, local reverberation processing, grouped source transformations, and dynamic range compression based on control information as an example.

[0215] Figure 13A The following is a schematic diagram of an audio signal rendering device according to an embodiment of the present application, Figure 13B The following is a schematic diagram of a detailed architecture of an audio signal rendering device according to an embodiment of the present application, as shown in Figure 13A The audio signal rendering device according to an embodiment of the present application can include a rendering interpreter, a rendering pre-processor, a signal format adaptive converter, a mixer, a grouped source processor, a dynamic range compressor, a loudspeaker rendering processor, and a binaural rendering processor. The audio signal rendering device according to an embodiment of the present application has flexible and universal rendering processing functions. The output of the decoder is not limited to a single signal format, such as a 5.1 multichannel format or a HOA signal of a certain order, but can also be a mixed form of three signal formats. For example, in a remote telephone conference application scenario involving multiple parties, some terminals send stereo channel signals, some terminals send object signals of a remote participant, and some terminals send high-order HOA signals. The audio signal decoded by the decoder from the code stream is a mixed signal of multiple signal formats. The audio rendering device according to an embodiment of the present application can support flexible rendering of mixed signals.

[0216] The rendering interpreter is configured to generate control information based on at least one of the content description metadata, the rendering format flag information, the loudspeaker configuration information, the application scenario information, the tracking information, the pose information, or the position information. The rendering pre-processor is configured to perform rendering pre-processing on the input audio signals as described in the above embodiments. The signal format adaptive converter is configured to perform signal format conversion on the input audio signals. The mixer is configured to perform local reverberation processing on the input audio signals. The group processor is configured to perform grouped source transformations on the input audio signals. The dynamic range compressor is configured to perform dynamic range compression on the input audio signals. The loudspeaker renderer is configured to perform loudspeaker rendering on the input audio signals. The binaural renderer is configured to perform binaural rendering on the input audio signals.

[0217] The detailed block diagram of the audio signal rendering apparatus is shown in FIG. 6. Figure 13B The rendering pre-processor can perform rendering pre-processing on audio signals of different signal formats respectively. The specific implementation of the rendering pre-processing can be referred to the embodiments shown in FIG. 3. Figure 6A The audio signals of different signal formats output by the rendering pre-processor are input into the signal format adaptive converter. The signal format adaptive converter performs format conversion or no conversion on the audio signals of different signal formats. For example, converts the channel-based audio signals into object-based audio signals (e.g., C to O as shown in FIG. 4), converts the channel-based audio signals into scene-based audio signals (e.g., C to HOA as shown in FIG. 5), converts the object-based audio signals into channel-based audio signals (e.g., O to C as shown in FIG. 6), converts the object-based audio signals into scene-based audio signals (e.g., O to HOA as shown in FIG. 7), converts the scene-based audio signals into channel-based audio signals (e.g., HOA to C as shown in FIG. 8), or converts the scene-based audio signals into scene-based audio signals (e.g., HOA to O as shown in FIG. 9). Figure 13B Figure 13B Figure 13B Figure 13B Figure 13B Figure 13B The audio signals output by the signal format adaptive converter are input into the mixer.

[0218] ​​​​​The mixer clusters audio signals of different formats to obtain group signals of different formats. A local reverb unit performs reverb processing on these group signals and inputs the processed audio signals to the group processor. The group processor performs real-time 3DoF processing, or 3DoF+ processing, or 6DoF processing on the group signals of different formats. The audio signal output from the group processor is input to a dynamic range compressor, which performs dynamic range compression on the audio signal output from the group processor and outputs the compressed audio signal to a speaker rendering processor or a binaural rendering processor. The binaural rendering processor performs direct convolution processing on the channel-based and object-based audio signals in the input audio signal, and performs spherical harmonic decomposition convolution on the scene-based audio signal in the input audio signal, outputting a binaural signal. The speaker rendering processor performs channel upmixing or downmixing on the channel-based audio signals in the input audio signal, performs energy mapping on the object-based audio signals in the input audio signal, and performs scene signal mapping on the scene-based audio signals in the input audio signal, outputting a speaker signal.

[0219] Based on the same inventive concept as the above method, this application also provides an audio signal rendering apparatus.

[0220] Figure 14 This is a schematic diagram of the structure of an audio signal rendering device according to an embodiment of this application, as shown below. Figure 14 As shown, the audio signal rendering device 1500 includes: an acquisition module 1501, a control information generation module 1502, and a rendering module 1503.

[0221] The acquisition module 1501 is used to acquire the audio signal to be rendered by decoding the received bitstream.

[0222] The control information generation module 1502 is used to acquire control information, which is used to indicate at least one of the following: content description metadata, rendering format flag information, speaker configuration information, application scenario information, tracking information, attitude information, or position information.

[0223] The rendering module 1503 is used to render the audio signal to be rendered according to the control information, so as to obtain the rendered audio signal.

[0224] The content description metadata is used to indicate a signal format of the audio signal to be rendered, the signal format comprising at least one of channel-based, scene-based or object-based; the rendering format flag information is used to indicate an audio signal rendering format, the audio signal rendering format comprising loudspeaker rendering or binaural rendering; the loudspeaker configuration information is used to indicate a layout of loudspeakers; the application scenario information is used to indicate renderer scene description information; the tracking information is used to indicate whether the rendered audio signal changes with a head rotation of a listener; the pose information is used to indicate an orientation and an amplitude of the head rotation; and the position information is used to indicate an orientation and an amplitude of a body movement of the listener.

[0225] In some embodiments, the rendering module 1503 is configured to perform at least one of:

[0226] performing pre-processing on the audio signal to be rendered according to the control information; or

[0227] performing signal format conversion on the audio signal to be rendered according to the control information; or

[0228] performing local reverberation processing on the audio signal to be rendered according to the control information; or

[0229] performing group processing on the audio signal to be rendered according to the control information; or

[0230] performing dynamic range compression on the audio signal to be rendered according to the control information; or

[0231] performing binaural rendering on the audio signal to be rendered according to the control information; or

[0232] performing loudspeaker rendering on the audio signal to be rendered according to the control information.

[0233] In some embodiments, the to-be-rendered audio signal comprises at least one of a channel-based audio signal, an object-based audio signal, or a scene-based audio signal, the obtaining module 1501 is further configured to: obtain, by decoding the bitstream, first reverberation information, the first reverberation information comprising at least one of first reverberation output loudness information, first time difference information of direct sound and early reflection sound, first reverberation duration information, first room shape and size information, or first sound scattering degree information. The rendering module 1503 is configured to: perform, according to the control information, control processing on the to-be-rendered audio signal to obtain a control-processed audio signal, the control processing comprising at least one of performing initial 3DoF processing on the channel-based audio signal, performing transformation processing on the object-based audio signal, or performing initial 3DoF processing on the scene-based audio signal, and perform, according to the first reverberation information, reverberation processing on the control-processed audio signal to obtain a first audio signal. Perform binaural rendering or loudspeaker rendering on the first audio signal to obtain the rendered audio signal.

[0234] In some embodiments, the rendering module 1503 is configured to: perform, according to the control information, signal format conversion on the first audio signal to obtain a second audio signal. Perform binaural rendering or loudspeaker rendering on the second audio signal to obtain the rendered audio signal.

[0235] In some embodiments, the signal format conversion comprises at least one of: converting channel-based audio signal in the first audio signal into scene-based or object-based audio signal; or converting scene-based audio signal in the first audio signal into channel-based or object-based audio signal; or converting object-based audio signal in the first audio signal into channel-based or scene-based audio signal.

[0236] In some embodiments, the rendering module 1503 is configured to: perform, according to the control information, signal format conversion on the first audio signal based on a signal format of the first audio signal and a processing capability of the terminal device.

[0237] In some embodiments, the rendering module 1503 is configured to: obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information comprising at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information. Perform, according to the control information and the second reverberation information, local reverberation processing on the second audio signal to obtain a third audio signal. Perform binaural rendering or loudspeaker rendering on the third audio signal to obtain the rendered audio signal.

[0238] In some embodiments, the rendering module 1503 is configured to: perform clustering processing on audio signals of different signal formats in the second audio signals respectively according to the control information, to obtain at least one of a channel-based group signal, a scene-based group signal, or an object-based group signal; and perform local reverberation processing on the at least one of the channel-based group signal, the scene-based group signal, or the object-based group signal according to the second reverberation information, to obtain the third audio signals.

[0239] In some embodiments, the rendering module 1503 is configured to: perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on audio signals of each signal format in the third audio signals according to the control information, to obtain fourth audio signals; and perform binaural rendering or loudspeaker rendering on the fourth audio signals, to obtain the rendered audio signals.

[0240] In some embodiments, the rendering module 1503 is configured to: perform dynamic range compression on the fourth audio signals according to the control information, to obtain fifth audio signals; and perform binaural rendering or loudspeaker rendering on the fifth audio signals, to obtain the rendered audio signals.

[0241] In some embodiments, the rendering module 1503 is configured to: perform signal format conversion on the audio signals to be rendered according to the control information, to obtain sixth audio signals; and perform binaural rendering or loudspeaker rendering on the sixth audio signals, to obtain the rendered audio signals.

[0242] In some embodiments, the signal format conversion includes at least one of: converting channel-based audio signals in the audio signals to be rendered into scene-based or object-based audio signals; or converting scene-based audio signals in the audio signals to be rendered into channel-based or object-based audio signals; or converting object-based audio signals in the audio signals to be rendered into channel-based or scene-based audio signals.

[0243] In some embodiments, the rendering module 1503 is configured to: perform signal format conversion on the audio signals to be rendered according to the control information, a signal format of the audio signals to be rendered, and a processing performance of the terminal device.

[0244] In some embodiments, the rendering module 1503 is configured to: obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information comprising at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information; perform local reverberation processing on the to-be-rendered audio signal according to the control information and the second reverberation information to obtain a seventh audio signal; and perform binaural rendering or loudspeaker rendering on the seventh audio signal to obtain the rendered audio signal.

[0245] In some embodiments, the rendering module 1503 is configured to: perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on the audio signal of each signal format in the to-be-rendered audio signal according to the control information to obtain an eighth audio signal; and perform binaural rendering or loudspeaker rendering on the eighth audio signal to obtain the rendered audio signal.

[0246] In some embodiments, the rendering module 1503 is configured to: perform dynamic range compression on the to-be-rendered audio signal according to the control information to obtain a ninth audio signal; and perform binaural rendering or loudspeaker rendering on the ninth audio signal to obtain the rendered audio signal.

[0247] It should be noted that the above obtaining module 1501, control information generation module 1502, and rendering module 1503 can be applied to the audio signal rendering process at the encoding end.

[0248] It should also be noted that the specific implementation process of the obtaining module 1501, the control information generation module 1502, and the rendering module 1503 can refer to the detailed description of the above method embodiments, and for the sake of brevity of the description, it will not be repeated here.

[0249] Based on the same inventive concept as the above method, the embodiments of the present application provide a device for rendering an audio signal, for example, an audio signal rendering device. Please refer to FIG. 16. Figure 15 As shown in FIG. 16, the audio signal rendering device 1600 comprises:

[0250] A processor 1601, a memory 1602, and a communication interface 1603 (wherein the number of processors 1601 in the audio signal encoding device 1600 can be one or more, and Figure 15 for example, one processor is taken as an example in the description below). In some embodiments of the present application, the processor 1601, the memory 1602, and the communication interface 1603 can be connected through a bus or other means, wherein, Figure 15 for example, the connection through the bus is taken as an example.

[0251] The memory 1602 can include read-only memory and random access memory, and provide instructions and data to the processor 1601. A portion of the memory 1602 can also include non-volatile random access memory (NVRAM). The memory 1602 stores operating systems and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.

[0252] The processor 1601 controls the operation of the audio coding device, and can also be referred to as a central processing unit (CPU). In specific applications, various components of the audio coding device are coupled together through a bus system, which can include a data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, all the buses are referred to as a bus system in the figure.

[0253] The method disclosed in the above embodiments of the present application can be applied in the processor 1601 or implemented by the processor 1601. The processor 1601 can be an integrated circuit chip with processing capability. In the implementation process, each step of the above method can be completed by integrated logic circuits or instruction of software form in the processor 1601. The processor 1601 mentioned above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory 1602, and the processor 1601 reads the information in the memory 1602 and combines the hardware to complete the steps of the above method.

[0254] The communication interface 1603 can be used to receive or send digital or character information, for example, can be an input / output interface, a pin or a circuit, etc. For example, the above-mentioned encoded code stream is received through the communication interface 1603.

[0255] Based on the same inventive concept as the above method, the embodiments of the present application provide an audio rendering device, comprising: a non-volatile memory and a processor coupled with each other, the processor invoking program codes stored in the memory to execute part or all steps of the audio signal rendering method as described in one or more of the above embodiments.

[0256] Based on the same inventive concept as the above method, the embodiments of the present application provide a computer readable storage medium, which stores program codes, wherein the program codes comprise instructions for executing part or all steps of the audio signal rendering method as described in one or more of the above embodiments.

[0257] Based on the same inventive concept as the above method, the embodiments of the present application provide a computer program product, which, when running on a computer, causes the computer to execute part or all steps of the audio signal rendering method as described in one or more of the above embodiments.

[0258] The processor mentioned in the above embodiments can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method embodiments can be completed by integrated logic circuits or instructions in the form of software in the processor. The processor can be a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware coding processor to execute, or a combination of hardware and software modules in the coding processor to execute. The software module can be located in a storage medium such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable memory (EEPROM), register, or other mature storage media in the art. The storage medium is located in the storage memory, and the processor reads information in the storage memory to combine the hardware to complete the steps of the above method.

[0259] The memory mentioned in each of the above embodiments can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct Rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.

[0260] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0261] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0262] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0263] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0264] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit.

[0265] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to make a computer device (personal computer, server, or network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program codes that can be stored in the medium.

[0266] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of audio signal rendering, characterized by, The method comprises: obtaining a to-be-rendered audio signal by decoding a received code stream; obtaining control information, the control information being used for indicating content description metadata, rendering format flag information and tracking information, or the control information being used for indicating content description metadata, rendering format flag information and tracking information, and at least one of speaker configuration information, application scenario information, posture information or position information; rendering the to-be-rendered audio signal according to the control information to obtain a rendered audio signal; wherein the rendering the to-be-rendered audio signal according to the control information comprises rendering pre-processing of the to-be-rendered audio signal according to the control information; the to-be-rendered audio signal comprises at least one of a channel-based audio signal, an object-based audio signal or a scene-based audio signal; the rendering pre-processing of the to-be-rendered audio signal according to the control information to obtain a rendered audio signal comprises: obtaining first reverberation information by decoding the code stream, wherein the reverberation information comprises at least one of reverberation output loudness information, time difference information of direct sound and early reflection sound, reverberation duration information, room shape and size information or sound scattering degree information; performing control processing on the to-be-rendered audio signal according to the control information to obtain a control-processed audio signal, the control processing comprising at least one of initial three-degree-of-freedom (3DoF) processing of the channel-based audio signal, transformation processing of the object-based audio signal or initial 3DoF processing of the scene-based audio signal; performing reverberation processing on the control-processed audio signal according to the first reverberation information to obtain a first audio signal; performing binaural rendering or speaker rendering on the first audio signal to obtain the rendered audio signal; wherein the content description metadata is used for indicating a signal format of the to-be-rendered audio signal, the signal format comprising at least one of a channel-based signal format, a scene-based signal format or an object-based signal format; the rendering format flag information is used for indicating an audio signal rendering format, the audio signal rendering format comprising speaker rendering or binaural rendering; the speaker configuration information is used for indicating a layout of speakers; the application scenario information is used for indicating renderer scene description information; the tracking information is used for indicating whether the rendered audio signal changes with head rotation of a listener; the posture information is used for indicating a direction and amplitude of the head rotation; and the position information is used for indicating a direction and amplitude of body movement of the listener.

2. The method of claim 1, wherein, the rendering the to-be-rendered audio signal according to the control information further comprises at least one of: performing signal format conversion on the to-be-rendered audio signal according to the control information; or performing local reverberation processing on the to-be-rendered audio signal according to the control information; or performing group processing on the to-be-rendered audio signal according to the control information; or performing dynamic range compression on the to-be-rendered audio signal according to the control information; or ​ The audio signal to be rendered is rendered in both ears according to the control information; or... The audio signal to be rendered is rendered using a speaker according to the control information.

3. The method of claim 1, wherein, The step of performing binaural rendering or speaker rendering on the first audio signal to obtain the rendered audio signal includes: The first audio signal is converted into a signal format according to the control information to obtain the second audio signal; Perform binaural rendering or speaker rendering on the second audio signal to obtain the rendered audio signal; The signal format conversion includes at least one of the following: converting the channel-based audio signal in the first audio signal into a scene-based or object-based audio signal; or converting the scene-based audio signal in the first audio signal into a channel-based or object-based audio signal; or converting the object-based audio signal in the first audio signal into a channel-based or scene-based audio signal.

4. The method of claim 3, wherein, The step of converting the signal format of the first audio signal according to the control information includes: Based on the control information, the signal format of the first audio signal, and the processing performance of the terminal device, the first audio signal is converted into a signal format.

5. The method of claim 3, wherein, The step of performing binaural rendering or speaker rendering on the second audio signal to obtain the rendered audio signal includes: Obtain second reverberation information, which is the reverberation information of the scene where the rendered audio signal is located; The second audio signal is subjected to local reverberation processing based on the control information and the second reverberation information to obtain the third audio signal; The third audio signal is subjected to binaural rendering or speaker rendering to obtain the rendered audio signal.

6. The method of claim 5, wherein, The step of performing local reverberation processing on the second audio signal according to the control information and the second reverberation information to obtain the third audio signal includes: Based on the control information, the audio signals of different signal formats in the second audio signal are clustered to obtain at least one of channel-based group signals, scene-based group signals, or object-based group signals. Based on the second reverberation information, at least one of the channel-based group signal, the scene-based group signal, or the object-based group signal is subjected to local reverberation processing to obtain the third audio signal.

7. The method according to claim 5 or 6, characterized in that, When rendering the audio signal to be rendered according to the control information, and further includes group processing of the audio signal to be rendered according to the control information, the step of performing binaural rendering or speaker rendering on the third audio signal to obtain the rendered audio signal includes: According to the control information, the group signal of each signal format in the third audio signal is subjected to 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom 6DoF processing to obtain the fourth audio signal. The fourth audio signal is subjected to binaural rendering or speaker rendering to obtain the rendered audio signal.

8. The method of claim 7, wherein, The step of performing binaural rendering or speaker rendering on the fourth audio signal to obtain the rendered audio signal includes: perform dynamic range compression on the fourth audio signal according to the control information to obtain a fifth audio signal; perform binaural rendering or loudspeaker rendering on the fifth audio signal to obtain the rendered audio signal.

9. The method of claim 1, wherein, The performing rendering on the to-be-rendered audio signal according to the control information to obtain the rendered audio signal comprises: performing signal format conversion on the to-be-rendered audio signal according to the control information to obtain a sixth audio signal; performing binaural rendering or loudspeaker rendering on the sixth audio signal to obtain the rendered audio signal. The signal format conversion comprises at least one of the following: converting a channel-based audio signal in the to-be-rendered audio signal into a scene-based or object-based audio signal; or converting a scene-based audio signal in the to-be-rendered audio signal into a channel-based or object-based audio signal; or converting an object-based audio signal in the to-be-rendered audio signal into a channel-based or scene-based audio signal.

10. The method of claim 9, wherein, The performing signal format conversion on the to-be-rendered audio signal according to the control information comprises: performing signal format conversion on the to-be-rendered audio signal according to the control information, a signal format of the to-be-rendered audio signal, and a processing performance of a terminal device.

11. The method of claim 1, wherein, The performing rendering on the to-be-rendered audio signal according to the control information to obtain the rendered audio signal comprises: obtaining second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located, the second reverberation information comprising at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information; performing local reverberation processing on the to-be-rendered audio signal according to the control information and the second reverberation information to obtain a seventh audio signal; performing binaural rendering or loudspeaker rendering on the seventh audio signal to obtain the rendered audio signal.

12. The method of claim 1, wherein, The performing rendering on the to-be-rendered audio signal according to the control information to obtain the rendered audio signal comprises: performing real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom 6DoF processing on an audio signal of each signal format in the to-be-rendered audio signal according to the control information to obtain an eighth audio signal; performing binaural rendering or loudspeaker rendering on the eighth audio signal to obtain the rendered audio signal.

13. The method of claim 1, wherein, The performing rendering on the to-be-rendered audio signal according to the control information to obtain the rendered audio signal comprises: performing dynamic range compression on the to-be-rendered audio signal according to the control information to obtain a ninth audio signal; performing binaural rendering or loudspeaker rendering on the ninth audio signal to obtain the rendered audio signal.

14. An audio signal rendering apparatus, characterized by comprise: an obtaining module, configured to obtain a to-be-rendered audio signal by decoding a received code stream; The control information generation module is configured to obtain control information, wherein the control information is used to indicate content description metadata, rendering format flag information, and tracking information, or the control information is used to indicate content description metadata, rendering format flag information, and tracking information, and at least one of speaker configuration information, application scenario information, posture information, or position information. The rendering module is configured to perform rendering on the to-be-rendered audio signal according to the control information to obtain a rendered audio signal. The rendering module is further configured to perform, before the rendering on the to-be-rendered audio signal according to the control information, rendering pre-processing on the to-be-rendered audio signal. The to-be-rendered audio signal includes at least one of a channel-based audio signal, an object-based audio signal, or a scene-based audio signal, and the obtaining module is further configured to obtain first reverberation information by decoding the bitstream, wherein the first reverberation information includes at least one of first reverberation output loudness information, first time difference information of direct sound and early reflection sound, first reverberation duration information, first room shape and size information, or first sound scattering degree information. The rendering module is configured to perform control processing on the to-be-rendered audio signal according to the control information to obtain a control-processed audio signal, wherein the control processing includes at least one of initial three-degree-of-freedom (3DoF) processing on the channel-based audio signal, transformation processing on the object-based audio signal, or initial 3DoF processing on the scene-based audio signal; perform reverberation processing on the control-processed audio signal according to the first reverberation information to obtain a first audio signal; and perform binaural rendering or speaker rendering on the first audio signal to obtain the rendered audio signal. The content description metadata is used to indicate a signal format of the to-be-rendered audio signal, wherein the signal format includes at least one of a channel-based signal format, a scene-based signal format, or an object-based signal format; the rendering format flag information is used to indicate an audio signal rendering format, wherein the audio signal rendering format includes speaker rendering or binaural rendering; the speaker configuration information is used to indicate a layout of speakers; the application scenario information is used to indicate renderer scene description information; the tracking information is used to indicate whether the rendered audio signal changes with head rotation of a listener; the posture information is used to indicate a direction and amplitude of the head rotation; and the position information is used to indicate a direction and amplitude of body movement of the listener.

15. The apparatus of claim 14, wherein, The rendering module is further configured to perform at least one of the following: perform signal format conversion on the to-be-rendered audio signal according to the control information; or perform local reverberation processing on the to-be-rendered audio signal according to the control information; or perform group processing on the to-be-rendered audio signal according to the control information; or perform dynamic range compression on the to-be-rendered audio signal according to the control information; or perform binaural rendering on the to-be-rendered audio signal according to the control information; or perform speaker rendering on the to-be-rendered audio signal according to the control information. ​ 16. The apparatus of claim 14, wherein, The rendering module is configured to perform signal format conversion on the first audio signal according to the control information, to obtain a second audio signal; and perform binaural rendering or loudspeaker rendering on the second audio signal, to obtain the rendered audio signal. The signal format conversion includes at least one of the following: converting a channel-based audio signal in the first audio signal into a scene-based or object-based audio signal; or converting a scene-based audio signal in the first audio signal into a channel-based or object-based audio signal; or converting an object-based audio signal in the first audio signal into a channel-based or scene-based audio signal.

17. The apparatus of claim 16, wherein, The rendering module is configured to perform signal format conversion on the first audio signal according to the control information, a signal format of the first audio signal, and a processing performance of the terminal device.

18. The apparatus of claim 16, wherein, The rendering module is configured to obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signal is located. The rendering module is configured to perform local reverberation processing on the second audio signal according to the control information and the second reverberation information, to obtain a third audio signal. The rendering module is configured to perform binaural rendering or loudspeaker rendering on the third audio signal, to obtain the rendered audio signal.

19. The apparatus of claim 18, wherein, The rendering module is configured to perform clustering processing on different signal formats of the second audio signal according to the control information, to obtain at least one of a channel-based group signal, a scene-based group signal, or an object-based group signal; and perform local reverberation processing on the at least one of the channel-based group signal, the scene-based group signal, or the object-based group signal according to the second reverberation information, to obtain the third audio signal.

20. The apparatus of claim 18 or 19, wherein, The rendering module is configured to perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on a group signal of each signal format in the third audio signal according to the control information, to obtain a fourth audio signal. The rendering module is configured to perform binaural rendering or loudspeaker rendering on the fourth audio signal, to obtain the rendered audio signal.

21. The apparatus of claim 20, wherein, The rendering module is configured to perform dynamic range compression on the fourth audio signal according to the control information, to obtain a fifth audio signal. The rendering module is configured to perform binaural rendering or loudspeaker rendering on the fifth audio signal, to obtain the rendered audio signal.

22. The apparatus of claim 14, wherein, The rendering module is configured to perform signal format conversion on the to-be-rendered audio signal according to the control information, to obtain a sixth audio signal; and perform binaural rendering or loudspeaker rendering on the sixth audio signal, to obtain the rendered audio signal. The signal format conversion includes at least one of the following: converting a channel-based audio signal in the to-be-rendered audio signal into a scene-based or object-based audio signal; or converting a scene-based audio signal in the to-be-rendered audio signal into a channel-based or object-based audio signal; or converting an object-based audio signal in the to-be-rendered audio signal into a channel-based or scene-based audio signal.

23. The apparatus of claim 22, wherein, The rendering module is configured to perform signal format conversion on the audio signals to be rendered according to the control information, the signal format of the audio signals to be rendered, and the processing performance of the terminal device.

24. The apparatus of claim 14, wherein, The rendering module is configured to: obtain second reverberation information, the second reverberation information being reverberation information of a scene in which the rendered audio signals are located, the second reverberation information including at least one of second reverberation output loudness information, second time difference information of direct sound and early reflection sound, second reverberation duration information, second room shape and size information, or second sound scattering degree information; perform local reverberation processing on the audio signals to be rendered according to the control information and the second reverberation information, to obtain seventh audio signals; perform binaural rendering or loudspeaker rendering on the seventh audio signals, to obtain the rendered audio signals.

25. The apparatus of claim 14, wherein, The rendering module is configured to: perform real-time 3DoF processing, or 3DoF+ processing, or six-degree-of-freedom (6DoF) processing on each signal format of the audio signals to be rendered according to the control information, to obtain eighth audio signals; perform binaural rendering or loudspeaker rendering on the eighth audio signals, to obtain the rendered audio signals.

26. The apparatus of claim 14, wherein, The rendering module is configured to: perform dynamic range compression on the audio signals to be rendered according to the control information, to obtain ninth audio signals; perform binaural rendering or loudspeaker rendering on the ninth audio signals, to obtain the rendered audio signals.

27. An audio signal rendering apparatus, characterized by comprise: a non-volatile memory and a processor coupled to each other, the processor invoking program code stored in the memory to execute the method according to any one of claims 1 to 13.

28. An audio signal rendering device, characterized by comprise: a renderer configured to execute the method according to any one of claims 1 to 13.

29. A computer-readable storage medium, characterized in that, comprise a computer program which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Methods, apparatus and systems for 6DOF audio rendering and data representations and bitstream structures for 6DOF audio rendering

    WO2019197404A1