System and method for rendering real-time spatial audio in a virtual environment

By processing mono audio signals through a real-time spatial audio rendering system to generate stereo audio, the problem of insufficient audio experience in virtual environments is solved, and more realistic audio source location and direction perception is achieved.

CN116095594BActive Publication Date: 2025-11-11AGORA LAB INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210666397.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-08
Filing Date
2022-06-13
Publication Date
2025-11-11
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In existing technologies, audio in real-time communication virtual environments is usually in mono format, which cannot effectively provide binaural effect and position awareness, resulting in an insufficient audio experience in virtual environments.

Method used

A real-time spatial audio rendering system is used to process mono audio signals through computer software applications to generate stereo audio. This includes determining the dynamic position of the audio source, converting the head-related impulse response, calculating the time difference between the ears, and controlling the gain. The system combines the room impulse response to generate a reverberation effect, and finally plays the stereo audio on a communication device.

Benefits of technology

It enables the provision of stereo audio to listeners in a virtual environment, enhances the perception of direction and distance of the audio source, and improves the audio experience in the virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116095594B_ABST
    Figure CN116095594B_ABST
Patent Text Reader

Abstract

This invention provides a novel real-time spatial audio rendering system, including a real-time spatial audio rendering computer software application that can run on a communication device. This application renders a mono audio source in a virtual room into stereo audio for a listener. The listener is movable. Stereo audio is rendered for each listener in the room. The real-time spatial audio rendering system has two different modes: with reverb and without reverb. Reverb can provide a sense of room dimension. First, a direct sound processing module generates direct sound stereo audio, which can reflect the direction and distance of spatial audio. When reverb is needed, a reverb processing module is also executed so that the final generated spatial audio can reflect the room dimension.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Application No. 17 / 520,956, filed on November 8, 2021. Technical Field

[0003] This invention relates to an audio rendering technique in real-time communication, and more specifically, to a real-time spatial audio rendering technique in a virtual environment. More specifically, this invention relates to a system and method for rendering real-time stereo audio in a virtual environment. Background Technology

[0004] In real-world communication, people can hear sounds from their source and distinguish their direction and distance. This is determined by the binaural effect. The binaural effect requires that the time delay and spectral energy distribution of the sound wave signals received by the listener's two ears be different. Therefore, spatial audio should have at least two channels (stereo audio) to provide users with a binaural effect in real-time communication environments (such as online gaming environments). Participants (or simply participants) are located in different room conditions in a real-time communication (RTC) virtual environment, such as an online conference room or virtual theater. They can also move from one place to another within their own room. There may be multiple audio sources in the room, such as voices, television, etc.

[0005] However, in real-time communication, many devices, such as laptops or mobile phones, may only support single-channel recording. Even if the device supports stereo recording, the audio codecs used by RTC applications may not support stereo audio. Therefore, audio in an RTC virtual environment is typically in mono format. Besides hardware and audio codec limitations, the location of each audio source can be variable in an RTC virtual environment. In other words, a new real-time spatial audio rendering system is needed for mono audio signals to generate stereo audio based on the real-time locations of the audio sources and the listener.

[0006] Therefore, a novel audio rendering system and method are needed to generate stereo audio for listeners in a virtual environment. A real-time spatial audio rendering system needs to transmit real-time stereo audio signals to each listener with minimal time delay, using the mono audio signal from the audio source, the real-time virtual positions of the listener and the audio source, and the listener's real-time orientation. The system renders the audio source and mixes it into a stereo playback format for transmission to listeners in the virtual room. Furthermore, listeners can distinguish the direction and distance of each audio source through stereo audio, making the virtual RTC environment closer to a real-world listening experience. Additionally, the system needs to generate stereo audio signals with reverberation effects. Summary of the Invention

[0007] Overall, the present invention provides a computer implementation method for rendering real-time spatial audio from a mono channel in a virtual environment. This method is executed by a real-time spatial audio rendering computer software application within a real-time spatial audio rendering system. Specifically, it includes: determining whether reverberation is applied to the spatial audio of a set of mono audio sources; determining the dynamic position set of each audio source in the mono audio source set relative to the listener in the virtual environment; obtaining a discrete set of head-related impulse responses (HRIRs); converting the discrete HRIR set into a continuous HRIR set; determining the interaural time difference of each mono audio source in the mono audio source set based on the dynamic position set; modifying the continuous HRIR according to the interaural time difference to generate a modified HRIR; applying gain control to the audio signal of each mono audio source in the mono audio source set to generate a modified audio signal; performing a convolution operation on the modified audio signal according to the modified HRIR to generate a spatial audio signal for each mono audio source in the mono audio source set; and combining the spatial audio signals of all mono audio sources in the mono audio source set to generate direct sound (reverberation-free) audio, which can be played by a communication device. Spatial audio is stereo audio. The method also includes compressing the level of the direct audio signal to a target range for playback by a communication device, wherein the spatial audio is stereo audio.

[0008] When setting reverberation, the method further includes: generating a binaural room impulse response (BRIR) based on the spatial dimensions of the room where the listener is located, the listener's position, and the aforementioned set of mono audio sources; using the BRIR to convolve the audio signals of each mono audio source in the set of mono audio sources to generate reverberated stereo audio for each mono audio source in the set of mono audio sources; combining the reverberated stereo audio from all mono audio sources in the set of mono audio sources to generate combined reverberated audio; and mixing the direct sound audio with the combined reverberated audio in both the left and right channels to generate the final spatial audio for playback on the communication device.

[0009] This invention also provides a real-time spatial audio rendering system, which includes a real-time spatial audio rendering computer software application running on a communication device. The real-time spatial audio rendering computer software application is capable of: determining whether a reverberation effect is set when rendering spatial audio from a set of mono audio sources; determining the set of dynamic positions of each audio source in the set of mono audio sources relative to a listener in a virtual environment; obtaining a discrete set of head-related impulse responses (HRIRs); converting the discrete HRIR set into a continuous HRIR set; determining the interaural time difference of each mono audio source in the set of mono audio sources; modifying the continuous HRIR based on the interaural time difference to generate a modified HRIR; applying gain control to the audio signal of each mono audio source in the set of mono audio sources to generate a modified audio signal; performing a convolution operation on the modified audio signal based on the modified HRIR to generate a spatial audio signal for each mono audio source in the set of mono audio sources; and combining the spatial audio signals of all mono audio sources in the set of mono audio sources to generate a direct sound audio signal that can be played by the communication device. In some implementations, spatial audio is stereo audio. Real-time spatial audio rendering computer software applications can also be used to compress the level of direct sound audio to a target range for playback by communication devices.

[0010] When configuring reverberation, the real-time spatial audio rendering computer software application can also: generate a binaural room impulse response (BRIR) based on a set of dimensional data of the room where the listener is located, the listener's position, and the aforementioned set of mono audio sources; use the BRIR to convolve the audio signals of each mono audio source in the set of mono audio sources to generate reverberated stereo audio for each mono audio source in the set of mono audio sources; combine the reverberated stereo audio from all mono audio sources in the set of mono audio sources to generate combined reverberated audio; and mix the direct sound audio with the combined reverberated audio in both the left and right channels to generate the final spatial audio for playback on a communication device. In a further embodiment, the real-time spatial audio rendering computer software application is also suitable for compressing the level of the final spatial audio to a target range. Attached Figure Description

[0011] This patent or application document contains at least one color drawing. The Patent Office will, upon request and upon payment of the relevant fees, provide a copy of this patent or application with color drawings.

[0012] The functional features of the invention will be specifically pointed out in the claims, and the invention itself, its structure, and its method of use can also be better understood by referring to the following drawings and related descriptions. All the drawings of this invention are also an integral part of this invention, wherein the same reference numerals denote the same parts:

[0013] Figure 1 This is a flowchart illustrating the process of a real-time spatial audio rendering system generating spatial audio according to an embodiment of the present invention.

[0014] Figure 2 This is an example block diagram of a real-time communication system including a real-time spatial audio rendering system, drawn according to an embodiment of the present invention.

[0015] Figure 3 This is an example block diagram of a communication device including a real-time spatial audio rendering system, drawn according to an embodiment of the present invention.

[0016] Figure 4 This is an example block diagram of a computer server including a real-time spatial audio rendering system, drawn according to an embodiment of the present invention.

[0017] Figure 5 This is a flowchart illustrating the process by which a computer spatial audio rendering system renders mono audio signals from one or more audio sources into reverberant stereo audio, according to an embodiment of the present invention.

[0018] Figure 6 This is a schematic diagram illustrating the dynamic position of a set of mono audio sources in a virtual environment relative to the listener and their orientation, as described in an embodiment of the present invention.

[0019] Figure 7 This is a flowchart drawn according to an embodiment of the present invention, illustrating the process by which a spatial audio rendering system renders a mono audio signal from one or more audio sources into a reverberant stereo audio signal.

[0020] Figure 8 This is a schematic diagram of a virtual room drawn according to an embodiment of the present invention.

[0021] Those skilled in the art will understand that the figures are not necessarily drawn to scale for the purpose of simply and clearly illustrating the various elements in the above figures. The dimensions of some components in the figures may be enlarged relative to other components to aid in understanding the invention. Furthermore, the specific order of certain elements, parts, components, modules, steps, operations, events, and / or processes described or illustrated herein may be changed in practical application. Those skilled in the art will understand that, for the sake of simplicity and clarity, those well-known and readily understood useful and / or necessary elements in commercially feasible embodiments may not be described herein in order to clearly present the various embodiments of the invention. Detailed Implementation

[0022] A novel real-time (RT) spatial audio rendering system can output stereo audio with or without reverberation. Reverberation conveys the sense of dimension in a virtual room. Depending on the use case, reverberation is not always necessary, as excessive reverberation can reduce intelligibility and is unsuitable in certain situations, such as multi-party virtual conferences conducted over the internet. In some implementations, the real-time spatial audio rendering system includes a computer software application (also referred to herein as a real-time spatial audio rendering computer software application) running on a communication device to convert mono audio signals from one or more audio sources into stereo audio for a listener, wherein the communication device is operated by either the listener or a computer server. While the computer server performs spatial audio rendering, the computer software application acquires input data from the listener's communication device via an internet connection, generates stereo audio, and forwards the stereo audio data to the listener's communication device over the internet for playback. The spatial audio rendering software application includes one or more computer programs written in a computer software programming language, such as C, C++, C#, Java, etc.

[0023] Figure 1 This illustrates the process by which a real-time spatial audio rendering software application provides spatial audio (such as stereo audio), the entire process represented by 100. Reference Figure 1 At point 102, the real-time spatial audio rendering software application determines whether to set a reverb effect. If not, at point 104, the application renders spatial audio without reverb. The generated spatial audio incorporates the listener's direction and distance. In other words, the spatial audio conveys a sense of direction and distance. This spatial audio is also referred to herein as direct sound audio, direct sound audio signal, or dry sound. If reverb is desired, at point 106, the application renders spatial audio with reverb. Users can set whether reverb is needed through a user input interface or configuration.

[0024] Figure 2 , 3 Figures 4 and 4 further illustrate communication equipment and computer servers. Figure 2 This is a schematic block diagram of a real-time communication system, represented by 200. 202 and 204 represent two example communication devices. 206 represents a computer server. Electronic devices 202-206 can all access the Internet 208.

[0025] Figure 3 Communication devices 202 (such as laptops, tablets, smartphones, etc.) are shown. Figure 3 This is a schematic block diagram of communication device 202. Device 202 includes a processor 302, a memory 304 adapted to the processor 302 and having a certain capacity, an audio output interface 306 adapted to the processor 302 (such as headphones), a network interface 308 (such as a WiFi network interface) adapted to the processor 302 and capable of accessing the Internet 208, and other interfaces 310 (such as video output interfaces and audio input interfaces). Device 202 also includes an operating system 322 running on the processor 302 (such as... (etc.). One or more computer software applications 324 (such as the novel real-time spatial audio rendering software application described above) are loaded and run on device 202. The computer software application 324 is implemented in a computer software programming language (such as C, C++, C#, Java, etc.).

[0026] Figure 4 Further details were provided for computer server 206. Figure 4 This is a schematic block diagram of computer server 206. Computer server 206 includes a processor 402, a memory 404 adapted to the processor 402 and having a certain capacity, and a network interface 406 (such as a WiFi network interface) adapted to the processor 402 and connected to the Internet 208. Computer server 206 also includes an operating system 422 (such as an operating system 422) running on the processor 402. One or more computer software applications 424 (such as the novel real-time spatial audio rendering software application described above) run on computer server 206. The novel real-time spatial audio rendering software application 424, as a server software application, is implemented using a computer software programming language (such as C, C++, C#, Java, etc.).

[0027] Figure 5The flowchart illustrates the process by which a spatial audio rendering software application 324 (or 424) renders mono audio signals from one or more audio sources into reverberant stereo audio, denoted as 500. At 502, the spatial audio rendering software application determines the set of dynamic positions of each audio source in a set of mono audio sources (i.e., one or more) relative to a listener. The dynamic position of each audio source is time-dependent because the listener can move. As the listener moves, the position of the audio source relative to the listener changes at different times. Figure 6 The concept of dynamic position is further elaborated.

[0028] Figure 6 This is a schematic diagram illustrating the dynamic position of the audio source relative to the listener and the listener's orientation. In this simulation scenario, there are two audio sources, P1 and P2, represented by points P1(α1,β1) and P2(α2,β2), respectively. The listener's position is represented by the origin of the coordinate system, and the listener's orientation is represented by the Y-axis. The dynamic positions of the two audio sources P1 and P2 at time t are represented by P1[t] and P2[t], respectively. Each dynamic position is represented by the azimuth angle α, elevation angle β, and distance d. The azimuth angle α is the angle between the horizontal plane and the counterclockwise direction of the Y-axis. The elevation angle β is the angle between the horizontal plane and the vertical / mid-plane formed by the X-axis and Y-axis. Therefore, the elevation angle β is positive in the Z-axis direction and negative in the opposite direction of the Z-axis. The distance d is the Eulerian distance between the audio source and the listener. Therefore, the time-dependent dynamic positions P1[t] and P2[t] of the two audio sources can be represented as (α1,β1,d1) and (α2,β2,d2), respectively. In the virtual environment system, the dynamic positions P1[t] and P2[t] are provided in real time.

[0029] Back Figure 5 At step 504, the spatial audio rendering software application obtains a discrete set of head-related impulse responses (HRIRs). In some embodiments, this set of discrete HRIRs can be pre-recorded and presented as a data table. For example, the discrete HRIR set can be measured at azimuth and elevation angles of 15 degrees and at distances of 1 meter. At each discrete angle (α, β) and distance, there is a set of HRIR data representing the left and right HRIRs. At step 506, the spatial audio rendering software application converts the discrete HRIR set into continuous HRIRs. In some embodiments, this conversion is achieved through interpolation, such as linear interpolation. In this invention, steps 504-506 are collectively referred to as determining continuous HRIRs.

[0030] In a real-time scenario, the distance between the listener and the audio source may change as the listener moves. Therefore, the distance between the audio source and the listener's ears also changes. This delay difference is crucial for the listener's spatial perception. Therefore, at 508, spatial audio rendering software applications determine the interaural time difference (ITD) for each mono audio source within the audio source set by calculating the distance from the audio source to each of the listener's ears and dividing that distance by the speed of sound. The ITD calculation formula is as follows:

[0031] Where a represents the listener's head circumference, c represents the speed of sound, and θ I θ is the interaural azimuth angle measured in radians. For an audio source on the listener's left, θ I The value ranges from 0 to π / 2; for the audio source to the right of the listener, θ I The value ranges from π / 2 to π.

[0032] At position 510, the spatial audio rendering software application modifies the continuous HRIR using the interaural time difference, generating the modified HRIR. In some implementations, some zero-valued samples can be added to the continuous HRIR. For example, if the audio source is on the left, the ITD is 1 ms, and the HRIR sampling rate is 48000 Hz, then 48 zero-valued samples can be added at the beginning of the right-side HRIR.

[0033] At position 512, the spatial audio rendering software application applies gain control to the mono audio signal from the audio source. Specifically, at position 512, the volume of the audio source is adjusted based on the distance between the mono audio source and the listener. The gain used to adjust the volume is applied to the audio signal from the audio source. The gain follows the volume propagation attenuation rule. In some implementations, the gain calculation formula is as follows:

[0034]

[0035] Where A(d) is the gain at a distance d, d ref It is a reference distance, A ref This is the reference gain. d ref and A ref These are predefined parameters, which means that at a distance d ref At this point, the gain applied to the mono audio signal is A. ref The modified audio signal is generated by multiplying the mono audio signal by A(d).

[0036] At point 514, the spatial audio rendering software application performs a convolution operation on the modified mono audio signal using the modified HRIR (right and left ear) to generate a stereo audio signal from the audio source. The stereo audio signal includes the right and left channels. At point 516, the spatial audio rendering software application combines the audio source set (such as...) Figure 6 The stereo audio signals from each audio source (P1 and P2) shown are used to generate a combined (or mixed) stereo audio signal, which is then played on the listener's communication device 202 (or 204). For example, the audio signals from all audio sources can be added together to achieve mixing. The combined stereo audio signal generated in step 514 is also referred to herein as direct sound audio, dry sound, direct sound stereo audio, direct sound stereo audio data, and direct sound stereo audio signal. It should be noted that if there is only one audio source, step 516 maintains the same audio signal. In a further embodiment, at 518, spatial audio rendering software is applied to compress the mixed audio signal to prevent it from being too loud. For example, at 518, a dynamic audio compressor is used to compress the level of the mixed spatial audio signal to within a target range to prevent the mixed spatial audio from being too loud. The compressed spatial audio at 518 is also referred to in this invention as compressed direct sound audio.

[0037] When spatial audio rendering requires room reverberation, reverberation based on binaural room impulse response (BRIR) is added during spatial audio rendering. Figure 7 A flowchart illustrating the process by which a spatial audio rendering software application renders a mono audio signal from one or more audio sources into a stereo audio signal with reverberation is shown, the overall process indicated by 700. At 702, the spatial audio rendering software application generates direct sound audio. In some embodiments, at 702, the spatial audio rendering software application performs steps 502-516. The stereo audio generated by step 516 is the direct sound audio.

[0038] At 704, the spatial audio rendering software application generates a BRIR based on the room's dimensions and the locations of the listener and the audio source. Figure 8This is a schematic diagram of a virtual room. In the virtual environment (also referred to as a room or virtual room in this invention), there are three audio sources: Audio1 (a1, b1, c1), Audio2 (a2, b2, c2), and Audio3 (a3, b3, c3). The listener (as shown in the avatar) is located at (a0, b0, c0). The width, length, and height of the room are represented by a, b, and c, respectively. In some implementations, a real-time BRIR can be generated using an image method or an image source (ISM). If the ISM method is used, the sound wave signal reflected by the wall when it hits a hard wall is considered as a sound wave from an image source behind the wall. The sound wave is reflected multiple times before reaching the listener's ear. Therefore, reverberation can be simulated by the sum of a finite number of image sources. Due to the different spatial dimensions of the room and the different positions of the listener and the mono audio sources, the reflection paths between the audio sources and the listener are also different. A BRIR set can be estimated for each audio source using the ISM method.

[0039] At point 706, the spatial audio rendering software application uses BRIR to convolve the mono audio signal from the audio source, generating reverberant stereo audio (also referred to as reverberant audio or reverberant audio signal in this invention). At point 708, the spatial audio rendering software application combines the reverberant stereo audio signals from all audio source signals in the generated audio source set (such as P1 and P2) to generate a combined reverberant stereo audio signal (or simply reverberant audio). In some embodiments, the combination operation can be achieved by adding the reverberant stereo audio signals of the audio source set using the following equation:

[0040]

[0041] Where S i This represents the reverberant stereo audio data of the i-th audio source, where n represents the number of audio sources.

[0042] At 710, the spatial audio rendering software application mixes the direct-sound stereo audio and the combined reverberant stereo audio for both the left and right channels, generating the final stereo audio for playback on device 202. In some embodiments, mixing simply adds the two types of audio data together. In a further embodiment, at 712, the spatial audio rendering software application also compresses the level of the final audio signal to a target range to prevent the playback sound from being too loud. For example, at 712, a dynamic audio compressor is used to compress the level of the final audio signal to within the target range.

[0043] Based on the above description, it is obvious that many other modifications and variations are possible with respect to the present invention. Therefore, please note that within the scope of the appended claims, the present invention can be implemented in ways different from those specifically described above.

[0044] The foregoing description of the present invention is for better illustration and explanation, and is not intended to be exclusive or to limit the invention to the specific forms described above. The foregoing description is intended to better explain the principles of the invention and their practical application, so that those skilled in the art can best utilize the invention to implement various embodiments and make various modifications for the intended specific uses. It should also be noted that the words "a" or "an" in this invention include both singular and plural forms. Conversely, where appropriate, the multiple elements mentioned in the invention should also include their singular forms.

[0045] The scope of this invention is not limited to the contents of the above description, but is defined by the claims. Furthermore, although the claims set forth below may seem narrow, it should be understood that the scope of this invention is much broader than that proposed by the claims. We will file broader claims in one or more applications claiming priority to this application. Any content disclosed in the above description and drawings that is not included within the scope of the claims is not disclosed herein, and we reserve the right to file one or more patent applications with respect to such content in the future.

Claims

1. A computer implementation method for rendering a mono audio source into real-time spatial audio in a virtual environment, the method being executed by a real-time spatial audio rendering computer software application in a real-time spatial audio rendering system, and the method comprising: 1) Determine if reverb was set when rendering a collection of mono audio sources as spatial audio; 2) Determine the set of dynamic positions of each audio source in the set of mono audio sources relative to the listener in the virtual environment; 3) Obtain a discrete HRIR set; 4) Convert the discrete HRIR set into a continuous HRIR set; 5) Based on the dynamic position set, determine the interaural time difference of each mono audio source in the mono audio source set; 6) Modify the continuous HRIR based on the interaural time difference to generate the modified HRIR; 7) Apply gain control to the audio signal of each mono audio source in the mono audio source set to generate a modified audio signal; 8) Perform convolution operation on the modified audio signal according to the modified HRIR to generate the spatial audio signal of each mono audio source in the mono audio source set; as well as 9) Combine the spatial audio signals of all mono audio sources in the mono audio source set to generate direct audio, which can be played by a communication device.

2. The method according to claim 1 further includes compressing the level of the direct audio signal to a target range for playback by the communication device.

3. The method according to claim 1, if reverberation is required, the method further includes: 1) Generate a BRIR based on the room dimensions of the listener's room, the listener's location, and the set of mono audio sources; 2) Using the BRIR, perform convolution operation on the audio signal of each mono audio source in the mono audio source set to generate reverberant stereo audio for each mono audio source in the mono audio source set. 3) Combine the reverberant stereo audio from all the mono audio sources in the mono audio source set to generate a combined reverberant audio; as well as 4) The direct audio frequency is mixed with the combined reverberation audio frequency in both the left and right channels to generate the final spatial audio for playback on the communication device.

4. The method of claim 3 further includes compressing the level of the final spatial audio to within the target range.

5. The method according to any one of claims 1 to 4, wherein the spatial audio is stereo audio.

6. A real-time spatial audio rendering system, the system comprising a real-time spatial audio rendering computer software application running on a communication device, the real-time spatial audio rendering computer software application being configured to: 1) Determine if reverb was set when rendering a collection of mono audio sources as spatial audio; 2) Determine the set of dynamic positions of each audio source in the set of mono audio sources relative to the listener in the virtual environment; 3) Obtain a discrete HRIR set; 4) Convert the discrete HRIR set into a continuous HRIR set; 5) Determine the interaural time difference of each mono audio source in the set of mono audio sources based on the dynamic position set; 6) Modify the continuous HRIR based on the interaural time difference to generate the modified HRIR; 7) Apply gain control to the audio signal of each mono audio source in the mono audio source set to generate a modified audio signal; 8) Perform convolution operation on the modified audio signal according to the modified HRIR to generate the spatial audio signal of each mono audio source in the mono audio source set; as well as 9) Combine the spatial audio signals of all mono audio sources in the mono audio source set to generate direct audio, which can be played by a communication device.

7. The real-time spatial audio rendering system according to claim 6, wherein the real-time spatial audio rendering system compresses the level of the direct audio signal to a target range for playback by the communication device.

8. The real-time spatial audio rendering system according to claim 6, if reverb is required, the real-time spatial audio rendering system is further configured as follows: 1) Generate a BRIR based on the room dimensions of the listener's room, the listener's location, and the set of mono audio sources; 2) Use the BRIR to perform convolution operation on the audio signal of each mono audio source in the mono audio source set to generate reverberant stereo audio of each mono audio source in the mono audio source set. 3) Combine the reverberant stereo audio from all mono audio sources in the mono audio source set to generate a combined reverberant audio; and 4) The direct audio frequency is mixed with the combined reverberation audio in both the left and right channels to generate the final spatial audio for playback on the communication device.

9. The real-time spatial audio rendering system according to claim 8, wherein the real-time spatial audio rendering system compresses the level of the final spatial audio to within the target range.

10. The real-time spatial audio rendering system according to any one of claims 6 to 9, wherein the spatial audio is stereo audio.

Citation Information

Patent Citations

  • Personalized virtual audio playback method based on binaural real-time measurement

    CN108616789A

  • System for and method of generating an audio image

    US20190261124A1