Audio rendering method and device, electronic equipment, storage medium and computer program product
By creating a virtual sound source corresponding to the physical sound source in a virtual scene and performing audio rendering based on the virtual sound source, the problem of low simulation degree of virtual audio in the existing technology is solved, and a high simulation degree effect of virtual audio is achieved.
Patent Information
- Application Number
- CN202511264805.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, audio rendering cannot accurately arrange virtual sound sources, resulting in a lack of accurate estimation of sound source locations and thus low virtual audio simulation.
By acquiring the physical location information of the physical sound source in the physical scene, and converting the physical location information based on the virtual objects included in the virtual scene to obtain virtual location information, a virtual sound source corresponding to the physical sound source is created in the virtual scene, and audio rendering is performed on the sound signal generated by the physical sound source based on the virtual sound source, ensuring that the position of the virtual sound source corresponds logically to the actual location of the physical sound source.
It improves the simulation of virtual audio, making the rendered audio highly consistent with the real listening experience in terms of spatial attributes such as orientation and distance, and significantly enhances the spatial realism of virtual audio.
Smart Images

Figure CN120935501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an audio rendering method, apparatus, electronic device, storage medium, and computer program product. Background Technology
[0002] Virtual reality is a computer technology that can create and experience simulated three-dimensional virtual environments. Its core is to generate highly realistic digital scenes through computers and combine them with special equipment (such as VR headsets, controllers, motion sensors, etc.) to simulate multiple sensory experiences such as human vision, hearing, and touch, so that users can have a sense of immersion in the environment psychologically and physiologically, as if they are in an interactive virtual world.
[0003] In related technologies, audio rendering typically involves dynamically adjusting the audio stream based on the head motion data of virtual objects to create a spatial sound experience. However, existing technologies cannot accurately position virtual sound sources, resulting in low fidelity in the rendered audio due to the lack of accurate estimation of sound source locations. Summary of the Invention
[0004] This application provides an audio rendering method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the simulation level of virtual audio.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides an audio rendering method, including:
[0007] Obtain the physical location information of the physical sound source in the physical scene;
[0008] Based on the virtual objects included in the virtual scene, the physical location information of the physical sound source is converted to obtain the virtual location information in the virtual scene corresponding to the physical location information;
[0009] Based on the virtual location information, a virtual sound source corresponding to the physical sound source is created in the virtual scene;
[0010] Based on the virtual sound source, the sound signal generated by the physical sound source is rendered to obtain rendered audio, and based on the rendered audio, a virtual audio for sending to the virtual object is determined.
[0011] This application provides an audio rendering apparatus, including:
[0012] The acquisition module is used to acquire the physical location information of the physical sound source in the physical scene;
[0013] The conversion module is used to convert the physical location information of the physical sound source based on the virtual objects included in the virtual scene, so as to obtain the virtual location information in the virtual scene corresponding to the physical location information;
[0014] A creation module is used to create a virtual sound source in the virtual scene that corresponds to the physical sound source based on the virtual orientation information;
[0015] The rendering module is used to perform audio rendering on the sound signal generated by the physical sound source based on the virtual sound source, to obtain rendered audio, and to determine the virtual audio to be sent to the virtual object based on the rendered audio.
[0016] In the above scheme, the conversion module is further used to obtain the virtual scene coordinate system of the virtual scene, wherein the virtual scene coordinate system takes the center position of the virtual scene as the origin; based on the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system, the virtual scene coordinate system is converted to obtain the object coordinate system with the virtual object as the origin; the transformation matrix between the object coordinate system and the physical scene coordinate system of the physical scene is determined, and based on the transformation matrix, the physical orientation information of the physical sound source is converted to obtain the virtual orientation information.
[0017] In the above scheme, the above conversion module is further used to determine the difference vector between the coordinates of the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system as the offset vector between the virtual scene coordinate system and the object coordinate system; and to perform a translation transformation on the virtual scene coordinate system according to the offset vector to obtain the object coordinate system with the virtual object as the origin.
[0018] In the above scheme, the transformation matrix includes a position transformation matrix and a direction transformation matrix. The above transformation module is further used to determine the position transformation matrix based on the position coordinates of the origin of the object coordinate system in the physical scene coordinate system; for each coordinate axis of the object coordinate system, determine the unit direction vector of the coordinate axis in the physical scene coordinate system; and construct a transformation matrix between the object coordinate system and the physical scene coordinate system based on the unit direction vector of each coordinate axis.
[0019] In the above scheme, the physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system, and the physical direction vector of the physical sound source in the physical scene coordinate system. The transformation matrix includes a position transformation matrix and a direction transformation matrix. The virtual orientation information includes virtual position coordinates and a virtual direction vector. The above conversion module is further used to determine the virtual position coordinates by multiplying the physical position coordinates by the position transformation matrix and to determine the virtual direction vector by multiplying the physical direction vector by the direction transformation matrix.
[0020] In the above scheme, the virtual orientation information includes virtual position coordinates and virtual direction vector. The above creation module is also used to create an initial virtual sound source corresponding to the physical sound source at the virtual position coordinates in the virtual scene; and to adjust the initial virtual sound source based on the virtual direction vector to obtain the virtual sound source.
[0021] In the above scheme, the creation module is further used to adjust the orientation of the initial sound source to be consistent with the virtual direction vector to obtain a candidate virtual sound source; obtain the acoustic feature parameters of the physical sound source, and adjust the acoustic feature parameters of the physical sound source based on the virtual scene to obtain the adjusted acoustic feature parameters; set the acoustic feature parameters of the candidate virtual sound source to the adjusted acoustic feature parameters to obtain the virtual sound source.
[0022] In the above scheme, the number of physical sound sources is at least one, and the virtual sound sources correspond one-to-one with the physical sound sources; the rendering module is further configured to, for each virtual sound source, perform audio rendering on the sound signal generated by the corresponding physical sound source based on the virtual sound source, to obtain the rendered audio corresponding to the virtual sound source; the rendering module is further configured to, when the number of virtual sound sources is multiple, synthesize the rendered audio corresponding to multiple virtual sound sources to obtain the virtual audio; when the number of virtual sound sources is one, determine the rendered audio corresponding to the virtual sound source as the virtual audio.
[0023] In the above scheme, the rendering module is further used to cluster the virtual sound sources to obtain at least one virtual sound source group, and the virtual audio corresponds one-to-one with the virtual sound source group; for each virtual sound source group, the rendered audio corresponding to each virtual sound source in the virtual sound source group is synthesized to obtain the virtual audio corresponding to the virtual sound source group.
[0024] In the above scheme, the rendering module is further configured to determine the relative orientation information between the virtual sound source and the virtual object based on the virtual orientation information corresponding to the physical orientation information in the virtual scene; and to perform audio rendering on the sound signal generated by the physical sound source based on the relative orientation information and the acoustic feature parameters of the virtual sound source to obtain the rendered audio.
[0025] In the above scheme, the rendering module is further configured to match the sound signal generated by the physical sound source with the acoustic feature parameters to obtain a matching result; when the matching result indicates that the sound signal does not match the acoustic feature parameters, the sound signal is adjusted based on the acoustic feature parameters to obtain a candidate sound signal; and the candidate sound signal is rendered based on the relative orientation information to obtain the rendered audio.
[0026] In the above scheme, the virtual object includes multiple audio receiving units, the relative orientation information includes the relative orientation sub-information between the virtual sound source and each audio receiving unit, and the rendered audio includes the sub-rendered audio corresponding to each audio receiving unit; the rendering module is further configured to perform audio rendering on the candidate sound signal for each audio receiving unit based on the relative orientation sub-information to obtain the sub-rendered audio corresponding to each audio receiving unit; when the number of virtual sound sources is multiple, determining the virtual audio to be sent to the virtual object based on the rendered audio includes: for each audio receiving unit, synthesizing the sub-rendered audio corresponding to each virtual sound source for that audio receiving unit to obtain the virtual audio of the audio receiving unit.
[0027] In the above scheme, the rendering module is further configured to acquire the acoustic characteristics of the candidate sound signal, and determine the target acoustic characteristics of the rendered audio based on the relative orientation information between the virtual sound source and the virtual object and the acoustic characteristics of the candidate sound signal; and perform audio rendering on the candidate sound signal based on the target acoustic characteristics to obtain the rendered audio.
[0028] In the above scheme, the rendering module is further configured to adjust the acoustic characteristics of the candidate sound signal based on the relative orientation information to obtain candidate rendered audio; select the target filter coefficient corresponding to the virtual sound source from multiple preset filter coefficients based on the relative orientation information; and perform audio rendering on the candidate rendered audio based on the target filter coefficient to obtain the rendered audio.
[0029] In the above scheme, the rendering module is further configured to determine the relative orientation information between the virtual sound source and the virtual object based on the virtual orientation information corresponding to the physical orientation information in the virtual scene; determine the rendering precision of the sound signal based on the relative orientation information between the virtual sound source and the virtual object; and perform audio rendering on the sound signal generated by the physical sound source according to the rendering precision of the sound signal based on the virtual sound source to obtain the rendered audio.
[0030] In the above scheme, the rendering module is further configured to obtain the sound source type of the physical sound source, compare the sound source type with a preset sound source type, and obtain a comparison result; when the comparison result indicates that the sound source type is the preset sound source type, based on the virtual orientation information, search for the preset rendering audio corresponding to the virtual orientation information in the audio library of the preset sound source type, and determine the found preset rendering audio as the rendering audio; the rendering module is further configured to perform audio rendering on the sound signal generated by the physical sound source based on the virtual sound source when the comparison result indicates that the sound source type is not the preset sound source type, to obtain the rendering audio.
[0031] This application provides an electronic device, including:
[0032] Memory is used to store executable instructions or computer programs.
[0033] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the audio rendering method provided in the embodiments of this application.
[0034] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the audio rendering method provided in this application.
[0035] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. An electronic device's processor reads the computer-executable instructions or computer program from the computer program or computer-readable storage medium, and executes the computer-executable instructions or computer program, causing the electronic device to perform the audio rendering method described above in this application.
[0036] The embodiments of this application have the following beneficial effects:
[0037] By acquiring the physical location information of the physical sound source in the physical scene and transforming it based on virtual objects included in the virtual scene to obtain virtual location information, the spatial relationships of the physical world are accurately mapped in the virtual scene. This ensures a logical correspondence between the position of the virtual sound source and the actual location of the physical sound source, laying the foundation for the spatial realism of the virtual audio. Based on the virtual location information corresponding to the physical location information in the virtual scene, a virtual sound source corresponding to the physical sound source is created in the virtual scene. Using the virtual sound source as a reference, the sound signal generated by the physical sound source is rendered to obtain the rendered audio. This allows the rendering process to fully integrate the spatial characteristics of the virtual scene (such as the distance and angle between the virtual sound source and virtual objects), so that the rendered audio naturally carries the acoustic characteristics consistent with the virtual location (such as volume attenuation at long distances and spectral differences at specific angles). By determining the virtual audio to be sent to the virtual object based on the rendered audio, and ensuring that the entire processing revolves around the perception logic of the virtual object, the final generated virtual audio not only retains the original sound characteristics of the physical sound source, but also accurately restores its spatial orientation relative to the virtual object in the virtual scene. This makes the sound received by the virtual object highly consistent with the real auditory experience in terms of spatial attributes such as orientation and distance, thereby significantly improving the simulation degree of the virtual audio. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the architecture of the audio rendering system provided in the embodiments of this application;
[0039] Figure 2 This is a schematic diagram of the structure of an electronic device for audio rendering provided in an embodiment of this application;
[0040] Figure 3 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 1 ;
[0041] Figure 4 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 2 ;
[0042] Figure 5 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 3 ;
[0043] Figure 6 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 4 ;
[0044] Figure 7 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 5 ;
[0045] Figure 8 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 5 ;
[0046] Figure 9 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 1 ;
[0047] Figure 10 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 2 ;
[0048] Figure 11 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 3 ;
[0049] Figure 12 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 4 ;
[0050] Figure 13 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 5 ;
[0051] Figure 14 This is a schematic diagram of the audio rendering method provided in the embodiments of this application. Figure 6 ;
[0052] Figure 15 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 7 . Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0054] In the following description, some embodiments are referred to, which describe a subset of all possible embodiments. However, it is understood that some embodiments may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0055] In the following description, the terms first, second, and third are used only to distinguish similar objects and do not represent a specific ordering of objects. It is understood that first, second, and third may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0057] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0058] 1) Physical Sound Sources: Physical sound sources refer to entities or vibration sources that exist in real physical space and can produce sound. Their core characteristic is the generation of sound waves that can be directly perceived by the human auditory system through physical vibrations (such as mechanical vibrations, air vibrations, etc.). They possess a real physical form and location, and sound propagation follows the laws of physical acoustics (such as air conduction, reflection, diffraction, etc.). Examples include human vocal organs (vocal cord vibration), musical instruments (string vibration, air column vibration), loudspeakers (diaphragm vibration), and natural sounds like wind / thunder (air vibration). Sound from physical sound sources can be collected as electrical signals using devices such as microphones for subsequent processing (such as converting them into raw material for virtual sound sources).
[0059] 2) Virtual Sound Source: A virtual sound source is an acoustic abstraction with spatial attributes defined in a virtual scene (such as a VR / AR environment or digital 3D space). Its core feature is simulating the spatial propagation characteristics of sound through digital signal processing technology, allowing listeners to perceive that the sound comes from a specific location in the virtual space. Virtual sound sources do not have a real physical form; they are digital objects that carry virtual orientation and acoustic parameters (such as directivity and distance attenuation characteristics) and are used to construct a virtual sound field. Virtual sound sources combine the sound signals (or digital audio signals) of physical sound sources with the spatial information of the virtual scene, simulating the propagation effect of sound in virtual space through audio rendering (such as HRTF convolution and reverberation processing), creating the auditory illusion that the sound comes from a virtual location. The parameters of the virtual sound source (such as 3D coordinates and motion state) must match the virtual scene. It is a bridge connecting physical sound and the virtual sound field, and is widely used in virtual reality, 3D audio, game sound effects, and other fields. Physical sound sources are real sound producers that follow physical laws; virtual sound sources are digital acoustic models used to simulate sound localization in virtual space, relying on signal processing technology to achieve the perceptual effect.
[0060] 3) Virtual Scene: This refers to the virtual scene displayed (or provided) by the application when it runs on the terminal. The virtual scene can be a simulation of the real world, a semi-simulated / semi-fictional virtual environment, or a purely fictional virtual environment. The virtual scene can be any of a two-dimensional, 2.5-dimensional, or three-dimensional virtual scene; this application does not limit the dimension of the virtual scene. For example, a virtual scene may include the sky, land, ocean, etc., and the land may include environmental elements such as deserts and cities. Users can control virtual objects to move within this virtual scene.
[0061] 4) Virtual Audio: Virtual audio refers to audio signals generated using digital signal processing technology in virtual scenes (such as VR / AR environments, digital 3D spaces, virtual conference systems, etc.), possessing spatial attributes and environmental adaptability. Its core characteristic is simulating the propagation effect of sound in virtual space, allowing listeners to perceive that the sound originates from a specific location within the virtual scene or conforms to the acoustic characteristics of the virtual environment. Generated based on digital signal processing, it does not directly correspond to sound waves in real physical space. Instead, it simulates spatial positioning, distance attenuation, and environmental reverberation through algorithms, serving as the auditory representation of the virtual sound field. Typically, it uses sound signals from physical sound sources (such as human voices or instrument sounds captured by a microphone) as raw material, combined with parameters of the virtual sound source (such as virtual orientation and directivity) and the acoustic environment of the virtual scene (such as room reverberation and obstacle occlusion), and then processed using audio rendering techniques (such as HRTF convolution, dynamic filtering, and reverberation overlay).
[0062] 5) Physical Location Information: This is a set of quantitative information describing the spatial position and orientation of a physical sound source in a real physical scene. It is used to accurately locate the spatial attributes of the physical sound source and mainly includes the following two core elements: Physical Position Coordinates: These refer to the specific spatial point of the physical sound source in the physical scene coordinate system (such as a three-dimensional rectangular coordinate system), usually represented by (x, y, z) coordinate values (units such as meters). For example, in a room scene, a physical sound source located at (3m, 2m, 1.5m) means that the sound source is 1.5 meters above the ground, 3 meters from the room origin (such as a corner) along the x-axis, and 2 meters from the y-axis. Physical Direction Vector: This refers to the orientation or radiation direction of the physical sound source in the physical scene coordinate system, usually represented by a three-dimensional vector (dx, dy, dz) (which can be normalized to a unit vector). For example, a physical direction vector of (0, 1, 0) means that the main radiation direction of the physical sound source is along the positive y-axis (e.g., when the speaker faces this direction, the sound radiation intensity is highest in this direction).
[0063] 6) Virtual Reality (VR): A computer technology that can create and experience simulated three-dimensional virtual environments. Its core is to generate highly realistic digital scenes through computers and combine them with special equipment (such as VR headsets, controllers, motion sensors, etc.) to simulate human visual, auditory, tactile and other sensory experiences, so that users can have a sense of immersion in the environment psychologically and physiologically, as if they are in an interactive virtual world.
[0064] 7) Virtual Objects: Virtual objects are interactive representations of various people and things within a virtual scene, or movable objects within the virtual scene. These movable objects can be virtual characters, virtual animals, anime characters, etc., such as people, animals, plants, oil drums, walls, and stones displayed in the virtual scene. A virtual object can be a virtual avatar representing the user within the virtual scene. A virtual scene can include multiple virtual objects, each with its own shape and volume, occupying a portion of the virtual scene's space. Optionally, the virtual object can be a user character controlled through client-side operations, an artificial intelligence (AI) trained and set up for virtual scene battles, or a non-user character (NPC) set up for interaction in the virtual scene.
[0065] 8) Acoustic Characteristics: Acoustic characteristics refer to the measurable and describable physical properties and auditory features exhibited by sound during propagation and perception. They are the core attributes reflecting the essence of sound. Acoustic characteristics can include volume and pitch: Volume (also known as loudness) is the subjective perception of the strength of a sound, mainly determined by the amplitude (vibration amplitude) of the sound signal, usually quantified in decibels (dB). The larger the amplitude, the higher the volume, reflecting whether the sound is loud or weak. Pitch is the subjective perception of the highness or lowness of a sound, mainly determined by the frequency (vibration speed) of the sound signal. The higher the frequency, the higher the pitch (such as high-pitched notes), and the lower the frequency, the lower the pitch (such as bass drums), reflecting the difference in pitch.
[0066] During the implementation of the embodiments of this application, the applicant discovered the following problems with the related technology:
[0067] In related technologies, audio rendering typically involves dynamically adjusting the audio stream based on the head motion data of virtual objects to create a spatial sound experience. However, existing technologies cannot accurately position virtual sound sources, resulting in low fidelity in the rendered audio due to the lack of accurate estimation of sound source locations.
[0068] This application provides an audio rendering method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the simulation degree of virtual audio. The exemplary application of the audio rendering system provided in this application is described below.
[0069] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the audio rendering system 100 provided in the embodiments of this application. The terminal (terminal 400 is shown as an example) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0070] Terminal 400 is used by users to access client 410 and display an audio playback interface on graphical interface 410-1 (graphical interface 410-1 is shown as an example). Terminal 400 and server 200 are interconnected via wired or wireless network.
[0071] In some embodiments, server 200 can be a standalone physical server, a server cluster or business system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, in-vehicle terminal, etc., but is not limited to these. The electronic device provided in this application embodiment can be implemented as a terminal or a server. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment.
[0072] In some embodiments, the server 200 converts the physical location information of the physical sound source based on the virtual objects included in the virtual scene to obtain virtual location information, and creates a virtual sound source in the virtual scene corresponding to the physical sound source based on the virtual location information. Based on the virtual sound source, the server performs audio rendering on the sound signal to obtain virtual audio, and sends the virtual audio to the terminal 400.
[0073] In other embodiments, the terminal 400 converts the physical location information of the physical sound source based on the virtual objects included in the virtual scene to obtain virtual location information, and creates a virtual sound source in the virtual scene corresponding to the physical sound source based on the virtual location information. Based on the virtual sound source, the terminal performs audio rendering on the sound signal to obtain virtual audio, and sends the virtual audio to the server 200.
[0074] In some embodiments, in the application scenario of remote collaborative conferencing, the physical scene is an offline conference room, and the physical sound source is the microphone of participant A (physical location information: coordinates (3, 2, 1.5) meters, located on the east side of the conference room, facing the center of the conference table); the virtual scene is a 3D virtual conference space, and the virtual object is the virtual avatar of remote participant B (located on the west side of the virtual conference table). First, the physical location information (position and orientation) of the microphone is obtained. Then, based on the position of the virtual object (the west coordinate of the virtual conference table), the physical coordinates are mapped to the virtual location information at a 1:1 ratio (the east coordinate of the virtual space is (3, 2, 1.5) meters, facing the center of the virtual conference table). Subsequently, a virtual sound source corresponding to the microphone (the sound source of virtual participant A) is created in the virtual scene. Finally, the speech collected by the microphone is rendered based on the virtual sound source—the volume is adjusted (attenuated to 55 decibels) according to the distance (6 meters) and angle (face-to-face) between the virtual sound source and the virtual object, the reverberation of the conference space is added, and the clarity of the human voice is preserved to obtain the rendered audio, thereby determining the virtual audio (the speech received by the virtual avatar of remote participant B from 6 meters to the east, facing him, so that B can perceive the same orientation and hearing in the virtual scene as offline).
[0075] In some embodiments, in a virtual reality game application scenario, the physical scene is the player's real living room, the physical sound source is the game controller held by the player (which emits a sound effect when a shot is triggered, and its physical location information is: coordinates (1, 0, 1.2) meters, located to the player's right, facing directly forward); the virtual scene is the battlefield map in the game, and the virtual object is the player's virtual character (located behind cover in the virtual battlefield). The physical location (position and orientation) of the controller is obtained through VR positioning devices. Then, based on the position of the virtual character (virtual cover coordinates), the physical coordinates are converted into virtual location information according to the ratio of 1 meter physical space = 5 meters virtual space (5 meters to the right of the virtual character, facing the direction of the virtual enemy). Subsequently, a corresponding virtual sound source (the sound source of the virtual props) is created in the virtual scene. Finally, the shooting sound effects emitted by the controller are rendered—according to the virtual location (close range, facing the enemy), the high-frequency sound is enhanced to have penetrating power, and the metallic reflection reverberation of the battlefield environment is added, so that the rendered audio presents a close-range, directional shooting sound to the right, and is determined to be the virtual audio received by the virtual object (player virtual character), allowing the player to obtain realistic auditory feedback on the location of props in VR.
[0076] In some embodiments, in the application scenario of remote virtual museum visits, the physical scene is a physical museum exhibition hall, and the physical sound source is an audio guide next to the display case (playing introductions of cultural relics, physical location information: coordinates (8, 5, 1.3) meters, located directly in front of the "Bronze Ware Display Case", facing the visitor passage); the virtual scene is a 3D replica of a virtual museum, and the virtual object is the virtual visitor's image (located in the virtual exhibition hall "Pottery Exhibition Area", 10 meters away from the Bronze Ware Display Case). First, the physical location of the audio guide is obtained through the exhibition hall positioning system. Then, based on the position of the virtual object, the physical coordinates are converted into virtual location information according to the principle of "aligning the physical exhibition hall with the virtual exhibition hall structure" (directly in front of the virtual bronze artifact display case, 10 meters away from the virtual object, facing the virtual visitor passage). Subsequently, a corresponding virtual sound source (virtual audio guide sound source) is created in the virtual scene. Finally, the audio of the audio guide is rendered—the volume is attenuated according to the 10-meter distance, the natural reverberation of the high space of the exhibition hall is added (extending the sound wave reflection time), and the mid-frequency clarity is slightly reduced because the virtual object is located to the side (not directly in front). The rendered audio is determined to be the virtual audio received by the virtual object, allowing remote users to perceive the audio guide sound coming from the direction of the bronze artifact display case 10 meters away, with a sense of spaciousness in the exhibition hall, simulating the auditory experience of a real visit.
[0077] In some embodiments, the physical scene is a factory workshop, and the physical sound source is a CNC machine tool (emitting mechanical sounds during operation; physical location information: coordinates (20, 15, 0.8) meters, located in the northeast corner of the workshop, facing due south); the virtual scene is a 3D digital twin model of the workshop, and the virtual object is the virtual perspective of a remote engineer (focused on the virtual model of the machine tool). First, the physical location of the machine tool is obtained through the workshop's IoT positioning device; then, based on the position of the virtual perspective (observed at zero distance from the virtual model of the machine tool), the physical coordinates are converted into virtual location information (the northeast corner of the virtual workshop, at zero distance from the virtual perspective, facing due south) using a 1:1 digital twin mapping; subsequently, a corresponding virtual sound source (the sound source of the virtual machine tool) is created in the virtual scene; finally, the mechanical sound of the machine tool is rendered—due to the close observation from the virtual perspective, the details of high-frequency mechanical friction sounds are preserved, distance attenuation is weakened, and hard reflection reverberation from the metal workshop is added to obtain the rendered audio, which is determined as the virtual audio received by the virtual object, allowing engineers to accurately determine the operating status of the machine tool (such as the location of abnormal friction sounds) through the virtual audio, thus achieving auditory assistance for remote operation and maintenance.
[0078] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device 500 for audio rendering provided in an embodiment of this application, wherein, Figure 2 The electronic device 500 shown can be Figure 1 Server 200 or terminal 400 in the middle, Figure 2The illustrated electronic device 500 includes at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.
[0079] Processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0080] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 430.
[0081] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0082] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0083] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0084] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, such as Bluetooth, WiFi, and Universal Serial Bus.
[0085] In some embodiments, the audio rendering apparatus provided in this application can be implemented in software. Figure 2 An audio rendering device 455 stored in memory 450 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: acquisition module 4551, conversion module 4552, creation module 4553, and rendering module 4554. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0086] In other embodiments, the audio rendering apparatus provided in this application can be implemented in hardware. As an example, the audio rendering apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio rendering method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0087] In some embodiments, the terminal or server can implement the audio rendering method provided in this application by running a computer program or computer-executable instructions. For example, the computer program can be a native program in the operating system (e.g., a dedicated audio rendering program) or a software module, such as an audio rendering module that can be embedded in any program (e.g., an instant messaging client, a photo album program, an electronic map client, a navigation client); or it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run. In summary, the above-mentioned computer program can be any form of application, module, or plugin.
[0088] The audio rendering method provided in this application will be described in conjunction with exemplary applications and implementations of the server or terminal provided in the embodiments of this application.
[0089] See Figure 3 , Figure 3 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 1 , will combine Figure 3Steps 101 to 105 are described below. The audio rendering method provided in this application embodiment can be implemented by the server or the terminal alone, or by the server and the terminal working together. The following description will take the implementation by the server alone as an example.
[0090] In step 101, the physical location information of the physical sound source in the physical scene is obtained.
[0091] In some embodiments, a physical sound source refers to an entity or vibration source that exists in real physical space and is capable of producing sound. Its core characteristic is that it generates sound waves that can be directly perceived by the human auditory system through physical vibrations (such as mechanical vibrations, air vibrations, etc.). It possesses a real physical form and location, and its sound propagation follows the laws of physical acoustics (such as air conduction, reflection, diffraction, etc.). Examples include human vocal organs (vocal cord vibration), musical instruments (string vibration, air column vibration), loudspeakers (diaphragm vibration), and natural sounds like wind / thunder (air vibration). The sound from a physical sound source can be collected as an electrical signal using devices such as microphones for subsequent processing (such as converting it into raw material for a virtual sound source).
[0092] In some embodiments, physical orientation information is a set of quantified information describing the spatial location and orientation of a physical sound source in a real physical scene. It is used to accurately locate the spatial attributes of the physical sound source and mainly includes the following two core elements: **Physical location coordinates:** These refer to the specific spatial position of the physical sound source in the physical scene coordinate system (such as a three-dimensional Cartesian coordinate system), usually represented by (x, y, z) coordinate values (units such as meters). For example, in a room scene, a physical sound source located at (3m, 2m, 1.5m) means that the sound source is 1.5 meters above the ground, 3 meters from the room's origin (such as a corner) along the x-axis, and 2 meters along the y-axis. **Physical direction vector:** This refers to the orientation or radiation direction of the physical sound source in the physical scene coordinate system, usually represented by a three-dimensional vector (dx, dy, dz) (which can be normalized to a unit vector). For example, a physical direction vector of (0, 1, 0) indicates that the main radiation direction of the physical sound source is along the positive y-axis (e.g., when the speaker faces this direction, the sound radiation intensity is highest in this direction). Physical orientation information directly reflects the spatial state of a physical sound source in real space. It is the basic data for converting the location and orientation of a sound source in the physical world to a virtual scene (such as generating virtual orientation information for a virtual sound source). It is widely used in technologies such as spatial audio conversion and virtual-real scene mapping.
[0093] In some embodiments, a physical scene refers to the real three-dimensional spatial environment in which a physical sound source exists. It is the actual location where the physical sound source exists and emits sound, possessing measurable spatial boundaries, geometric structure, and physical properties (such as size, material, and obstacle distribution). The physical scene is composed of real physical space and follows physical laws (e.g., the reflection and attenuation characteristics of sound propagation are determined by the scene's material and size). Its spatial coordinate system can be used to accurately locate the position (e.g., three-dimensional coordinates) and direction (e.g., direction vector) of the physical sound source. The physical scene provides the context for the existence of the physical sound source. The physical orientation information (e.g., position, direction) of the physical sound source needs to be defined based on the scene's coordinate system (for example, in the physical scene of a conference room, the position of the physical sound source can be described as 3 meters from the east wall, 2 meters from the south wall, and 1.5 meters from the ground). The parameters of the physical scene (e.g., spatial dimensions, obstacles) affect the natural propagation of the sound from the physical sound source. Simultaneously, its coordinate system serves as the benchmark for converting physical orientation information into virtual orientation information in the virtual scene, acting as a spatial reference connecting the physical and virtual worlds. The physical scene is the real spatial environment in which the physical sound source exists and is the fundamental carrier for defining physical orientation information.
[0094] As an example, the location of the speaker (physical sound source) in a conference room scenario is obtained. The physical scenario is a rectangular conference room measuring 8 meters long, 6 meters wide, and 3 meters high (a three-dimensional Cartesian coordinate system is established with the southwest corner of the room as the origin: x-axis east, y-axis north, z-axis upward). The physical sound source is the person speaking (vocal cord vibration is the physical sound source). The acquisition process is as follows: Position coordinate acquisition: Using an infrared positioning camera on the ceiling, the location marker worn by the speaker is captured, and their three-dimensional coordinates in the coordinate system are calculated as (x = 3.5 meters, y = 2.0 meters, z = 1.6 meters) – that is, 3.5 meters from the west wall, 2.0 meters from the south wall, and 1.6 meters from the ground. Direction vector acquisition: The speaker's facial orientation is captured by the camera, and combined with a head posture sensor, their facing direction is determined to be 10° east of north, converted to a unit direction vector of (dx = 0.98, dy = 0.17, dz = 0.0) (the positive x-axis direction is east, so dx is the principal component). Final physical location information: position (3.5, 2.0, 1.6), direction (0.98, 0.17, 0.0).
[0095] As an example, the location of an instrument (physical sound source) in a stage scene is obtained. The physical scene is a stage 12 meters wide and 8 meters deep (with the bottom left corner of the stage as the origin, the x-axis extends to the right along the width of the stage, the y-axis extends inward along the depth of the stage, and the z-axis is perpendicular to the stage and upward). The physical sound source is a violin on the stage (the vibration of its strings is the physical sound source). The acquisition process is as follows: Position coordinate acquisition: Using a UWB (Ultra-Wideband) positioning base station laid on the stage floor, the signal from the positioning tag installed on the bottom of the violin is received, and its coordinates are calculated as (x = 5.2 meters, y = 3.0 meters, z = 1.1 meters) – that is, 5.2 meters from the left side of the stage, 3.0 meters from the front edge of the stage, and 1.1 meters above the stage surface. Direction vector acquisition: Using an attitude sensor (such as an IMU) installed at the tail of the violin, its orientation is detected as the right side of the stage (positive x-axis direction), and the direction vector is (dx = 1.0, dy = 0.0, dz = 0.0). Final physical location information: position (5.2, 3.0, 1.1), direction (1.0, 0.0, 0.0).
[0096] In step 102, based on the virtual objects included in the virtual scene, the physical location information of the physical sound source is converted to obtain the virtual location information in the virtual scene corresponding to the physical location information.
[0097] In some embodiments, a virtual scene is a virtual scene displayed (or provided) by an application running on a terminal. This virtual scene can be a simulation of the real world, a semi-simulated / semi-fictional virtual environment, or a purely fictional virtual environment. The virtual scene can be any of a two-dimensional, 2.5-dimensional, or three-dimensional virtual scene; this application embodiment does not limit the dimension of the virtual scene. For example, a virtual scene may include sky, land, ocean, etc., and the land may include environmental elements such as deserts and cities. Users can control virtual objects to move within this virtual scene.
[0098] In some embodiments, a virtual object is an image of various people and objects that can be interacted with in a virtual scene, or an animated object in a virtual scene. This animated object can be a virtual character, virtual animal, cartoon character, etc., such as people, animals, plants, oil drums, walls, stones, etc., displayed in the virtual scene. The virtual object can be a virtual avatar representing the user within the virtual scene. A virtual scene can include multiple virtual objects, each with its own shape and volume, occupying a portion of the space in the virtual scene. Optionally, the virtual object can be a user character controlled through client operations, an artificial intelligence (AI) trained and set up for virtual scene battles, or a non-user character (NPC) set up for virtual scene interaction. A virtual object can be a virtual character engaging in adversarial interaction within the virtual scene. The number of virtual objects participating in interaction in the virtual scene can be pre-set or dynamically determined based on the number of clients joining the interaction.
[0099] In some embodiments, virtual orientation information refers to the spatial attribute data corresponding to the physical orientation information (such as position coordinates and direction vectors) of the physical sound source in the virtual scene after spatial transformation, used to characterize the spatial position and orientation of the virtual sound source in the virtual environment. Based on the physical orientation information of the physical scene, combined with the coordinate system rules of the virtual scene, the spatial distribution and mapping relationship of virtual objects (such as scaling, coordinate offset, direction calibration, etc.), spatial parameters adapted to the virtual scene are obtained through spatial transformation algorithms (such as coordinate transformation, rotation matrix operations, etc.). Virtual orientation information includes virtual position coordinates: the corresponding point of the physical sound source position in the coordinate system of the virtual scene (e.g., the position (3, 2, 1) in physical space, which becomes (6, 4, 2) in the virtual scene after 1:2 scaling). Virtual direction vector: the direction of the physical sound source after direction calibration, corresponding to the direction in the virtual scene (e.g., the physical direction is east, which is mapped to the positive x-axis direction in the virtual scene). Virtual location information serves as the spatial identity card of a virtual sound source in a virtual scene. It is the core basis for realizing spatial rendering of virtual audio (such as HRTF convolution and distance attenuation processing), ensuring that virtual objects can perceive the spatial position of the virtual sound source, and ultimately allowing users to obtain an immersive experience that conforms to the logic of the virtual scene.
[0100] In some embodiments, see Figure 4 , Figure 4 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 2 , Figure 3 Step 102 shown can be achieved through Figure 4Steps 1021 to 1024 shown are implemented.
[0101] In step 1021, the virtual scene coordinate system of the virtual scene is obtained, and the virtual scene coordinate system takes the center position of the virtual scene as the origin.
[0102] In some embodiments, the virtual scene coordinate system is a three-dimensional coordinate system used to locate the spatial positions of all elements (such as virtual sound sources, virtual objects, virtual environment components, etc.) in a virtual scene. Its core feature is that the origin is the center of the virtual scene, and the coordinate axes define the three spatial dimensions. The virtual scene coordinate system provides a unified quantitative standard for all spatial information within the virtual scene, serving as a spatial scale for converting physical location information to virtual location information. For example, the physical location of a sound source, after conversion, needs to be represented by specific coordinate values in this coordinate system to determine the position of its corresponding virtual sound source in the virtual scene. The relationship between the virtual scene coordinate system and the virtual scene: the coordinate system's scope must cover the entire virtual scene to ensure that the positions of all virtual elements can be accurately described, and the setting of its origin (the center of the virtual scene) makes spatial positioning more consistent with the overall layout logic of the virtual scene.
[0103] In step 1022, based on the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system, the virtual scene coordinate system is transformed to obtain an object coordinate system with the virtual object as the origin.
[0104] In some embodiments, the object coordinate system is a local three-dimensional coordinate system with the virtual object itself as the origin, used to describe the relative spatial relationship between the internal components of the virtual object or its associated elements (such as a virtual sound source). Its coordinate axis directions are typically bound to the virtual object's own posture (e.g., the x-axis along the front of the object, the y-axis along the right side of the object, and the z-axis perpendicular to the object upwards), and the coordinate values reflect the position or orientation relative to the virtual object itself (e.g., 0.5 meters in front of the virtual person and 0.3 meters to the right of the virtual device).
[0105] In some embodiments, the object coordinate system is a local three-dimensional coordinate system with the virtual object itself as the origin, used to describe the relative spatial relationship between the internal components of the virtual object or its associated elements (such as a virtual sound source). Its coordinate axis directions are typically bound to the virtual object's own posture (e.g., the x-axis along the front of the object, the y-axis along the right side of the object, and the z-axis perpendicular to the object upwards), and the coordinate values reflect the position or orientation relative to the virtual object itself (e.g., 0.5 meters in front of the virtual person and 0.3 meters to the right of the virtual device).
[0106] As an example, see Figure 11 , Figure 11 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 3 Based on the center position Ox of the virtual scene and the position coordinates Oy of the virtual object in the virtual scene coordinate system, the virtual scene coordinate system Txyz is transformed to obtain the object coordinate system Rxyz with the virtual object as the origin.
[0107] In some embodiments, the virtual scene coordinate system and the object coordinate system represent a mapping relationship between local and global coordinates, achieved through spatial transformation. The virtual scene coordinate system provides a global spatial reference for locating the absolute position of virtual objects within the entire virtual scene. The object coordinate system provides a local perspective, facilitating the description of the relative position of a virtual sound source to a virtual object (e.g., the virtual sound source is 3 meters to the left and in front of the virtual object), and serves as the direct basis for spatial audio rendering (e.g., calculating the relative position information between the virtual sound source and the virtual object). The object coordinate system is a local coordinate system obtained by translating the virtual scene coordinate system from the origin and rotating it. The two coordinate systems achieve spatial mapping through the position and posture of the virtual object, serving global positioning and local relative relationship calculations, respectively.
[0108] In some embodiments, step 1022 above can be implemented as follows: the difference vector between the coordinates of the center position of the virtual scene and the position coordinates is determined as the offset vector between the virtual scene coordinate system and the object coordinate system; the virtual scene coordinate system is translated according to the offset vector to obtain the object coordinate system with the virtual object as the origin.
[0109] In some embodiments, the offset vector is a spatial vector from the origin of the virtual scene coordinate system to the origin of the object coordinate system, obtained by the server through coordinate difference calculation. The expression for the offset vector can be:
[0110]
[0111] in, Used to indicate the offset vector, O obj The coordinates used to indicate the center position of the virtual scene, O scene Used to indicate location coordinates.
[0112] As an example, the difference vector describes the spatial displacement from the center of the virtual scene (global origin) to the position of the virtual object (local origin), including direction and distance information. For example, if the center of the virtual scene is (0, 0, 0) and the position of the virtual object is (3, 2, 1), then the offset vector is (3, 2, 1), indicating that the origin of the object's coordinate system is located at x-axis +3, y-axis +2, z-axis +1 in the virtual scene coordinate system.
[0113] In some embodiments, the virtual scene coordinate system is translated according to the offset vector to obtain an object coordinate system with the virtual object as the origin. Based on the offset vector, the virtual scene coordinate system is translated to obtain the object coordinate system. The origin of the virtual scene coordinate system, Osce ne, is moved along the offset vector d until it coincides with the position Oobj of the virtual object—that is, the origin of the object coordinate system, Olocal = Oobj. The translation transformation only changes the position of the origin; the coordinate axes of the object coordinate system are completely identical to those of the virtual scene coordinate system (e.g., the x-axis remains east-west and the y-axis remains north-south), without any rotation or scaling. For any point P in the virtual scene (coordinates (xp, yp, zp)), its coordinates P′ in the object coordinate system are calculated by the server as follows: P′ = (xp - xo, yp - yo, zp - zo) (essentially, the global coordinates of the object coordinate system's origin are subtracted from the global coordinates, thus eliminating the influence of the offset vector).
[0114] In some embodiments, the object coordinate system is centered on the virtual object. When calculating the position of the virtual sound source relative to the virtual object, the server can directly use the relative coordinates in the object coordinate system (e.g., (2, 0, 0) in the object coordinate system means the virtual sound source is 2 units directly in front of the virtual object), without relying on global coordinates, thus reducing computational complexity. When the virtual object moves (e.g., when a user walks in VR), the server recalculates the offset vector (based on the new Oobj) in real time and updates the origin of the object coordinate system, ensuring that the local coordinates are always bound to the real-time position of the virtual object. If there are multiple virtual objects in the virtual scene, the server calculates the offset vector and object coordinate system independently for each object, so that the local spatial relationships of each object (e.g., the relative position with other virtual elements) can be described independently, avoiding confusion in the global coordinates.
[0115] As an example, the virtual scene coordinate system has the origin Oscene = (0, 0, 0) (the center of the scene). The position of virtual object A is OobjA = (5, 3, 0) (5 meters east and 3 meters north of the scene center). The server calculates the offset vector: dA = (5-0, 3-0, 0-0) = 5, 3, 0. After translation, the origin of object coordinate system A is (5, 3, 0). The coordinates of a virtual sound source in the virtual scene coordinate system are (7, 3, 0). Therefore, its coordinates in object coordinate system A are (7-5, 3-3, 0-0) = (2, 0, 0), which is 2 meters directly east of virtual object A.
[0116] In this way, the difference vector between the coordinates of the virtual scene's center and the coordinates of the virtual object in the virtual scene's coordinate system is determined as the offset vector. Based on this offset vector, the virtual scene coordinate system is translated to obtain the object's coordinate system. Through explicit offset vector calculation and coordinate system translation, a local spatial reference system centered on the virtual object can be quickly established in the virtual scene. This simplifies the description and calculation of the relative positional relationships between the virtual object and other virtual elements (such as virtual sound sources), avoiding the complex coordinate calculations that might occur when directly using the global coordinate system. This significantly reduces the computational load on the server when handling spatial interactions (such as virtual sound source orientation determination and collision detection), improving operational efficiency. It ensures a precise association between the object's coordinate system and the virtual scene's coordinate system. When a virtual object moves in the virtual scene, the server only needs to update the object's coordinate system in real time by updating the offset vector, guaranteeing the dynamic consistency between local and global coordinates. This allows spatial logic based on the object's coordinate system (such as spatial rendering of virtual audio and behavioral interactions of virtual objects) to adjust in real time as the virtual object's position changes, enhancing the accuracy and real-time performance of spatial relationships in the virtual scene.
[0117] In step 1023, the transformation matrix between the object coordinate system and the physical scene coordinate system is determined.
[0118] In some embodiments, the transformation matrix is a mathematical matrix (usually a 4x4 homogeneous matrix) used to describe the spatial mapping relationship between the object coordinate system and the physical scene coordinate system. The transformation of coordinates of any point in the two coordinate systems can be achieved through a single matrix operation. Its core function is to integrate spatial transformation parameters such as translation, rotation, and scaling to transform coordinates in one coordinate system (such as relative coordinates in the object coordinate system) into coordinates in another coordinate system (such as absolute coordinates in the physical scene coordinate system).
[0119] As an example, see Figure 12 , Figure 12 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 4 Determine the transformation matrix between the object coordinate system Rxyz and the physical scene coordinate system Wxyz.
[0120] In some embodiments, in the association between the object coordinate system and the physical scene coordinate system, the elements of the transformation matrix are determined by the spatial relationship between the two: if only the positional difference exists, the matrix includes translation parameters (corresponding to the offset vectors of the origins of the two coordinate systems); if the directional difference exists, the matrix includes rotation parameters (corresponding to the angle between the coordinate axes of the two coordinate systems); if the size difference exists, the matrix includes scaling parameters (corresponding to the unit scale of the two coordinate systems). This matrix enables rapid coordinate transformation (e.g., converting the relative position of a virtual sound source perceived by a virtual object in the object coordinate system to its absolute position in the physical scene coordinate system), and is a core mathematical tool for connecting virtual space and physical space and ensuring the consistency of their spatial logic.
[0121] In some embodiments, the transformation matrix includes a position transformation matrix and a direction transformation matrix. Step 1023 can be implemented as follows: determine the position transformation matrix based on the position coordinates of the origin of the object coordinate system in the physical scene coordinate system; determine the unit direction vector of each coordinate axis in the physical scene coordinate system for each coordinate axis of the object coordinate system; and construct a transformation matrix between the object coordinate system and the physical scene coordinate system based on the unit direction vector of each coordinate axis.
[0122] In some embodiments, the transformation matrix between the object coordinate system and the physical scene coordinate system is a 4x4 homogeneous matrix (compatible with translation and rotation operations in three-dimensional space). Its structure can be divided into two parts: position transformation matrix: responsible for describing the translation relationship between the origins of the two coordinate systems, i.e., the position offset of the origin of the object coordinate system in the physical scene coordinate system; direction transformation matrix: responsible for describing the rotation relationship between the coordinate axes of the two coordinate systems (i.e., the directions in which the x, y, and z axes of the object coordinate system point in the physical scene coordinate system).
[0123] In some embodiments, the core of the position transformation matrix is the absolute position of the object coordinate system origin in the physical scene coordinate system, used to describe the origin offset between the two coordinate systems. Let the origin of the object coordinate system be Ololocal, and its position coordinates in the physical scene coordinate system be (x0, y0, z0) (i.e., the translation vector from the physical scene origin to Ololocal). In a homogeneous coordinate system, the position transformation matrix is a 4x4 matrix, and the translation parameters only affect the last column (the first three rows correspond to the translation amounts along the x, y, and z axes, and the fourth row indicates the homogeneous coordinates). When the translation parameter matrix acts alone, it can translate a point (x, y, z) in the object coordinate system to (x+x0, y+y0, z+z0) in the physical scene coordinate system, handling only the position offset without changing the direction.
[0124] In some embodiments, for determining the unit direction vectors of each coordinate axis of the object coordinate system in the physical scene coordinate system, the core of direction transformation is to clarify the direction of the x, y, and z axes of the object coordinate system in the physical scene coordinate system. This requires defining a unit direction vector (a vector of length 1 describing the direction) for each coordinate axis. Let the unit direction vector of the x-axis of the object coordinate system in the physical scene coordinate system be ux = (ux1, ux2, ux3) (for example, if the object's x-axis points to the northeast direction of the physical scene, the vector might be (0.707, 0.707, 0); the unit direction vector of the y-axis be uy = (uy1, uy2, uy3); and the unit direction vector of the z-axis be uz = (uz1, uz2, uz3). The x, y, and z axes of the object coordinate system are mutually perpendicular (right-handed coordinate system), therefore the three direction vectors must satisfy orthogonality (the dot product of any pairwise vectors is 0) and the right-hand rule (uz = ux × uy). To ensure logical consistency in direction, the direction vectors are constructed into a direction transformation matrix. This matrix is a 3x3 rotation matrix whose columns (or rows, depending on the coordinate system convention) consist of the unit direction vectors of the three coordinate axes of the object's coordinate system, describing the rotational relationship between the two coordinate systems. When used alone, the direction transformation matrix converts vectors along the object's x, y, and z axes into corresponding direction vectors in the physical scene coordinate system. For example, the vector (1, 0, 0) along the x-axis in the object's coordinate system becomes ux in the physical scene after matrix transformation, achieving a rotational mapping of direction.
[0125] As an example, the complete transformation matrix is a 4x4 homogeneous matrix that integrates the orientation transformation matrix (3x3 rotation part) and the position transformation matrix (translation part), as follows:
[0126]
[0127] In this way, the transformation matrix is decomposed into a position transformation matrix and an orientation transformation matrix. The position transformation matrix is determined by the position coordinates of the origin of the object coordinate system in the physical scene coordinate system. At the same time, the overall transformation matrix is constructed based on the unit direction vectors of each coordinate axis of the object coordinate system in the physical scene coordinate system. The decompositional construction method decomposes the complex spatial transformation into two independent dimensions: position offset and orientation rotation. This reduces the complexity of matrix calculation, enabling the server or processing system to complete the transformation between coordinate systems more efficiently and reducing the consumption of computing resources. By clearly defining the position of the origin of the object's coordinate system and the direction vectors of each coordinate axis, the spatial relationship between the two coordinate systems can be accurately captured, ensuring the accuracy of coordinate transformation and avoiding the accumulation of errors caused by the coupling of the overall transformation matrix parameters. This provides a reliable mathematical foundation for the accurate mapping between virtual and physical spaces (such as matching the position of a virtual sound source with the physical scene). It has good scalability and real-time performance. When the position of the origin or the direction of the coordinate axes of the object's coordinate system changes (such as when a virtual object moves or rotates), the overall transformation matrix can be quickly reconstructed simply by updating the position transformation matrix or the direction vector, without having to recalculate the entire matrix. This ensures the efficiency and consistency of coordinate system transformation in dynamic scenes (such as VR interaction and real-time simulation), laying a solid foundation for improving the immersion and interaction accuracy of virtual scenes.
[0128] In step 1024, the physical orientation information of the physical sound source is transformed based on the transformation matrix to obtain the virtual orientation information.
[0129] In some embodiments, the physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system and the physical direction vector of the physical sound source in the physical scene coordinate system. The transformation matrix includes a position transformation matrix and a direction transformation matrix. The virtual orientation information includes virtual position coordinates and a virtual direction vector.
[0130] In some embodiments, step 1024 above can be implemented as follows: the product of the physical position coordinates and the position transformation matrix is used to determine the virtual position coordinates; the product of the physical direction vector and the direction transformation matrix is used to determine the virtual direction vector.
[0131] In some embodiments, determining the virtual position coordinates by multiplying the physical position coordinates and the position transformation matrix, and determining the virtual direction vector by multiplying the physical direction vector and the direction transformation matrix, are core steps in achieving precise mapping from physical space parameters to virtual space parameters through matrix operations. Essentially, this involves using the mathematical expression of position translation and direction rotation using matrices to complete the parameter transformation between the two coordinate systems. The physical position coordinates refer to the three-dimensional absolute coordinates of the physical sound source in the physical scene coordinate system (denoted as Pphys = (xp, yp, zp)), while the position transformation matrix is a 4x4 homogeneous matrix (denoted as Mpos, containing the position of the object coordinate system origin in the physical scene (x0, y0, z0)) describing the offset relationship between the object coordinate system and the origin of the physical scene coordinate system. The product of the two essentially transforms the absolute position in the physical scene into a relative position in the object coordinate system (i.e., the virtual position coordinates Pvirt) through translation transformation.
[0132] In some embodiments, the physical direction vector refers to the unit direction vector of the physical sound source in the physical scene coordinate system (denoted as dphys = (dpx, dpy, dpz), describing the orientation of the sound source), while the direction transformation matrix is a 3x3 rotation matrix (denoted as Mdir, composed of the direction vectors of each coordinate axis of the object coordinate system in the physical scene) describing the rotation relationship between the coordinate axes of the object coordinate system and the physical scene coordinate system. The product of the two essentially transforms the direction vector in the physical scene into a direction vector in the object coordinate system (i.e., the virtual direction vector dvirt) through rotation. By multiplying the physical position coordinates with the position transformation matrix, a translation transformation from absolute position to relative position is achieved; by multiplying the physical direction vector with the direction transformation matrix, a rotation transformation from physical direction to virtual direction is achieved. These two steps, targeting the position and direction dimensions respectively, ensure through the mathematical rigor of matrix operations that the parameters of the physical space can be accurately mapped to the virtual object coordinate system. This provides an accurate spatial parameter foundation for subsequent spatial interactions in the virtual scene (such as HRTF rendering of virtual audio and the virtual object's determination of the physical sound source's orientation), ultimately improving the realism and interactive accuracy of the virtual scene.
[0133] Thus, by multiplying the physical position coordinates by the position transformation matrix to obtain the virtual position coordinates, and by multiplying the physical direction vector by the direction transformation matrix to obtain the virtual direction vector, matrix operations are used to convert physical spatial parameters to virtual spatial parameters. Leveraging the specific functions of the position and direction transformation matrices, position translation and direction rotation—two types of spatial transformations—can be accurately separated and processed, avoiding conversion errors caused by parameter confusion. This ensures that the virtual position coordinates and virtual direction vectors accurately reflect the relative position and orientation of the physical sound source in the virtual scene, providing a reliable spatial data foundation for subsequent processing such as virtual audio rendering and virtual object interaction. The matrix-based conversion method is mathematically rigorous and efficient. When the spatial state of the physical scene or virtual object changes (such as the movement of the physical sound source or the rotation of the virtual object), only the corresponding transformation matrix needs to be updated to quickly recalculate the virtual parameters without reconstructing the entire conversion logic. This significantly improves the real-time performance and flexibility of spatial parameter mapping in dynamic scenes, thereby enhancing the immersion and interactive response speed of the virtual scene.
[0134] Thus, the initial coordinate system with the center of the virtual scene as the origin provides a unified benchmark for the entire virtual space, ensuring the integrity of spatial positioning. Transforming the coordinate system into an object coordinate system with the virtual object as the origin allows for a more direct focus on the virtual object's perspective, better aligning with its perceptual logic. This makes subsequent orientation calculations more intuitive in reflecting the relative relationship (such as distance and angle) between the virtual sound source and the virtual object. By determining the transformation matrix between the object coordinate system and the physical scene coordinate system, a precise mathematical connection between the physical and virtual spaces is established, avoiding orientation deviations caused by coordinate system differences. This ensures that the actual orientation (such as position and direction) of the physical sound source can be accurately mapped into the virtual scene. Ultimately, the virtual orientation information obtained based on this transformation matrix allows the position of the virtual sound source in the virtual scene to form a logically consistent correspondence with the position of the physical sound source in the physical scene. This provides a precise spatial basis for subsequent audio rendering, making the virtual audio received by the virtual object more closely resemble a real auditory experience in terms of spatial attributes such as orientation and distance, effectively enhancing the spatial realism of the virtual scene and the user's immersion.
[0135] In step 103, based on the virtual orientation information, a virtual sound source corresponding to the physical sound source is created in the virtual scene.
[0136] In some embodiments, a virtual sound source refers to an acoustic abstraction with spatial attributes defined in a virtual scene (such as a VR / AR environment or digital 3D space). Its core feature is simulating the spatial propagation characteristics of sound through digital signal processing technology, allowing listeners to perceive that the sound originates from a specific location in the virtual space. A virtual sound source does not have a real physical form; it is a digital object carrying virtual orientation and acoustic parameters (such as directivity and distance attenuation characteristics) used to construct a virtual sound field. The virtual sound source combines the sound signal (or digital audio signal) of the physical sound source with the spatial information of the virtual scene, simulating the propagation effect of sound in the virtual space through audio rendering (such as HRTF convolution and reverberation processing), creating the auditory illusion for listeners that the sound originates from a virtual location. The parameters of the virtual sound source (such as 3D coordinates and motion state) must match the virtual scene; it serves as a bridge connecting physical sound and the virtual sound field, and is widely used in virtual reality, 3D audio, game sound effects, and other fields. A physical sound source is a real sound producer that follows physical laws; a virtual sound source is a digital acoustic model used to simulate sound localization in virtual space, relying on signal processing technology to achieve the perceptual effect.
[0137] In some embodiments, the essence is to convert the quantized orientation parameters of the physical sound source into the quantized orientation parameters of the virtual sound source through a preset mapping rule, ensuring that the two are completely consistent in spatial orientation perception. The specific mapping rule must meet the following quantization requirements: accurate mapping of coordinate parameters: if the virtual scene is mapped to the real space in a 1:1 ratio, then the three-dimensional coordinates (X2, Y2, Z2) of the virtual sound source must be completely equal to the three-dimensional absolute coordinates (X1, Y1, Z1) of the physical sound source, that is, X1 = X2, Y1 = Y2, Z1 = Z1, with the error controlled within ±0.01 virtual units; if it is a proportional mapping (let the proportional coefficient be k, where k is a positive number), then the virtual coordinates must satisfy X2 = X1 × k, Y2 = Y1 × k, Z2 = Z1 × k, ensuring that the relative coordinate relationship is without deviation (e.g., if the physical sound source is 2 meters to the right and 0.5 meters above in the real space, and the proportional coefficient k = 2, then the virtual coordinates are 4 virtual units to the right and 1 virtual unit above). Complete consistency of azimuth parameters: The horizontal azimuth of the virtual sound source relative to the listener's virtual character must be exactly equal to the horizontal azimuth of the physical sound source relative to the real listener (within ±0.1° error); the virtual vertical azimuth must be exactly equal to the real vertical azimuth (within ±0.1° error). For example, if the physical sound source has a horizontal azimuth of 45° and a vertical azimuth of 30°, then the virtual sound source must also have a horizontal azimuth of 45° and a vertical azimuth of 30°, ensuring that the direction of the sound source perceived by the listener in the virtual scene is completely consistent with the real scene. Proportional matching of distance parameters: The distance between the virtual sound source and the listener's virtual character must satisfy a proportional mapping relationship with the distance between the physical sound source and the real listener (the proportional coefficient must be exactly consistent with the coordinate mapping coefficient k). If the real distance is d1 and the virtual distance is d2, then d2 = d1 × k, with the error controlled within ±0.01 virtual units. For example, if the real distance is 2.5 meters and k = 1.5, then the virtual distance needs to be 3.75 virtual units to ensure that the distance of the sound source perceived by the listener through virtual audio matches the real scene proportionally (or is consistent 1:1).
[0138] As an example, see Figure 10 , Figure 10 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 2 ,like Figure 10 In the virtual scene shown, based on the virtual object 21 included in the virtual scene, the physical location information of the physical sound source is converted to obtain the virtual location information corresponding to the physical location information in the virtual scene. Based on the virtual location information, a virtual sound source 22 corresponding to the physical sound source is created in the virtual scene.
[0139] In some embodiments, the virtual orientation information includes virtual position coordinates and virtual direction vectors, see [link to relevant documentation]. Figure 5 , Figure 5 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 3 , Figure 3 Step 103 shown can be achieved through Figure 3 Steps 1031 to 1032 shown are implemented.
[0140] In step 1031, an initial virtual sound source corresponding to the physical sound source is created at the virtual position coordinates in the virtual scene.
[0141] In some embodiments, within a virtual scene, an initial virtual sound source corresponding to a physical sound source is created based on virtual position coordinates calculated using physical position coordinates and a position transformation matrix. This initial virtual sound source is a digital object in the virtual scene used to associate the acoustic properties of the physical sound source. Its spatial position in the virtual scene is strictly limited to the calculated virtual position coordinates, ensuring that the relative spatial relationship between the physical sound source and the physical sound source in the physical scene is accurately reproduced in the virtual scene. For example, if a physical sound source is located 2 meters in front of an object in the physical scene, and its virtual position coordinates in the virtual scene are (5, 3, 1) obtained through coordinate transformation, then the initial virtual sound source will be created at that coordinate point. This allows interactive objects in the virtual scene (such as virtual humans or user perspectives) to perceive that the sound originates from this virtual location, laying the foundation for subsequent spatial audio rendering (such as simulating sound direction and distance attenuation) using virtual direction vectors, thereby achieving spatial association and acoustic matching between the physical sound source and the virtual scene.
[0142] As an example, see Figure 13 , Figure 13 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 5 ,exist Figure 13 At the virtual position coordinates (x1, y1, z1) in the virtual scene shown, an initial virtual sound source 23 corresponding to the physical sound source is created.
[0143] In step 1032, the initial virtual sound source is adjusted based on the virtual direction vector to obtain the virtual sound source.
[0144] In some embodiments, adjusting the initial virtual sound source based on the virtual direction vector to obtain the final virtual sound source essentially involves calibrating the spatial orientation of the virtual sound source so that its acoustic radiation characteristics in the virtual scene are consistent with the actual orientation of the physical sound source in the physical scene. Although the initial virtual sound source has been anchored in space using virtual position coordinates, it has not yet reflected the directional attributes of the physical sound source. The virtual direction vector is precisely the mapping result of the physical sound source's orientation in the virtual scene (derived from the physical direction vector through a direction transformation matrix). Setting the radiation direction of the initial virtual sound source to the direction pointed to by the virtual direction vector makes the sound energy distribution of the virtual sound source conform to that directional characteristic—for example, if the virtual direction vector points to the northeast of the virtual scene, the adjusted virtual sound source will have the highest sound radiation intensity in the northeast direction, while other directions will attenuate as the angle increases, just like the effect of a physical sound source emitting sound towards the northeast in a real scene. This adjustment ensures that the virtual sound source not only corresponds to the physical sound source in location, but also accurately replicates its directional characteristics. This provides accurate directional parameters for subsequent spatial audio rendering (such as simulating the differences in human hearing perception of sound from different directions using HRTF technology). Ultimately, this allows the direction of sound heard by the user in the virtual scene to be consistent with the direction of real sound in the physical scene, enhancing the realism and spatial consistency of the virtual auditory experience.
[0145] In some embodiments, step 1032 above can be implemented as follows: the orientation of the initial sound source is adjusted to be consistent with the virtual direction vector to obtain a candidate virtual sound source; the acoustic feature parameters of the physical sound source are obtained, and the acoustic feature parameters of the physical sound source are adjusted based on the virtual scene to obtain the adjusted acoustic feature parameters; the acoustic feature parameters of the candidate virtual sound source are set as the adjusted acoustic feature parameters to obtain the virtual sound source.
[0146] In some embodiments, candidate virtual sound sources are obtained by adjusting the orientation of the initial virtual sound source to match the virtual direction vector. Although the initial virtual sound source has its spatial position anchored based on virtual coordinates, it does not yet reflect the directional attributes of the physical sound source. The virtual direction vector is a precise mapping of the physical sound source's orientation in the virtual scene (derived from the physical direction vector through a direction transformation matrix). Adjusting the orientation means ensuring that the sound radiation direction of the candidate virtual sound source matches the virtual direction vector—for example, if the virtual direction vector points directly forward (positive z-axis direction) in the virtual scene, the sound from the candidate virtual sound source will primarily propagate along that direction, simulating the spatial characteristics of a physical sound source emitting sound directly forward, providing basic directional parameters for subsequent acoustic rendering. Acoustic characteristic parameters of the physical sound source (such as volume, pitch, timbre, Doppler effect coefficients, etc.) are obtained and adjusted in conjunction with the environmental characteristics of the virtual scene (such as the size, material, and reverberation parameters of the virtual scene) to obtain acoustic characteristic parameters adapted to the virtual scene. For example, the volume attenuation characteristics of a physical sound source in an open physical space may differ from the attenuation pattern in a closed room in a virtual scene. The volume parameters need to be adjusted according to the acoustic model of the virtual scene (e.g., an attenuation index of 1.5 for indoor environments). Alternatively, if specific frequencies of environmental noise exist in the virtual scene, the timbre spectrum envelope of the physical sound source needs to be fine-tuned to match the virtual sound field. Replacing the acoustic characteristic parameters of the candidate virtual sound source with the adjusted parameters yields the final virtual sound source. This step ensures that the virtual sound source not only corresponds to the physical sound source in spatial location and direction, but also that its acoustic characteristics are fully adapted to the environmental logic of the virtual scene. This guarantees that the sound heard by the user in the virtual scene not only reproduces the essential characteristics of the physical sound source but also conforms to the acoustic laws of the virtual environment (e.g., the sound heard in a virtual cave has obvious reverberation). This achieves a precise and natural mapping of the physical sound source to the virtual scene, enhancing the immersion and realism of the virtual auditory experience.
[0147] As an example, see Figure 14 , Figure 14 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 6 ,exist Figure 14 The orientation of the initial sound source 23 is adjusted to be consistent with the virtual direction vector, as shown, to obtain the candidate virtual sound source 24.
[0148] As an example, suppose there's a physical speaker (physical sound source) playing music, facing east in a physical scene; the corresponding virtual scene is a virtual concert hall. An initial virtual sound source has been created in the virtual concert hall at a virtual location corresponding to the physical speaker's position, but it doesn't yet have a defined orientation. Based on the virtual direction vector (the mapping of the physical speaker's eastward direction in the virtual scene, assuming it corresponds to the right side of the stage in the virtual concert hall), we adjust the initial virtual sound source's orientation to point towards the right side of the virtual stage. This gives us a candidate virtual sound source—it now behaves like a physical speaker, radiating sound primarily in a specific direction. For example, in the virtual concert hall, virtual listeners on the right side of the stage will perceive the direction of this sound source more clearly than those on the left. We obtain the acoustic characteristics of the physical speaker: for example, a medium volume, a bright timbre, and no significant reverberation. However, the virtual scene is a concert hall, and a real concert hall would produce natural reverberation, and the larger space might require a slightly higher volume to ensure a comfortable listening experience. Therefore, based on the environmental characteristics of a virtual concert hall, we adjusted these parameters: appropriately increasing the volume and adding a slight reverberation effect to the timbre to better match the acoustic atmosphere of a concert hall, resulting in the adjusted acoustic characteristic parameters. Replacing the original acoustic parameters of the candidate virtual sound source with these adjusted parameters yields the final virtual sound source. This virtual sound source not only emits sound in a specific direction like a physical speaker, but its volume and timbre are also adapted to the environment of a virtual concert hall, allowing users to hear sound in the virtual scene that both replicates the essential characteristics of a physical speaker and conforms to the listening logic of a virtual concert hall.
[0149] In this way, the orientation of the initial sound source is adjusted to match the virtual direction vector to obtain candidate virtual sound sources. Then, the acoustic characteristic parameters of the physical sound source are adjusted in conjunction with the virtual scene and assigned to the candidate virtual sound sources to obtain the final virtual sound source. By adjusting the orientation, it is ensured that the orientation of the virtual sound source in the virtual scene accurately corresponds to the actual orientation of the physical sound source in the physical scene. This ensures that the spatial directivity of sound in the virtual environment is consistent with that in the real world, avoiding the destruction of immersion caused by directional misalignment. At the same time, the targeted adjustment of acoustic characteristic parameters based on the virtual scene allows the volume, timbre, reverberation, and other characteristics of the virtual sound source to adapt to the environmental characteristics of the virtual scene (such as the spaciousness of a virtual hall or the closedness of a virtual room). This preserves the essential acoustic characteristics of the physical sound source while conforming to the acoustic logic of the virtual scene. This step-by-step optimization method ultimately achieves a natural and accurate mapping of the physical sound source to the virtual scene, making the virtual sound highly integrated with the virtual environment in terms of spatial orientation and acoustic characteristics. This greatly enhances the user's auditory immersion and the realism of the experience in the virtual scene.
[0150] In this way, an initial virtual sound source corresponding to the physical sound source is created at the virtual location coordinates in the virtual scene, and then adjusted based on the virtual direction vector to obtain the final virtual sound source. By creating the initial virtual sound source at the precisely calculated virtual location coordinates, it is ensured that the spatial position of the physical sound source in the physical scene and its mapped position in the virtual scene form a strict correspondence, laying the foundation for the spatial positioning of virtual sound. Furthermore, the adjustment based on the virtual direction vector ensures that the orientation of the virtual sound source is consistent with the actual sound emission direction of the physical sound source in the virtual scene, making the radiation characteristics of the sound conform to the logic of real space. This two-step construction method not only achieves accurate coordinate mapping from physical space to virtual space, but also ensures the true restoration of the sound direction characteristics. Ultimately, it allows the virtual sound source to be accurately positioned and correctly transmit directional information in the virtual scene, providing users with a virtual auditory experience consistent with the acoustic laws of the physical scene, effectively enhancing the immersion and spatial realism of the virtual scene.
[0151] In step 104, based on the virtual sound source, the sound signal generated by the physical sound source is rendered to obtain the rendered audio.
[0152] In some embodiments, audio rendering refers to the process of processing the original sound signal generated by the physical sound source using a virtual sound source as a spatial reference, combined with the acoustic environmental characteristics of the virtual scene, to ultimately generate rendered audio that conforms to the listening logic of the virtual scene. This process is mainly achieved by adjusting acoustic characteristics. Specifically, the spatial attributes of the virtual sound source in the virtual scene, such as its position and orientation, serve as the core basis for audio rendering: for example, based on the distance between the virtual sound source and the listener (or virtual object) in the virtual scene, the volume attenuation of the sound is adjusted (the further away, the weaker the volume), simulating the characteristics of sound changing with distance in real space; based on the orientation of the virtual sound source and the relative angle with the listener, the volume ratio of the left and right channels is adjusted, and subtle phase differences are added, so that the listener can perceive the direction of the sound (e.g., when the virtual sound source is on the left, the volume of the left channel is stronger); at the same time, combined with the environmental attributes of the virtual scene (e.g., whether the virtual scene is a spacious hall or a small room), the reverberation effect of the sound is adjusted (the reverberation is stronger and lasts longer in a hall), echo delay, etc., so that the sound carries the spatial imprint unique to the virtual environment. By adjusting these acoustic properties, the original sound signals from the physical sound source are given spatial attributes of the virtual scene. The final rendered audio allows users to hear sounds in the virtual scene that retain the essential characteristics of the physical sound source while fully integrating into the acoustic logic of the virtual environment, creating an immersive auditory experience.
[0153] In some embodiments, the above-described audio rendering process, based on the virtual sound source, renders the sound signal generated by the physical sound source to obtain the rendered audio. Several parallel technical solutions are presented below. These parallel solutions are independent of each other and possess their own complete technical architecture, each capable of independently solving the technical problem addressed by this invention and achieving the corresponding technical effect. These parallel solutions can be arbitrarily combined according to actual application scenarios, needs, and technical conditions, and this combination does not limit the scope of protection claimed in this application. The scope of protection of this application is clearly defined by the claims, covering the set of technical features presented when each parallel solution is implemented individually, as well as the set of technical solutions formed by combining each parallel solution in any reasonable manner, provided it complies with the relevant provisions of the Patent Law. Whether it is the implementation of a single solution or the synergistic use of multiple solutions, as long as it falls within the boundaries defined by the claims, it is protected by this application. The specific implementation process of audio rendering is described in detail below.
[0154] In some embodiments, the number of physical sound sources is at least one, and the virtual sound sources correspond one-to-one with the physical sound sources. Step 104 above can be implemented in the following way: for each virtual sound source, based on the virtual sound source, perform audio rendering on the sound signal generated by the corresponding physical sound source to obtain the rendered audio corresponding to the virtual sound source.
[0155] In some embodiments, when there are one or more physical sound sources, each physical sound source generates a corresponding virtual sound source in the virtual scene, meaning there is a one-to-one correspondence between virtual and physical sound sources. For each virtual sound source, the sound signal emitted by its corresponding physical sound source is processed using audio rendering, resulting in a unique rendered audio for that virtual sound source. For example, if there are two physical sound sources—a speaker emitting sound to the east and a microphone emitting sound to the south—then there will be two corresponding virtual sound sources in the virtual scene, one for the speaker and one for the microphone. Then, for the virtual sound source corresponding to the speaker, the sound signal emitted by the speaker is rendered using its position and orientation information in the virtual scene, resulting in rendered audio for the speaker. Simultaneously, for the virtual sound source corresponding to the microphone, the sound signal collected by the microphone is rendered, resulting in rendered audio for the microphone. In this way, the sound from each physical sound source can be processed independently in the virtual scene, preserving its individual characteristics while adapting to the virtual scene environment, ensuring that multiple sounds can coexist naturally and without interference in the virtual scene.
[0156] Thus, when there is at least one physical sound source and each virtual sound source corresponds one-to-one with a physical sound source, and audio rendering is performed on the corresponding physical sound source's sound signal for each virtual sound source, the one-to-one correspondence ensures that the sound characteristics of each physical sound source can be independently and accurately mapped in the virtual scene. This avoids auditory confusion caused by the mixing of multiple sound source signals. For example, the voices of people speaking and the sounds of musical instruments from different physical locations can still maintain their own spatial positioning and acoustic characteristics in the virtual scene. Rendering each virtual sound source separately allows for personalized processing based on its specific location, orientation, and virtual environment characteristics (such as distance from the audience and the material of the virtual space). For example, distant sound sources can be made weaker and have more reverberation, while nearby sound sources can be made clearer and have a stronger sense of direction. This allows multiple rendered audios to form a layered sound field in the virtual scene that conforms to the logic of real space. Ultimately, this allows users to accurately distinguish the location and characteristics of different physical sound sources in the virtual scene, greatly enhancing the realism and immersion of the virtual auditory experience in multi-sound-source scenarios.
[0157] In some embodiments, see Figure 6 , Figure 6 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 4 , Figure 6 Step 104 shown can be achieved through Figure 3 Steps 1041A to 1042A shown are implemented.
[0158] In step 1041A, the relative orientation information between the virtual sound source and the virtual object is determined based on the virtual orientation information in the virtual scene corresponding to the physical orientation information.
[0159] In some embodiments, relative orientation information refers to quantitative data used to describe the spatial relationship between a virtual sound source and a virtual object, calculated based on the virtual orientation information (position, direction) of the virtual sound source and the spatial attributes of the virtual object in a virtual scene. It primarily includes two key parameters: distance and pitch angle. Distance: This refers to the straight-line spatial distance between the virtual sound source and the virtual object in the virtual scene. It is calculated using their position coordinates in the virtual scene coordinate system (or object coordinate system) and reflects the distance between the virtual sound source and the virtual object (e.g., the virtual sound source is 5 meters in front of the virtual object). Pitch angle: This refers to the angle between the line connecting the virtual sound source and the virtual object and the horizontal plane where the virtual object is located. It describes the vertical position of the virtual sound source relative to the virtual object (e.g., a 30° angle above the virtual object indicates the sound source is diagonally above the object, while a 20° angle indicates the sound source is diagonally below the object).
[0160] In some embodiments, the physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system and the physical direction vector of the physical sound source in the physical scene coordinate system. The virtual orientation information includes virtual position coordinates and virtual direction vector. The relative orientation information includes the distance between the virtual sound source and the virtual object and the pitch angle between the virtual sound source and the virtual object.
[0161] In some embodiments, the relative orientation information between the virtual sound source and the virtual object is determined based on the virtual orientation information corresponding to the physical orientation information in the virtual scene. This can be achieved by: determining the coordinate distance between the virtual position coordinates and the physical position coordinates as the distance between the virtual sound source and the virtual object; and determining the angle between the virtual direction vector and the physical direction vector as the pitch angle between the virtual sound source and the virtual object.
[0162] In some embodiments, when the physical orientation information includes the physical position coordinates (position parameters) and physical direction vector (direction parameters) of the physical sound source in the physical scene coordinate system, and the virtual orientation information correspondingly includes the virtual position coordinates (position parameters) and virtual direction vector (direction parameters), the process of determining the relative orientation information (distance and pitch angle) between the virtual sound source and the virtual object can be specifically broken down into two parts: For determining the distance, the core is to calculate the coordinate distance between the virtual position coordinates and the physical position coordinates. Here, the coordinate distance refers to the straight-line distance between two coordinate points in three-dimensional space, that is, by calculating the straight-line length between the two points through the spatial coordinate difference between the virtual position coordinates (the position of the virtual sound source in the virtual scene) and the physical position coordinates (the position of the physical sound source in the physical scene), so as to reflect the distance of the virtual sound source relative to the virtual object. For example, if the virtual position coordinates are (8, 5, 2) and the physical position coordinates are (352), then the coordinate distance between the two is 5 units, that is, the distance between the virtual sound source and the virtual object is 5 units. For determining the pitch angle, it is obtained by calculating the angle between the virtual direction vector and the physical direction vector. The virtual direction vector is the mapping of the physical direction vector in the virtual scene. The angle between the two intuitively reflects the vertical relationship between the virtual sound source and the virtual object: if the angle is 30°, it means that the virtual sound source is 30° above the virtual object; if the angle is -20° (or 20° downward angle), it means that the virtual sound source is 20° below the virtual object.
[0163] Thus, the distance between the virtual sound source and the virtual object is determined by the coordinate distance between the virtual and physical location coordinates, and the pitch angle is determined by the angle between the virtual and physical direction vectors. By directly linking the physical and virtual location coordinates and direction vectors, this ensures that the relative orientation information accurately reflects the real spatial relationship between the sound source and the object in the physical space, avoiding distance or angle distortion caused by complex conversion logic. For example, a sound source and an object that are physically 5 meters apart will maintain a relative distance of 5 meters in the virtual scene, and a sound source that is physically at a 30° pitch angle will also be presented at a 30° pitch angle in the virtual scene. Distance and pitch angle, as the core of relative orientation information, provide clear and accurate spatial parameters for subsequent audio rendering, enabling the rendered audio to adjust volume attenuation according to distance and optimize spectral characteristics according to pitch angle. This ensures that the sound orientation heard by the user in the virtual scene is highly consistent with that in the physical scene, greatly enhancing the realism and spatial immersion of the auditory experience.
[0164] In step 1042A, based on the relative orientation information and the acoustic feature parameters of the virtual sound source, the sound signal generated by the physical sound source is rendered to obtain the rendered audio.
[0165] In some embodiments, rendering audio refers to spatializing the original sound signal generated by the physical sound source by using the relative orientation information (distance, pitch angle, etc.) between the virtual sound source and the virtual object, and the acoustic characteristic parameters of the virtual sound source (such as volume, timbre, reverberation characteristics, etc.) as the core basis. The relative orientation information is used to simulate the propagation characteristics of sound in space—for example, adjusting the attenuation of sound based on distance (the greater the distance, the weaker the volume), and changing the spectral distribution of sound based on the pitch angle (such as making the high-frequency components of the sound source above more prominent). Simultaneously, the acoustic characteristic parameters of the virtual sound source (such as reverberation effects adapted to the virtual scene, basic volume, etc.) are used to modify the original sound signal so that it retains the essential characteristics of the physical sound source while accurately reflecting its spatial position and acoustic environment characteristics relative to the virtual object in the virtual scene.
[0166] In some embodiments, step 1042A above can be implemented as follows: matching the sound signal generated by the physical sound source with the acoustic feature parameters to obtain a matching result; when the matching result indicates that the sound signal does not match the acoustic feature parameters, adjusting the acoustic features of the sound signal based on the acoustic feature parameters to obtain a candidate sound signal; and performing audio rendering on the candidate sound signal based on the relative orientation information to obtain the rendered audio.
[0167] In some embodiments, acoustic feature parameters are a set of quantitative indicators used to describe the physical properties and auditory attributes of sound signals. They cover multiple parameters that reflect the essential characteristics of sound, mainly including: basic parameters such as frequency (the pitch of the sound, in Hertz), amplitude (the strength of the sound, corresponding to the volume), and duration (the duration of the sound); spectral features such as spectral envelope (the energy distribution profile of the sound frequency components), formants (the characteristic frequency regions that determine the timbre of vowels); and transient features (the rapid changes at the beginning and end of the sound), timbre features (the unique quality of sound determined by the harmonic distribution, such as distinguishing the sound of a piano from a violin), and reverberation features (the sustained effect formed by the reflection of sound in space), etc. These parameters together constitute the characteristic fingerprint of the sound, which is used in sound signal processing to accurately describe and compare sound characteristics. When the sound signal of the physical sound source does not match the preset acoustic characteristic parameters (such as the actual sound frequency being too high or the timbre deviating from the expectation), the sound signal can be adjusted in a targeted manner based on these parameters (such as correcting the frequency range and optimizing the spectrum distribution) to make it meet the acoustic requirements of the virtual scene for the sound source. This provides candidate signals that meet the characteristic standards for subsequent audio rendering combined with spatial orientation, ultimately ensuring that the rendered audio not only fits the spatial positioning but also has the acoustic quality that meets the scene settings.
[0168] In some embodiments, the original sound signal generated by the physical sound source is matched and compared with the acoustic characteristic parameters (such as preset volume, timbre, frequency range, etc.) of the virtual sound source. This matching checks whether the acoustic properties of the original sound signal conform to the virtual scene's settings for the sound source—for example, if the acoustic characteristic parameters of the virtual sound source require a medium volume and no noise, while the original sound signal from the physical sound source is too loud and contains electrical noise, the matching result will show a mismatch. When the matching result is a mismatch, the original sound signal needs to be adjusted according to the acoustic characteristic parameters of the virtual sound source to obtain candidate sound signals. For example, in the above example, the excessively loud volume will be reduced to a medium level, and noise reduction processing will be used to eliminate electrical noise, ensuring that the adjusted sound signal is consistent with the parameter requirements in terms of acoustic characteristics, thus ensuring that it conforms to the virtual scene's basic settings for the sound source. Based on the relative orientation information (distance, pitch angle, etc.) between the virtual sound source and the virtual object, audio rendering is performed on the candidate sound signals. For example, if the relative orientation information shows that the virtual sound source is 10 meters in front of the virtual object at an elevation angle of 30°, the volume of the candidate sound signal will be attenuated according to the distance (the attenuation amount corresponding to a distance of 10 meters), and the spectral distribution of the sound will be adjusted according to the elevation angle (enhancing high-frequency components to simulate the listening experience of the sound source above). The final rendered audio can not only meet the acoustic characteristics requirements of the virtual scene, but also accurately reflect its spatial position, allowing users to obtain a realistic auditory experience in the virtual scene.
[0169] As an example, suppose the physical sound source is a microphone playing a speech. The virtual scene sets the acoustic characteristics of the virtual sound source corresponding to this microphone as follows: volume 60 dB, no background noise, prominent mid-frequency, and relative location information showing that the virtual sound source is 5 meters in front of the virtual object (such as a virtual audience member) at an elevation angle of 10°. Matching the original sound signal generated by the microphone with the above acoustic characteristics reveals that the original sound signal volume is 75 dB (higher than the parameter requirements) and contains slight air conditioner noise (not meeting the requirement of no background noise), resulting in a mismatch. Adjusting the original sound signal based on the acoustic characteristics: reducing the volume from 75 dB to 60 dB, eliminating the air conditioner noise through noise reduction, and appropriately enhancing the mid-frequency range to make the sound clearer, yields a candidate sound signal that meets the parameter requirements. The candidate sound signals are rendered by combining relative orientation information: based on a distance of 5 meters in front, a slight volume attenuation (simulating long-distance propagation loss) and a very weak reverberation (simulating the natural acoustic effect of an open space) are applied to the candidate sound signals; based on an elevation angle of 10°, the high-frequency components of the sound are finely adjusted (so that the virtual audience perceives the sound coming from a direction slightly above the horizontal plane). The final rendered audio not only conforms to the acoustic settings of the virtual scene, but also allows the virtual audience to clearly perceive that the speech sound comes from a position of 5 meters in front and slightly upward, restoring a realistic sense of auditory space.
[0170] In this way, the sound signal from the physical sound source is matched with its acoustic feature parameters. When there is a mismatch, candidate sound signals are obtained through adjustment. These candidate signals are then combined with relative location information for audio rendering. Through matching and adjustment, it is possible to ensure that the sound signal is consistent with the virtual scene settings in terms of basic acoustic features (such as volume, timbre, and noise reduction). This avoids the destruction of the acoustic logic of the virtual scene due to deviations in the original signal of the physical sound source (such as excessive volume or noise), providing a clean audio foundation that meets the requirements for subsequent rendering. Rendering based on relative location information further endows the sound with spatial attributes, allowing the adjusted sound to present a listening experience that conforms to spatial laws based on parameters such as distance and pitch angle (e.g., distant sounds are weaker, and high-frequency sounds from above are more prominent). This two-step processing method, which first calibrates the basic features and then endows them with spatial attributes, not only ensures the standardization of the acoustic features of the virtual sound source but also achieves accurate spatial positioning of the sound in the virtual scene. Ultimately, the rendered audio not only fits the virtual scene settings but also has a realistic sense of spatial immersion, greatly enhancing the realism of the user's auditory experience in the virtual environment.
[0171] In some embodiments, the virtual object includes a plurality of audio receiving units, the relative orientation information includes relative orientation sub-information between the virtual sound source and each of the audio receiving units, and the rendered audio includes sub-rendered audio corresponding to each of the audio receiving units.
[0172] In some embodiments, the above-mentioned audio rendering of the candidate sound signal based on the relative orientation information to obtain the rendered audio can be achieved in the following way: for each audio receiving unit, audio rendering of the candidate sound signal based on the relative orientation sub-information is performed to obtain the sub-rendered audio corresponding to each audio receiving unit.
[0173] In some embodiments, when a virtual object contains multiple audio receiving units (such as two receiving units simulating human ears, or multiple receiving points simulating a multi-microphone array), the relative orientation information is refined into relative orientation sub-information between the virtual sound source and each audio receiving unit (i.e., independent parameters such as the distance and elevation angle between each receiving unit and the virtual sound source). The final rendered audio is also correspondingly split into sub-rendered audio specific to each receiving unit. The specific implementation of audio rendering of candidate sound signals based on relative orientation information is as follows: for each audio receiving unit on the virtual object, the relative orientation sub-information corresponding to the unit is used individually (e.g., the virtual sound source is 3 meters away from the left receiving unit at an elevation angle of 5°, and 3.2 meters away from the right receiving unit at an elevation angle of 3°) to perform targeted processing on the candidate sound signals—for example, when rendering the left receiving unit, the volume attenuation and spectral characteristics are adjusted according to its corresponding distance and angle; when rendering the right receiving unit, different parameters are adjusted according to the distance and angle of the right unit. In this way, each audio receiving unit can obtain sub-rendered audio that conforms to its spatial relationship with the virtual sound source. Ultimately, these sub-rendered audios together constitute the complete rendered audio, enabling virtual objects to accurately determine the spatial position of the virtual sound source through the perceptual differences of multiple receiving units (such as subtle differences in volume and timbre heard by the left and right ears), just like real objects, thus enhancing the stereoscopic and realistic feel of virtual hearing.
[0174] As an example, suppose the virtual object is a virtual human with an audio receiving unit on each side of its head (simulating the left and right ears of a human). The virtual sound source is a virtual speaker playing music in the virtual scene. The relative orientation information shows that the virtual speaker is 4 meters away from the left receiving unit (left ear) at an elevation angle of 8°, and 4.3 meters away from the right receiving unit (right ear) at an elevation angle of 6°. During audio rendering, the left and right receiving units are processed separately: For the left receiving unit, based on the relative azimuth information of 4 meters and an elevation angle of 8°, the candidate sound signal (adjusted to match the acoustic characteristics of the virtual scene) is processed—due to the slightly closer distance and slightly larger elevation angle, the volume of the left channel is slightly stronger, while the high-frequency components are enhanced to simulate the slightly higher position of the sound perceived by the left ear; For the right receiving unit, based on the information of 4.3 meters and an elevation angle of 6°, the same candidate sound signal is adjusted—due to the slightly farther distance, the volume of the right channel is about 5% weaker than the left channel, the smaller elevation angle results in slightly fewer high-frequency components than the left channel, and a very slight time delay is added (simulating the time difference between the two ears receiving sound). Finally, the left receiving unit obtains the left sub-rendered audio adapted to its location, and the right receiving unit obtains the right sub-rendered audio adapted to its location, together forming the complete rendered audio. When the virtual human receives these two sub-rendered audios, it can, like a real person, determine the position of the virtual speaker slightly further and slightly higher to the left front by the difference in sound between the left and right ears, achieving realistic spatial auditory perception.
[0175] Thus, when a virtual object contains multiple audio receiving units, and the final rendered audio is generated by rendering sub-rendered audio for each unit based on corresponding relative orientation sub-information, the multiple audio receiving units simulate the characteristics of multiple receiving points in a real auditory system (such as human binaural ears or microphone arrays). The dedicated rendering for each unit can accurately reproduce the spatial differences of the virtual sound source relative to different receiving points. For example, different receiving units have different distances and angles from the virtual sound source, and the corresponding sub-rendered audio will show subtle differences in volume, spectrum, and time difference. This difference is like the natural difference between sounds heard from different locations in a real scene. The combination of these sub-rendered audios allows the virtual object to more accurately locate the spatial position of the virtual sound source through multi-dimensional auditory information, greatly enhancing the stereoscopic sense and directional recognition of virtual hearing. At the same time, this unit-based processing method also provides flexible support for the differentiation of multiple sound sources and the realization of surround sound effects in complex virtual scenes. Ultimately, it allows users to obtain a spatial auditory experience in the virtual environment that is highly consistent with the real world, enhancing the overall immersion.
[0176] In some embodiments, the above-mentioned audio rendering of the candidate sound signal based on the relative orientation information to obtain the rendered audio can be achieved in the following manner: obtaining the acoustic characteristics of the candidate sound signal, and determining the target acoustic characteristics of the rendered audio based on the relative orientation information between the virtual sound source and the virtual object and the acoustic characteristics of the candidate sound signal; and performing audio rendering on the candidate sound signal based on the target acoustic characteristics to obtain the rendered audio.
[0177] In some embodiments, the existing acoustic characteristics of the candidate sound signal (such as current volume, frequency distribution, and whether it has reverberation) are acquired and combined with the relative orientation information (distance, pitch angle, and other spatial parameters) between the virtual sound source and the virtual object to jointly determine the target acoustic characteristics that the rendered audio needs to achieve. The target acoustic characteristics here are a comprehensive standard that integrates the original sound features and spatial environment requirements—for example, if the acoustic characteristics of the candidate sound signal are a volume of 50 dB and mid-frequency dominance, and the relative orientation information shows that the virtual sound source is 10 meters in front of the virtual object at a pitch angle of 15°, then the target acoustic characteristics might be determined as a volume reduction to 35 dB (simulating attenuation over a 10-meter distance), slightly weaker high frequencies (adapting to the spectral changes caused by the pitch angle), preservation of mid-frequency dominance, and the addition of slight natural reverberation (simulating the environmental impact of long-distance propagation). The key is to ensure that the target characteristics neither lose the essential characteristics of the candidate sound signal nor fail to accurately reflect its positional attributes in the virtual space. Based on the determined target acoustic characteristics, specific audio rendering processing is performed on the candidate sound signal. For example, if the target characteristics require a volume of 35 dB, slightly weak high frequencies, and a slight reverberation, the candidate signal will be reduced from 50 dB to 35 dB using a volume attenuation algorithm, the energy of the high frequency band will be weakened using an equalizer, and then an ambient reflection sound that matches the distance of 10 meters will be added using a reverberation algorithm. The final rendered audio will fully meet the requirements of the target acoustic characteristics.
[0178] In some embodiments, acoustic properties refer to the measurable and describable physical attributes and auditory characteristics exhibited by sound during propagation and perception, and are the core attributes reflecting the essence of sound. Acoustic properties may include volume and pitch: Volume (also known as loudness) is the subjective perception of the strength of a sound, mainly determined by the amplitude (vibration amplitude) of the sound signal, usually quantified in decibels (dB). The larger the amplitude, the higher the volume, reflecting whether the sound is loud or weak; Pitch is the subjective perception of the highness or lowness of a sound, mainly determined by the frequency (vibration speed) of the sound signal. The higher the frequency, the higher the pitch (such as treble notes), and the lower the frequency, the lower the pitch (such as bass drums), reflecting the difference in pitch. By acquiring the acoustic characteristics of candidate sound signals, such as volume and pitch, and combining them with the relative orientation information (such as distance and angle) between the virtual sound source and the virtual object, the target acoustic characteristics of the rendered audio can be determined. For example, based on the relative orientation at a distance, the target volume needs to be reduced according to the attenuation law; based on the spatial characteristics of the virtual scene, the target pitch may need to retain a specific frequency range to adapt to environmental reflections, etc., so that the acoustic performance of the rendered audio not only conforms to the original sound characteristics but also fits the spatial logic of the virtual scene.
[0179] In some embodiments, the above-mentioned audio rendering of the candidate sound signal based on the target acoustic characteristics to obtain the rendered audio can be achieved by adjusting the acoustic characteristics of the candidate sound signal to the target acoustic characteristics to obtain the rendered audio.
[0180] As an example, suppose the candidate sound signal is an audio clip of a piano performance with the following acoustic characteristics: volume 70 dB, bright timbre (prominent high-frequency components), and no reverberation. The relative location information between the virtual sound source and the virtual object (such as a virtual listener) is 8 meters away and 20° at an elevation angle (the virtual sound source is diagonally above the virtual listener). The acoustic characteristics of the candidate sound signal (70 dB, prominent high frequencies, no reverberation) are obtained, and combined with the relative location information to determine the target acoustic characteristics: considering the 8-meter distance, the sound will naturally attenuate, so the target volume is set at 55 dB; due to the 20° elevation angle, the high frequencies of the sound from the diagonally above will be more pronounced in real-world listening, so the prominent high-frequency characteristic is retained but slightly enhanced; simultaneously, the 8-meter distance will cause slight environmental reflections, so the target acoustic characteristics also include the addition of weak reverberation (0.5 seconds). In summary, the target acoustic characteristics are determined to be a volume of 55 dB, slightly stronger high frequencies than the original signal, and a weak 0.5-second reverberation. The candidate sound signal is adjusted based on the target acoustic characteristics: the volume is reduced from 70 dB to 55 dB using the volume adjustment tool, the energy of the high frequency band is moderately boosted using the equalizer, and a 0.5-second weak reverb is added using the reverb effect to make the acoustic characteristics of the candidate sound signal completely match the target acoustic characteristics. The final rendered audio not only retains the bright timbre of the piano, but also presents a spatial listening experience from 8 meters above by adapting the distance and angle, allowing virtual listeners to have a realistic listening experience.
[0181] As an example, based on the relative orientation information between the virtual sound source and the virtual object and the acoustic characteristics of the candidate sound signal, the target acoustic characteristics of the rendered audio are determined. Specifically, this can be real-time volume adjustment: based on the distance attenuation formula V = V0 × (d0 / d)α (V0 is the initial volume, d0 is the reference distance, d is the current distance, and α is the attenuation index, taken as 1.5 for indoor scenes and 2.0 for outdoor scenes). For example, when the user is 1 meter away from the sound source, the volume is 0.8, and when the user moves to 2 meters, it automatically attenuates to 0.4. Dynamic pitch correction: according to the Doppler effect formula f′ = f × (v + vr) / (v + vs) (f is the original frequency, v is the speed of sound, vr is the receiver's speed, and vs is the sound source's speed). When the user approaches a 440Hz sound source at a speed of 10m / s, the perceived frequency becomes 440 × (340 + 10) / (340 + 0) ≈ 453Hz.
[0182] In this way, the determination of the target acoustic characteristics fully integrates the original acoustic features of the candidate signal (such as timbre, basic volume, etc.) with the spatial information of the virtual space (such as attenuation caused by distance, and spectral changes caused by angle), ensuring that the target characteristics do not deviate from the essential attributes of sound, and can accurately reflect its spatial position in the virtual scene, avoiding the problem of losing the characteristics of the sound source itself due to simply emphasizing spatial attributes. By directly adjusting the candidate signal to the target characteristics, the rendering process has a clear direction and standard, and can efficiently and accurately achieve the adaptation of acoustic characteristics—for example, ensuring that the bright timbre of the piano sound is not destroyed, while adjusting parameters such as volume and reverberation to reflect its spatial sense of coming from an oblique position 8 meters away. This allows the rendered audio to perfectly integrate into the spatial logic of the virtual scene while retaining the true characteristics of the sound source, ultimately presenting users with an auditory experience that is both realistic and consistent with the virtual environment setting, effectively improving the naturalness and immersion of virtual audio.
[0183] In some embodiments, the above-mentioned audio rendering of the candidate sound signal based on the relative orientation information to obtain the rendered audio can be achieved in the following manner: adjusting the acoustic characteristics of the candidate sound signal based on the relative orientation information to obtain candidate rendered audio; selecting the target filter coefficient corresponding to the virtual sound source from a plurality of preset filter coefficients based on the relative orientation information; and performing audio rendering on the candidate rendered audio based on the target filter coefficient to obtain the rendered audio.
[0184] In some embodiments, the acoustic characteristics of candidate sound signals are adjusted based on relative orientation information to obtain candidate rendered audio. This acoustic characteristic adjustment primarily targets the basic spatial properties of the sound, such as adjusting the volume according to the distance between the virtual sound source and the virtual object (the greater the distance, the more obvious the volume attenuation), adjusting the spectral distribution of the sound according to the pitch angle (e.g., enhancing high frequencies for upper sound sources and weakening high frequencies for lower sound sources), and adding basic reverberation based on the openness / enclosure of the virtual scene. For example, if the relative orientation information shows that the virtual sound source is 15 meters in front of the virtual object at a 10° pitch angle, the volume of the candidate sound signal will first be reduced from the original 60 dB to 40 dB (simulating natural attenuation at a 15-meter distance), and the high-frequency components will be slightly enhanced (to suit the listening experience at a 10° pitch angle), while a slight reverberation will be added (simulating sound reflection in an open space), resulting in candidate rendered audio that is initially adapted to the spatial properties.
[0185] In some embodiments, a target filter coefficient is selected from multiple preset filter coefficients based on relative orientation information. The preset filter coefficients are parameters pre-set for different spatial orientations (such as different distances and angle combinations), and their function is to simulate the frequency response characteristics of sound at specific orientations—for example, distant sounds often experience more significant high-frequency attenuation (corresponding to a high-frequency suppression filter coefficient), while mid-frequency components from side sound sources may have a slight enhancement (corresponding to a mid-frequency optimization filter coefficient). Once the relative orientation information is determined, the system matches the corresponding preset scenario: for example, the aforementioned 5-meter, 10° elevation angle scenario might correspond to a preset mid-to-long distance + small elevation angle filter coefficient (this coefficient is preset to moderately suppress high frequencies and fine-tune mid-frequency gain), and thus selects this coefficient as the target filter coefficient.
[0186] In some embodiments, candidate rendered audio is processed based on target filtering coefficients to obtain the final rendered audio. The target filtering coefficients perform more refined optimization of the frequency components of the candidate rendered audio, making it conform to the true auditory characteristics at that location. For example, when processing candidate rendered audio with the aforementioned mid-to-long distance + small elevation angle filtering coefficients, it further weakens the sharp components in the high frequencies (avoiding excessive harshness from distant sounds) and slightly boosts the fundamental frequencies of vocals or instruments in the mid-range (ensuring clarity of core sound information), ultimately allowing the candidate rendered audio to more accurately reproduce subtle auditory differences at that location, beyond its basic spatial attributes.
[0187] In some embodiments, preset filter coefficients refer to the left and right ear filter coefficients corresponding to a specific spatial orientation, obtained from the Head Related Transfer Function (HRTF) library. The HRTF library is a database built through extensive acoustic measurements, containing filtering characteristic data of sound waves reaching the human ear from different directions (such as different distances, angles, and pitch angles), resulting from the reflection, scattering, and absorption of sound waves by physiological structures such as the head, auricle, and torso. These data are extracted into filter coefficients corresponding to the left and right ears, i.e., preset filter coefficients—wherein, the left ear filter coefficient corresponds to the acoustic filtering characteristics when sound reaches the left ear from a specific orientation, and the right ear filter coefficient corresponds to the characteristics when sound reaches the right ear from the same orientation. When it is necessary to simulate the auditory perception of sound at a specific orientation in a virtual scene, the preset filter coefficients corresponding to that orientation can be directly retrieved from the HRTF library for filtering the audio signal, thereby accurately reproducing the sound differences heard by the human ear at that orientation (such as subtle differences in volume, time, and spectrum between the left and right ears), which is a core parameter for achieving realistic auditory localization in virtual space.
[0188] As an example, suppose in the virtual scene, the virtual sound source is a virtual bird playing birdsong, and the virtual object is a virtual pair of ears simulating human hearing. The relative orientation information is that the virtual bird is 30° to the right front of the virtual object, 5 meters away, and at an elevation angle of 15°. Based on the relative orientation information, the candidate sound signal of the birdsong (already adjusted to conform to the basic acoustic characteristics of the virtual scene) is acoustically adjusted: based on the distance of 5 meters, the volume is reduced from the original 65 dB to 50 dB (simulating mid-distance attenuation); based on the elevation angle of 15°, the high-frequency components are slightly enhanced (simulating the listening experience of sound from above); combined with the open space attribute to the right front, a very weak natural reverberation is added to obtain a candidate rendering audio that is initially adapted to the macroscopic characteristics of the space. Based on the relative orientation information of 30° to the right front and 15° elevation angle, the corresponding target filter coefficient is selected from the HRT F library (the source of preset filter coefficients). The HRTF library pre-stores the left and right ear filter coefficients for this orientation. The right ear filter coefficients, due to the sound source being to the right front, retain more mid-frequency energy and have less delay, while the left ear filter coefficients, due to head obstruction, have slightly weaker high frequencies and a slight time difference. This set of coefficients serves as the target filter coefficients. The selected target filter coefficients are used to process the candidate rendered audio: the right ear filter coefficients make the bird calls received by the right ear clearer and slightly louder, while the left ear filter coefficients add subtle high-frequency attenuation and delay to the sound received by the left ear. After this processing, the final rendered audio not only reflects the spatial sense of 5 meters away and slightly upward through pre-processing, but also restores the difference in hearing between the left and right ears caused by a 30° angle to the right front through the HRTF filter coefficients. This allows the virtual object (simulating a human) to accurately distinguish that the bird calls are coming from a specific location to the right front, achieving a realistic spatial auditory experience.
[0189] Thus, by adjusting the acoustic characteristics of candidate sound signals based on relative location information (such as volume attenuation, spectrum adaptation, and basic reverberation addition), the sound can be quickly made to conform to the macroscopic location characteristics of the virtual space, laying a foundation for subsequent processing that conforms to spatial logic. Selecting target filter coefficients from preset filter coefficients (such as filter coefficients corresponding to the location in the HRTF library) utilizes pre-measured acoustic data to capture subtle auditory differences caused by human physiological structures (such as the head and auricle) at different locations, compensating for the shortcomings of basic adjustments in detail reproduction. Finally, secondary rendering of the candidate audio using the target filter coefficients allows the sound to incorporate subtle acoustic features specific to the location (such as volume differences, time differences, and spectrum differences between the left and right ears) in addition to macroscopic spatial attributes. This ensures both overall matching of the sound with the location of the virtual space and restores subtle differences in real hearing through professional filter coefficients, ultimately giving the rendered audio both spatial realism and auditory delicacy, significantly improving the accuracy of sound positioning in the virtual scene and the user's auditory immersion.
[0190] In some embodiments, see Figure 7, Figure 7 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 5 , Figure 7 Step 104 shown can be achieved through Figure 3 Steps 1041B to 1043B shown are implemented.
[0191] In step 1041B, the relative orientation information between the virtual sound source and the virtual object is determined based on the virtual orientation information in the virtual scene corresponding to the physical orientation information.
[0192] In some embodiments, relative orientation information refers to quantitative data used to describe the spatial relationship between a virtual sound source and a virtual object, calculated based on the virtual orientation information (position, direction) of the virtual sound source and the spatial attributes of the virtual object in a virtual scene. It primarily includes two key parameters: distance and pitch angle. Distance: This refers to the straight-line spatial distance between the virtual sound source and the virtual object in the virtual scene. It is calculated using their position coordinates in the virtual scene coordinate system (or object coordinate system) and reflects the distance between the virtual sound source and the virtual object (e.g., the virtual sound source is 5 meters in front of the virtual object). Pitch angle: This refers to the angle between the line connecting the virtual sound source and the virtual object and the horizontal plane where the virtual object is located. It describes the vertical position of the virtual sound source relative to the virtual object (e.g., a 30° angle above the virtual object indicates the sound source is diagonally above the object, while a 20° angle indicates the sound source is diagonally below the object).
[0193] In some embodiments, the physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system and the physical direction vector of the physical sound source in the physical scene coordinate system. The virtual orientation information includes virtual position coordinates and virtual direction vector. The relative orientation information includes the distance between the virtual sound source and the virtual object and the pitch angle between the virtual sound source and the virtual object.
[0194] In some embodiments, the relative orientation information between the virtual sound source and the virtual object is determined based on the virtual orientation information corresponding to the physical orientation information in the virtual scene. This can be achieved by: determining the coordinate distance between the virtual position coordinates and the physical position coordinates as the distance between the virtual sound source and the virtual object; and determining the angle between the virtual direction vector and the physical direction vector as the pitch angle between the virtual sound source and the virtual object.
[0195] In some embodiments, when the physical orientation information includes the physical position coordinates (position parameters) and physical direction vector (direction parameters) of the physical sound source in the physical scene coordinate system, and the virtual orientation information correspondingly includes the virtual position coordinates (position parameters) and virtual direction vector (direction parameters), the process of determining the relative orientation information (distance and pitch angle) between the virtual sound source and the virtual object can be specifically broken down into two parts: For determining the distance, the core is to calculate the coordinate distance between the virtual position coordinates and the physical position coordinates. Here, the coordinate distance refers to the straight-line distance between two coordinate points in three-dimensional space, that is, by calculating the straight-line length between the two points through the spatial coordinate difference between the virtual position coordinates (the position of the virtual sound source in the virtual scene) and the physical position coordinates (the position of the physical sound source in the physical scene), so as to reflect the distance of the virtual sound source relative to the virtual object. For example, if the virtual position coordinates are (8, 5, 2) and the physical position coordinates are (352), then the coordinate distance between the two is 5 units, that is, the distance between the virtual sound source and the virtual object is 5 units. For determining the pitch angle, it is obtained by calculating the angle between the virtual direction vector and the physical direction vector. The virtual direction vector is the mapping of the physical direction vector in the virtual scene. The angle between the two intuitively reflects the vertical relationship between the virtual sound source and the virtual object: if the angle is 30°, it means that the virtual sound source is 30° above the virtual object; if the angle is -20° (or 20° downward angle), it means that the virtual sound source is 20° below the virtual object.
[0196] In step 1042B, the rendering accuracy of the sound signal is determined based on the relative orientation information between the virtual sound source and the virtual object.
[0197] In some embodiments, rendering accuracy refers to the degree to which the generated rendered audio matches the relative positional information (distance, pitch angle, and other spatial relationships) of virtual sound sources and virtual objects in the virtual scene during audio rendering of sound signals. It is a key indicator for measuring whether the rendering effect accurately conforms to the virtual spatial logic. Based on the relative positional information of virtual sound sources and virtual objects (such as requiring more delicate processing when close to the source and with small angle differences, while simplifying appropriately when far away and with large angle differences), the fineness of parameter adjustment during audio rendering is determined. For example, when the virtual sound source is near the virtual object (close distance) and the angle change is subtle, the rendering accuracy requirement is higher, requiring precise adjustment of the volume attenuation curve, left and right ear time difference, spectral details, etc., to restore the subtle positional differences of sound at close range. When the virtual sound source is far away and the angle is clear, the rendering accuracy can be appropriately reduced, focusing on ensuring the accuracy of macroscopic characteristics such as volume attenuation and basic reverberation. The setting of rendering precision directly affects the realism of virtual hearing: high-precision rendering allows users to clearly perceive subtle changes in the position of the sound source, while low-precision rendering reduces processing complexity while ensuring a basic sense of space. The balance between the two needs to be dynamically adjusted based on relative orientation information to achieve the optimal combination of effect and efficiency.
[0198] In some embodiments, the above-mentioned determination of the rendering precision of the sound signal based on the relative orientation information between the virtual sound source and the virtual object can be achieved by: querying an index entry containing relative orientation information from a preset mapping relationship, and determining the preset rendering precision in the index entry as the rendering precision of the sound signal.
[0199] In some embodiments, a direct association between relative orientation information and rendering precision is achieved through a preset mapping relationship. The specific process is as follows: First, a set of relative orientation information-rendering precision mapping relationships (i.e., preset mapping relationships) is pre-established in the system. This set includes multiple index entries, each corresponding to a specific set of relative orientation information (e.g., distance 3 meters, elevation angle 10°, distance 10 meters, elevation angle 30°, etc.) and a matching preset rendering precision (e.g., high precision, medium precision, low precision). When it is necessary to determine the rendering precision, it is only necessary to search for the index entry containing the relative orientation information in the preset mapping relationship based on the relative orientation information of the current virtual sound source and the virtual object (e.g., the distance and angle actually measured). Once found, the preset rendering precision corresponding to that entry (e.g., if a distance of 3 meters and an elevation angle of 10° are found, corresponding to high precision) is directly determined as the rendering precision of the current sound signal. This approach uses predefined correspondences to quickly match the appropriate rendering precision without complex real-time calculations. It ensures the adaptability of precision selection to spatial orientation (e.g., high precision for close distances to reproduce subtle differences, and low precision for distant distances to simplify processing) and improves the system's response efficiency, making the rendering process more efficient and stable.
[0200] In step 1043B, based on the virtual sound source, the sound signal generated by the physical sound source is rendered according to the rendering precision of the sound signal to obtain the rendered audio.
[0201] In some embodiments, using a virtual sound source as the core reference and combining it with a determined sound signal rendering precision, the original sound signal generated by the physical sound source undergoes targeted audio processing to obtain rendered audio that meets the requirements of the virtual scene. The virtual sound source implies that the rendering process must conform to the attributes of the virtual sound source in the virtual scene (such as virtual identity, acoustic characteristics, etc.) to ensure that the processing direction is consistent with the virtual scene's settings. The rendering precision of the sound signal clarifies the level of detail in the processing—if the rendering precision is high, more detailed adjustments are needed to the spatial parameters of the sound signal (such as the volume attenuation curve caused by distance, spectral details at different angles, left and right ear time differences, etc.). For example, a near-distance sound source needs to accurately simulate a 0.1 dB level volume change and a 0.5° angle difference in the spectral differences. If the precision is medium or low, some details can be simplified, focusing on ensuring the accuracy of macroscopic spatial characteristics such as volume and basic reverberation to balance effect and resource consumption.
[0202] As an example, suppose in a physical scene, there is a microphone (physical sound source) playing a lecture. Its physical location information is: physical position coordinates (10, 5, 1.5) (unit: meters, based on the physical scene coordinate system, with the origin at the corner of the conference room), and physical direction vector (0, 1, 0) (pointing towards the front of the conference room). The virtual scene is a virtual conference space, and the physical microphone corresponds to the virtual sound source (the virtual lecturer's voice source) in the virtual scene. The virtual object is the virtual participant (located in the virtual scene). Mapping the physical location information to the virtual scene: the physical position coordinates (10, 5, 1.5) correspond to the virtual position coordinates (10, 5, 1.5) in the virtual scene coordinate system (keeping the coordinates consistent and simplifying the mapping), and the physical direction vector (0, 1, 0) corresponds to the virtual direction vector (0, 1, 0) (pointing towards the front of the virtual conference space). At this point, the virtual object (virtual participant) is positioned at coordinates (5, 5, 1.5) in the virtual scene. Therefore, the relative orientation information between the virtual sound source and the virtual object can be determined: a distance of 5 meters (straight-line distance between two points) and a pitch angle of 0° (same height, on the same horizontal plane). Based on this relative orientation information, the rendering precision is determined: in the system's preset mapping relationship, a distance of 5 meters and a pitch angle of 0° correspond to medium-high precision (because the distance is moderate and they are on the same horizontal plane, requiring a more detailed reproduction of the difference in hearing between the left and right ears). Therefore, the rendering precision of the sound signal is determined to be medium-high precision. Based on a virtual sound source (virtual lecturer's voice source), the original sound signal (lecture audio) generated by the physical microphone is rendered with medium to high precision: Since the virtual sound source is a virtual lecturer, the clarity of their voice (basic acoustic characteristics) needs to be preserved; according to the medium to high precision requirements, the acoustic characteristics are first adjusted: based on a distance of 5 meters, the original volume (65 decibels) is attenuated to 50 decibels, and a slight reverberation is added (simulating sound reflection in a medium-sized space in a conference room); then, the processing is refined: since the pitch angle is 0° (same horizontal plane), the difference in hearing between the left and right ears needs to be restored - the virtual object is located 5 meters to the left of the virtual sound source, so the HRTF filter coefficient of the left sound source is applied to the sound signal (the volume of the left channel is slightly higher than that of the right channel, and the right ear signal is delayed by 0.002 seconds), while preserving the mid-frequency clarity of the human voice (consistent with the lecturer's voice characteristics). The final rendered audio allows virtual attendees to clearly perceive the voice of the virtual lecturer, located 5 meters in front of them and slightly to the left. This not only fits the spatial logic of the virtual scene but also, due to the high-precision processing, restores the subtle differences in auditory perception between the left and right ears, enhancing the immersive experience of the virtual meeting.
[0203] In this way, the physical-to-virtual orientation mapping ensures that the spatial relationships of sound in the virtual scene are highly consistent with the physical world, avoiding the disconnect between virtual hearing and physical reality. This lays a realistic spatial foundation for subsequent rendering. Because the rendering precision is dynamically determined by relative orientation information (e.g., high precision for close distances and adaptive precision for long distances), it can both restore subtle differences in sound orientation through fine processing in key scenes (such as close-range interaction) to enhance auditory realism, and reduce system resource consumption through appropriate simplification in non-critical scenes, achieving a balance between effect and efficiency. By combining the virtual sound source attributes with corresponding precision for rendering, the processed audio not only fits the virtual scene settings (e.g., the virtual lecturer's voice needs to be clearly distinguishable) but also accurately reflects its spatial location. This allows users to perceive the accurate orientation of sound in the virtual environment and obtain an auditory experience that conforms to the scene logic, ultimately significantly enhancing the immersion and efficiency of virtual hearing.
[0204] In some embodiments, before performing audio rendering on the sound signal generated by the physical sound source based on the virtual sound source to obtain the rendered audio, the following processing can be performed: obtaining the sound source type of the physical sound source, comparing the sound source type with a preset sound source type, and obtaining a comparison result; when the comparison result indicates that the sound source type is the preset sound source type, based on the virtual orientation information, searching for the preset rendered audio corresponding to the virtual orientation information in the audio library of the preset sound source type, and determining the found preset rendered audio as the rendered audio.
[0205] In some embodiments, the sound source type of the physical sound source is obtained and compared with a preset sound source type. Here, the sound source type refers to the sound attribute classification of the physical sound source (such as ambient sound, human voice, instrument sound, etc.), and the preset sound source type is a type predefined by the system that is suitable for replacing real-time rendering with preset audio (e.g., common ambient sounds such as wind and rain, or fixed scene sound effects such as keyboard clicks). The purpose of the comparison is to determine whether the current physical sound source belongs to a type that can be optimized—for example, if the physical sound source is a speaker playing rain sounds (sound source type is ambient sound-rain sound), and the preset sound source type includes ambient sound-rain sound, then the comparison result is the preset sound source type. When it is confirmed to be the preset sound source type, the preset rendering audio is searched from the corresponding audio library based on the virtual location information. The "Preset Audio Library for Preset Sound Source Types" is a collection of rendered audio pre-stored for each preset type, processed according to different virtual orientation information (distance, angle, etc.). For example, the ambient sound / rain sound audio library pre-stores rain sound rendered audio corresponding to different orientations such as 5 meters away, 10 meters away at 0° elevation angle, and 15° elevation angle (including volume, reverberation, spectrum, and other characteristics at that orientation). At this point, based on the virtual orientation information (such as virtual position and direction) mapped from the physical sound source to the virtual scene, the corresponding entry in the audio library is matched to find the preset rendered audio for that orientation. The found preset rendered audio is directly determined as the final rendered audio. Since these preset audios are pre-optimized for specific types and orientations, their effects already meet the requirements of the virtual scene; therefore, there is no need to render the original sound signal of the physical sound source in real time, and it can be directly called upon.
[0206] As an example, suppose in a virtual office scenario, the physical sound source is a running printer (its physical location corresponds to the virtual orientation information in the virtual scene: 8 meters away from the virtual object, directly in front, and at an elevation angle of 0°). First, the sound source type of this physical sound source is identified as office equipment - printer operating sound, and it is compared with the system's preset sound source types (including office equipment - printer operating sound, ambient sound - air conditioner sound, etc.). The comparison result shows that it belongs to the preset sound source type. At this point, the system will search for the corresponding entry in the preset audio library of office equipment - printer operating sound based on the virtual orientation information (8 meters, directly in front, elevation angle of 0°). This audio library pre-stores printer rendering audio from different orientations, among which the preset rendering audio corresponding to 8 meters, directly in front, and at an elevation angle of 0° has been pre-processed to have a moderate volume (simulating attenuation at 8 meters), slight office reverberation, and slightly weaker high-frequency mechanical sounds (consistent with the perception of hearing at a distance). Finally, the system directly determines this preset rendering audio as the final rendering audio, without needing to render the original sound signal generated by the printer in real time. This ensures that the rendering effect matches the auditory logic of that location in the virtual scene, and also reduces the computing power consumption of real-time calculation by calling preset resources, thus improving the system response speed.
[0207] In this way, by comparing the physical sound source type with the preset sound source type, high-frequency, stable sound sources (such as common office equipment sounds, ambient sounds, etc.) can be quickly identified. For these sound sources, the corresponding virtual location rendering audio from the preset audio library can be directly called, which can significantly reduce the computing power required for real-time rendering and improve system processing efficiency. Especially in complex virtual scenes with multiple sound sources running concurrently, it can effectively avoid latency problems caused by real-time computing overload. The preset rendering audio is finely tuned in advance for specific types and locations. Its acoustic characteristics (such as volume attenuation, reverberation adaptation, spectrum distribution, etc.) are more in line with the listening logic of the virtual scene. Compared with the parameter deviations that may occur in real-time rendering, it can ensure the consistency of rendering effect of the same sound source in the same location and enhance the stability of the virtual auditory experience.
[0208] In some embodiments, the above-mentioned audio rendering of the sound signal generated by the physical sound source based on the virtual sound source to obtain rendered audio can be achieved in the following way: when the comparison result indicates that the sound source type is not the preset sound source type, audio rendering of the sound signal generated by the physical sound source based on the virtual sound source is performed to obtain rendered audio.
[0209] In some embodiments, when the comparison result between the physical sound source's sound source type and the preset sound source type is a mismatch (i.e., the sound source type does not belong to the predefined types of callable preset audio), a regular real-time audio rendering process is initiated: using the virtual sound source as the core reference (including its attributes and acoustic characteristics in the virtual scene), the original sound signal generated by the physical sound source is processed in a targeted manner. Specifically, the relative orientation information (distance, angle, etc.) between the virtual sound source and virtual objects, and the environmental characteristics of the virtual scene (such as space size and acoustic materials), are combined to adjust the acoustic characteristics of the sound signal, such as volume, spectrum, reverberation, and directional filtering, in real time, ultimately generating rendered audio that conforms to the logic and auditory requirements of the virtual scene. For non-preset types of sound sources (usually sound sources with varied characteristics, infrequent occurrences, or those requiring precise reproduction of their original features, such as the voice of a specific person or the sound of a special instrument), real-time calculations ensure that the rendering effect not only fits the spatial relationship of the virtual scene but also preserves its unique sound characteristics. This avoids distortion or insufficient adaptability caused by relying on preset audio, thereby covering more types of sound sources while ensuring the realism and flexibility of the virtual auditory experience.
[0210] Thus, when the sound source type does not belong to the preset sound source type, the rendered audio is obtained by real-time audio rendering of the sound signal of the physical sound source based on the virtual sound source. For non-preset sound sources with variable characteristics, infrequent occurrences, or those requiring precise reproduction of unique features (such as personalized human voices, special instrument sounds, etc.), real-time rendering can avoid sound distortion or insufficient adaptability to the virtual scene caused by relying on fixed preset audio, ensuring that the unique acoustic characteristics of these sound sources (such as the unique timbre of an individual's voice, the exclusive sound quality of a special instrument) are fully preserved. At the same time, real-time processing with the virtual sound source as a reference allows the sound signal to accurately match its spatial attributes in the virtual scene (such as relative orientation, environmental characteristics, etc.), so that the rendered audio not only fits the logical settings of the virtual scene, but also truly reflects the characteristics of the sound source itself. In this way, while covering a wider variety of sound source types, it ensures the authenticity, uniqueness, and scene adaptability of the virtual auditory experience, further enhancing the overall immersive experience.
[0211] In step 105, based on the rendered audio, virtual audio is determined for sending to the virtual object.
[0212] In some embodiments, a virtual object refers to the final perceiving subject of virtual audio, that is, the specific virtual entity in the virtual scene that receives and perceives the virtual audio. Once the virtual audio to be sent to the virtual object is determined based on the rendered audio, it is clearly defined that the receiving object of the virtual audio is the virtual object itself. This means that the virtual audio is specifically generated for the virtual object and adapted to its perceptual needs—for example, if the virtual object is a virtual character simulating human hearing, the virtual audio will be optimized according to the character's auditory characteristics (such as binaural position, sensitivity to sound, etc.) so that it can be accurately perceived by the virtual character; if the virtual object is a virtual device with a specific audio receiving device (such as a virtual microphone array), the virtual audio will be adapted to the device's receiving parameters (such as pickup range, frequency response, etc.). This setting ensures that the virtual audio matches the perceptual logic of the receiver (virtual object), enabling the virtual audio to be effectively received by the virtual object and transformed into auditory information that conforms to its characteristics. This lays the foundation for subsequent interaction of the virtual object based on the audio (such as locating the sound source, responding to sound commands, etc.), ultimately realizing a closed-loop logic for audio transmission and perception in the virtual scene.
[0213] In some embodiments, a virtual object, as the ultimate perceptual subject of virtual audio, refers to a virtual entity in a virtual scene that possesses quantifiable perceptual parameters. Its core characteristic is that it has a clear set of audio perceptual attribute parameters. This set of parameters defines the range of audio reception and processing capabilities of the virtual object through preset technical indicators. These parameter sets include, but are not limited to, audio perceptual range parameters, frequency response parameters, spatial positioning accuracy parameters, and sensitivity threshold parameters, which constitute the technical basis for the virtual object to receive virtual audio.
[0214] In some embodiments, the audio-based interactive capability of virtual objects can be quantified as follows: sound source positioning accuracy: horizontal azimuth error ≤ ±1°, vertical azimuth error ≤ ±2°, distance error ≤ ±3% (relative to the actual virtual distance); command recognition rate: for virtual audio with voice commands, the correct command recognition rate of virtual objects is ≥99% (under a signal-to-noise ratio ≥20dB); interaction response latency: the time from receiving audio to triggering an interactive action is ≤100ms. The above technical parameter system, by clarifying the adaptation rules between the perceptual attributes of virtual objects and virtual audio, constructs a closed-loop quantitative logic for audio "generation-transmission-reception-perception" in virtual scenes, ensuring that virtual audio can be accurately received by virtual objects and used in subsequent interaction processes.
[0215] As an example, see Figure 15 , Figure 15 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 7 The virtual audio received by the virtual audio source 31 is virtual object 32.
[0216] In some embodiments, see Figure 8 , Figure 8 This is a flowchart illustrating the audio rendering method provided in the embodiments of this application. Figure 5 , Figure 8 Step 105 shown can be achieved through Figure 3 Steps 1051 to 1052 shown are implemented.
[0217] In step 1051, when there are multiple virtual sound sources, the rendered audio corresponding to the multiple virtual sound sources is synthesized to obtain the virtual audio.
[0218] In some embodiments, when there are multiple virtual sound sources, the process of synthesizing the rendered audio corresponding to multiple virtual sound sources to obtain virtual audio is to integrate multi-directional and multi-type sound information to construct a comprehensive audio signal that conforms to the overall auditory logic of the virtual scene. Each virtual sound source has completed independent audio rendering based on its relative position information with the virtual object, sound source type, etc., to obtain the corresponding rendered audio. These rendered audios each contain the spatial attributes of the sound source (such as volume caused by distance, spectral differences caused by angle) and its own acoustic characteristics (such as the timbre of human voices, the pitch of musical instruments). For example, there may be three virtual sound sources in the virtual scene: A is the dialogue sound 5 meters to the right front (the rendered audio is a clear mid-frequency dominant sound, with slightly higher volume in the right ear), B is the ambient music 10 meters to the left rear (the rendered audio is a low-volume, reverberant, wide-frequency sound), and C is the keyboard typing sound 3 meters directly in front (the rendered audio is a short, high-frequency sound, with equal volume in both ears).
[0219] In some embodiments, synthesis rules can be set according to the acoustic logic of the virtual scene, mainly including: volume balancing: adjusting the weights of each sound source according to its importance or the needs of the virtual scene. For example, dialogue (A) has the highest weight and its clarity must be ensured, while ambient music (B) has a lower weight and exists as background sound; phase alignment: ensuring that the audio signals of different sound sources are synchronized on the time axis to avoid auditory confusion caused by delay (e.g., keyboard sounds and dialogue sounds need to maintain a natural time difference, which conforms to the order of sound propagation in real scenes); frequency avoidance: when the frequency ranges of multiple sound sources overlap (e.g., dialogue sounds and the sound of a certain instrument are both concentrated in the mid-frequency range), the distortion after signal superposition is reduced by slightly adjusting the spectrum distribution to ensure the recognizability of each.
[0220] In some embodiments, multiple rendered audio files are superimposed according to synthesis rules: first, the volume of each rendered audio file is adjusted based on weights (e.g., A's volume remains at 70%, B's is reduced to 40%, and C's remains at 60%); then, the time-synchronized signals are superimposed on the same time axis, and the waveforms are fused using digital signal processing techniques (e.g., when the mid-frequency sound wave of A is superimposed with the broadband sound wave of B, their respective characteristic frequencies are preserved); finally, the synthesized signal is optimized as a whole, such as eliminating noise generated by superposition and adjusting the overall dynamic range to ensure that the synthesized audio has moderate loudness and clear layers. After synthesis, the originally independent multiple rendered audio files are merged into a unified virtual audio. For example, in the final virtual audio, the user can clearly distinguish the dialogue voice from the right front (dominant), while simultaneously feeling the background music from the left rear (creating atmosphere) and the keyboard sounds from the front (adding detail). Each sound maintains its own spatial positioning and characteristics, while together forming a coherent and natural virtual auditory environment.
[0221] In step 1052, when the number of virtual sound sources is one, the rendered audio corresponding to the virtual sound source is determined as the virtual audio.
[0222] As an example, suppose there is a virtual classroom scene where the virtual object is a virtual student (responsible for receiving virtual audio). When there are multiple virtual audio sources, for example, three: ① a virtual teacher (located 3 meters directly in front of the virtual student, rendered as a clear lecture, slightly weaker in the right ear and slightly stronger in the left ear to match the frontal orientation); ② virtual book-turning sound (from 2 meters to the left rear of the virtual student, rendered as a soft paper-rubbing sound with slight distance attenuation); ③ virtual birdsong outside the window (from 10 meters to the right front of the virtual student, rendered as a low-volume, high-frequency, crisp call). The system will synthesize these three rendered audios: first, adjust the weights according to scene logic (80% of the lecture volume is retained for clarity, 50% for the book-turning sound as close-up detail, and 30% for the birdsong as distant background); then, ensure time synchronization (e.g., a natural transition between the pauses in the book-turning sound and the lecture); finally, use audio overlay technology to fuse the waveforms of the three, eliminating frequency conflicts (e.g., reducing the mid-frequency overlap between the birdsong and the lecture to avoid noise). The synthesized virtual audio allows the virtual student to simultaneously perceive the clear lecture sound directly in front, the details of turning pages in the book to the left rear, and the distant birdsong to the right front, with each sound layered clearly and spatially accurately positioned.
[0223] Continuing the previous example, when there is only one virtual sound source: for instance, if there is only one virtual teacher in the scene (located 3 meters directly in front), the corresponding rendered audio has been adjusted according to its location to a clear lecture sound that is slightly weaker in the right ear and slightly stronger in the left ear. In this case, there is no need for synthesis; the rendered audio is directly designated as the virtual audio. The virtual students receive this single, precise lecture sound that reflects the teacher's location, ensuring the simplicity and accuracy of auditory information.
[0224] Thus, when multiple virtual sound sources exist, by synthesizing their respective rendered audio (such as weight adjustment, time synchronization, and frequency optimization), multi-directional and multi-type sound information can be organically integrated. This preserves the spatial positioning and acoustic characteristics of each sound source (such as the clear dominance of a lecture, the close-up details of a book-turning sound, and the distant background of birdsong), while also creating a layered and naturally harmonious overall auditory environment. This allows virtual objects to perceive rich sound information that conforms to the scene logic, avoiding auditory confusion caused by the mixing of multiple sound sources. When there is only one virtual sound source, its rendered audio can be directly determined as virtual audio. This simplifies the processing flow, reduces system resource consumption, and improves efficiency while ensuring sound accuracy (such as the sense of location of a single lecture).
[0225] In some embodiments, the above-mentioned synthesis of rendered audio corresponding to multiple virtual sound sources to obtain virtual audio can be achieved in the following way: clustering the virtual sound sources to obtain at least one virtual sound source group, wherein the virtual audio corresponds one-to-one with the virtual sound source group; for each virtual sound source group, synthesizing the rendered audio corresponding to each virtual sound source in the virtual sound source group to obtain the virtual audio corresponding to the virtual sound source group.
[0226] In some embodiments, the core of clustering virtual sound sources and grouping them to synthesize virtual audio is to logically categorize complex multi-sound-source scenes and then perform targeted processing to improve synthesis efficiency and audio layering. The clustering of virtual sound sources is based on their correlation within the virtual scene, which can include three types of logic: spatial location correlation: grouping virtual sound sources that are spatially close together (e.g., setting a distance threshold of 3 meters, sound sources smaller than this distance are considered "near-distance clusters"). For example, in a virtual living room, the sound of the television in the sofa area, the clinking of water glasses on the coffee table, and conversations on the sofa are all clustered into the sofa area group because their spatial distance is within 2 meters. Sound source type correlation: grouping virtual sound sources of the same category together (e.g., those that are both ambient sounds, both that are instrument sounds, etc.). For example, in a virtual concert, violins, cellos, and violas, all string instruments, are clustered into the string group, while trumpets and trombones, both brass instruments, are clustered into the brass group. Scene Function Association: Virtual sound sources serving the same scene function are grouped together (e.g., in a virtual classroom, the teacher's lecturing voice and the sound of chalk writing on the blackboard both serve teaching interaction and are grouped into the teaching group, while the sound of rain outside the window and the sound of the classroom air conditioner are both environmental background and are grouped into the background group). Each virtual sound source group formed after clustering represents a set of sounds that are synergistic in space, type, or function, and each group will generate an independent virtual audio.
[0227] In some embodiments, when synthesizing the rendered audio within each group for each virtual sound source to form a corresponding virtual audio, it is necessary to consider both intra-group coordination and inter-group differentiation. Intra-group coordination: Adjust the synthesis rules according to the correlation of the sound sources within the group. For example, in the string group, when synthesizing the rendered audio of the violin (high-frequency dominant) and cello (low-frequency dominant), their respective frequency characteristics are preserved and the volume ratio is adjusted (the violin is slightly louder to highlight the melody), while phase alignment is used to ensure natural sound blending; in the sofa area group, the weight of conversation (core information) is set to 70%, the weight of television sound (auxiliary information) is set to 50%, and the weight of the clinking of water glasses (detail information) is set to 30%, ensuring that the primary and secondary sounds within the group are distinct; Inter-group differentiation: Create a contrast between the virtual audio of different groups through differences in overall parameters. For example, the virtual audio of the teaching group retains high definition (reducing reverberation), while the virtual audio of the background group increases environmental reverberation and reduces the overall volume (accounting for 30% of the teaching group), so that the two groups of sounds form a foreground-background hierarchy in terms of auditory perception.
[0228] As an example, suppose there are multiple virtual sound sources in a virtual coffee shop scene: ① the sound of the coffee machine operating at the bar (4 meters to the right front), ② the sound of the bartenders talking (3.5 meters to the right front), ③ the sound of customers talking at the next table (6 meters to the left front), ④ the sound of cups and saucers clinking at the next table (5.8 meters to the left front), and ⑤ the sound of traffic outside the window (8 meters behind). These virtual sound sources are clustered: based on spatial location and functional association, the coffee machine sound (①) and the bartenders' conversation (②) are only 0.5 meters apart and both belong to the bar service area, so they are clustered into the bar group. The customer conversation (③) and the sound of cups and saucers clinking (④) are 0.2 meters apart and both belong to the customer dining area, so they are clustered into the next table group. The traffic sound outside the window (⑤) is independent of the indoor scene and is clustered into the environment group.
[0229] Continuing from the previous example, the virtual sound source groups are synthesized as follows: The bar group is synthesized by combining the coffee machine sound (high-frequency mechanical sound) and the waiter's conversation (mid-frequency human voice). Each sound retains its own frequency characteristics, and the weights are adjusted (60% for conversation to ensure clarity, and 40% for environmental details). Time synchronization ensures natural sound superposition, resulting in the "bar group virtual audio" (service scene sound in the right front area). The neighboring table group is synthesized by combining the customer conversation (dominant information) and the clinking of cups and saucers (auxiliary details). The conversation accounts for 70% of the volume, and the cup and saucer sounds account for 30%. The overlapping mid-frequency portion of both is weakened to avoid noise, resulting in the neighboring table group virtual audio (dining scene sound in the left front area). The environment group is synthesized by using only one sound source: traffic noise. Its rendered audio is directly used as the environment group virtual audio (background street sound in the distance). Three virtual sound source groups generate three virtual audios, which together constitute the complete auditory environment of the coffee shop. The virtual objects can clearly distinguish the service interaction at the bar in front of the right, the dining conversation at the neighboring table in front of the left, and the street background outside the window behind. The sounds in each area maintain coordination within the group and form a clear spatial and functional distinction, making the virtual auditory experience closer to the hierarchical logic of the real scene.
[0230] In this way, the clustering process groups multiple virtual sound sources based on spatial location, sound source type, or scene function, significantly simplifying the complexity of multi-source synthesis and avoiding the sound chaos caused by directly superimposing all sound sources. This allows for more efficient processing of complex scenes (e.g., simplifying the processing of 10 independent sound sources to processing 3 groups). When synthesizing each group, specific rules (e.g., adjusting weights, optimizing frequency coordination) can be formulated based on the correlation of sound sources within the group (e.g., same region, same type). This ensures the synergy of sounds within the group (e.g., the coffee machine sound and human voice in the bar group blend naturally) while creating clear hierarchical distinctions through differences in parameters between groups (e.g., clearer sound in the bar group, weaker sound in the environment group), giving the virtual audio a "regionalized" and "functional" auditory structure. The virtual audio corresponding to multiple groups together constructs a logically clear and layered virtual auditory environment, enabling virtual objects to perceive sound information from different regions or functions more naturally (e.g., distinguishing between bar service sounds and conversations at neighboring tables). This significantly improves the immersion and recognizability of virtual hearing in complex scenes while also ensuring processing efficiency.
[0231] In some embodiments, when the number of virtual sound sources is multiple, determining the virtual audio to be sent to the virtual object based on the rendered audio can be achieved by: for each audio receiving unit, synthesizing the sub-rendered audio corresponding to each virtual sound source for the audio receiving unit to obtain the virtual audio of the audio receiving unit.
[0232] In some embodiments, when there are multiple virtual sound sources, to determine the virtual audio to be sent to the virtual object based on the rendered audio, it can be achieved by grouping and synthesizing audio receiving units. The core is to process and synthesize the corresponding virtual audio for each audio receiving unit on the virtual object (such as the left and right ears simulating human hearing, or multiple microphone array units of a virtual device). Specifically, firstly, the audio receiving units on the virtual object are identified (for example, a virtual human has two receiving units, a left ear and a right ear, and a virtual robot has multiple receiving units such as a forward microphone and a side microphone); then, each virtual sound source generates sub-rendered audio for each receiving unit based on its relative orientation information (such as distance and angle) with each audio receiving unit—these sub-rendered audios are adapted to the perceptual characteristics of the receiving unit (e.g., the sub-audio received by the left ear will show high-frequency attenuation caused by head occlusion, while the right ear does not have this feature); finally, for each audio receiving unit, the sub-rendered audios corresponding to all virtual sound sources are synthesized (e.g., adjusting the volume weight of each sub-audio, synchronizing the time axis, and optimizing frequency superposition) to obtain the virtual audio specific to that receiving unit.
[0233] As an example, the virtual object is a virtual person (containing two receiving units, left and right ears), and there are two virtual sound sources: A (speech 5 meters to the right front) and B (footsteps 3 meters to the left rear). Sound source A generates sub-rendered audio for the left ear (slightly lower volume, slightly weaker high frequencies) and sub-rendered audio for the right ear (slightly higher volume, clearer high frequencies); sound source B generates sub-rendered audio for the left ear (higher volume, no significant attenuation) and sub-rendered audio for the right ear (lower volume, attenuated high frequencies). Subsequently, the virtual audio for the left ear is a synthesis of the left ear sub-audio of A and the left ear sub-audio of B (emphasizing the clarity of the footsteps to the left rear), and the virtual audio for the right ear is a synthesis of the right ear sub-audio of A and the right ear sub-audio of B (emphasizing the clarity of the speech to the right front). This method ensures that the virtual audio of each receiving unit accurately reflects its relative relationship with all sound sources. By leveraging the audio differences among multiple receiving units, more realistic spatial auditory positioning is achieved, enhancing the immersive experience of the virtual scene.
[0234] Thus, by processing and synthesizing sub-rendered audio corresponding to all virtual sound sources separately for each receiving unit (such as the left and right ears of a virtual human, or different microphones of a virtual device), the unique spatial relationship between each receiving unit and multiple sound sources can be accurately reproduced. For example, the virtual audio received by the left ear will naturally reflect the occlusion effect of the head on the right sound source, while the right ear will show the attenuation characteristics of the left sound source. This difference makes the virtual object's perception of sound closer to the auditory logic of the real physical world. Unit-based synthesis avoids the mixing of sound information from different receiving units, ensuring that the virtual audio of each unit clearly reflects its unique acoustic characteristics (such as position and sensitivity), making sound localization in multi-sound-source scenarios more accurate (such as being able to distinguish the directional difference between a voice speaking from the right front and a footstep from the left rear). The synergistic effect of the virtual audio of each receiving unit not only preserves the rich information of multiple sound sources, but also constructs a three-dimensional spatial auditory experience through the auditory differences between units, greatly improving the realism and immersion of sound perception in virtual scenes, while providing reliable audio basis for virtual objects to perform precise interactions based on sound (such as turning towards the sound source).
[0235] Thus, by acquiring the physical location information of the physical sound source in the physical scene and transforming it based on the virtual objects included in the virtual scene to obtain virtual location information, the spatial relationships of the physical world are accurately mapped in the virtual scene. This ensures a logical correspondence between the position of the virtual sound source and the actual location of the physical sound source, laying the foundation for the spatial realism of the virtual audio. Based on the virtual location information corresponding to the physical location information in the virtual scene, a virtual sound source corresponding to the physical sound source is created in the virtual scene. Using the virtual sound source as a reference, the sound signal generated by the physical sound source is rendered to obtain the rendered audio. This allows the rendering process to fully integrate the spatial characteristics of the virtual scene (such as the distance and angle between the virtual sound source and virtual objects), so that the rendered audio naturally carries the acoustic characteristics that conform to the virtual location (such as volume attenuation at long distances and spectral differences at specific angles). By determining the virtual audio to be sent to the virtual object based on the rendered audio, and ensuring that the entire processing revolves around the perception logic of the virtual object, the final generated virtual audio not only retains the original sound characteristics of the physical sound source, but also accurately restores its spatial orientation relative to the virtual object in the virtual scene. This makes the sound received by the virtual object highly consistent with the real auditory experience in terms of spatial attributes such as orientation and distance, thereby significantly improving the simulation degree of the virtual audio.
[0236] The following will describe an exemplary application of the embodiments of this application in a real virtual reality application scenario.
[0237] Virtual social scenarios: In virtual meeting platforms such as Meta Horizon Workrooms, spatial audio effects are achieved for multi-person conversations, allowing users to determine the speaker's location through sound direction, enhancing the realism of communication. VR game interaction applications: In VR games, precise spatial audio cues indicate enemy location (e.g., footsteps coming from 3 meters to the left and rear), enhancing immersion and interactivity. Medical training simulation applications: In surgical training VR systems, spatial audio simulates the direction of surgical instrument sounds (e.g., the sound of a scalpel cutting tissue coming from 45° in front), helping trainees develop spatial awareness. Cultural heritage restoration applications: In VR tours, spatial audio recreates the natural sound effects of cultural heritage (e.g., sound coming from the direction of a cave entrance) and the spatial relationship between the sounds and historical explanations, enhancing the immersive experience of the cultural experience.
[0238] Virtual Reality (VR) is a computer technology that enables the creation and experience of simulated three-dimensional virtual environments. Its core lies in generating highly realistic digital scenes using computers and combining them with specialized equipment (such as VR headsets, controllers, and motion sensors) to simulate various human sensory experiences, including sight, hearing, and touch. This creates a psychologically and physiologically immersive experience for the user, as if they were in an interactive virtual world. Vision: Dual-screen display technology in the headset provides binocular parallax, simulating stereoscopic vision. Hearing: Spatial audio technologies such as HRTF (Head-Related Transfer Function) enable sound to achieve three-dimensional positioning. Haptic (in some devices): Force feedback controllers, haptic gloves, and other devices simulate the pressure and vibration of touching objects. Virtual reality technology is widely used in gaming, education and training (such as virtual surgery simulation), medical rehabilitation, industrial simulation (such as virtual assembly), and virtual social interaction, providing users with experiences that transcend the limitations of physical space.
[0239] See Figure 9 , Figure 9 This is a schematic diagram of the principle of the audio rendering method provided in the embodiments of this application. Figure 1 The embodiments of this application are mainly implemented in the following ways: audio encoding and rich media information collection ( Figure 9 The audio encoding and rich media information collection shown (including location information, direction information, and audio features) (i.e., the physical and virtual location information described above), and the sound source arrangement in the VR scene ( Figure 9 The sound source setup shown includes creating audio objects and adjusting audio object properties, and spatial audio rendering. Figure 9 The spatial audio rendering shown includes audio rendering adjustments and head detection and positioning. The process described above will be explained in detail below.
[0240] In some embodiments, see Figure 9Regarding audio encoding and rich media information collection, in the audio encoding stage, this application embodiment not only focuses on the encoding quality of the audio itself, but also emphasizes the accurate representation of the sound source location. To this end, rich media is introduced to describe the sound source location. This rich media information is a comprehensive data description that records in detail the sound source's position, direction, and other possible audio characteristics in three-dimensional space. Position information: The sound source's position in virtual space is accurately represented using three-dimensional coordinates (X, Y, Z). This provides accurate positioning data for subsequently placing sound sources in the VR scene. Direction information: Besides position, the direction of the sound source is also very important. For example, a speaker facing the user and a speaker facing away from the user, even if their positions are the same, will provide completely different auditory experiences. Therefore, the main sound direction of the sound source is recorded to simulate realistic sound propagation effects during rendering. Audio characteristics: This includes the sound source's volume, pitch, timbre, etc. This information helps us more realistically simulate the sound source's performance in virtual space. For example, a deep drum sound and a sharp flute sound, even if they are in the same position and direction, need to have their audio characteristics considered during rendering to achieve a more realistic auditory effect.
[0241] In some embodiments, see Figure 9 For the placement of sound sources in a VR scene, after obtaining the rich media information of the sound sources, the next step is to place the sound sources in the VR scene based on this information. Creating audio objects: First, a corresponding audio object (i.e., the virtual sound source described above) is created for each sound source in the virtual space of the VR scene. These audio objects are positioned based on the location data in the rich media information. Adjusting audio object properties: Next, the properties of these audio objects are adjusted based on the direction data and audio characteristics in the rich media information. For example, the sound direction, volume, and timbre of the audio objects (i.e., the virtual sound sources described above) are set to ensure that their performance in the virtual space matches the sound sources in the real world (i.e., the physical sound sources described above).
[0242] In some embodiments, see Figure 9For spatial audio rendering, when a user wears a VR device and enters a virtual scene, our system renders spatial audio in real time based on the user's head position and orientation, as well as the location and attributes of the sound source. Head detection and positioning: The system detects the user's head position and orientation in real time. This data is crucial for dynamically adjusting the audio rendering effect. Audio rendering adjustment: Based on the user's head position and orientation, as well as the location and attributes of the sound source, the system dynamically adjusts the audio rendering effect. For example, when the user turns towards a sound source, the system increases the volume of that sound source and adjusts the balance of the left and right channels to simulate the effect of sound coming from the sound source. Simultaneously, it adjusts the reverberation effect based on the distance to the sound source to simulate the attenuation effect of sound propagation in the air.
[0243] In some embodiments, the location of the sound source is given in rich media form during audio encoding, and the speaker (i.e., the virtual sound source described above) is positioned in the VR scene according to the rich media information to generate spatial audio. The following describes the collection and application of sound source location information in rich media form. First, a multi-dimensional data acquisition system is implemented, including three-dimensional coordinate positioning: the absolute position (after conversion) of the sound source in virtual space is accurately recorded using a Cartesian coordinate system (X, Y, Z), with the error controllable to the centimeter level. For example, the position coordinates of a gunshot in a game (10.5, 2.3, 1.8) can correspond to a specific location in a real scene. Direction vector description: the direction of sound source emission is represented by unit vectors (dx, dy, dz). For example, the direction vector of a speaker facing the user is (0, 0, 1), and that facing away from the user is (0, 0, -1). Combined with head detection data, the directional difference of sound can be simulated. Audio characteristic matrix: a multi-dimensional audio feature vector is constructed, including volume (0.0-1.0), pitch (Hz), timbre (spectral feature parameters), etc. For example, the timbre of a bass drum can be described by parameters such as the proportion of low-frequency energy (70% of the energy in the 60-250Hz frequency band) and the formant frequencies (500Hz, 1200Hz).
[0244] In some embodiments, the data structure and encoding implementation in this application are designed using the AudioSource class: the audio source data is encapsulated using a Python class, as shown in the pseudocode below:
[0245]
[0246]
[0247] Doppler effect coefficient = 1.0 / / Used to adjust the Doppler frequency deviation intensity
[0248] In some embodiments, the data transmission protocol is:
[0249] In VR devices, rich media data is transmitted via the UDP protocol, using a custom data packet format.
[0250] |Header (8B)||Timestamp (8B)||Location data (12B)||Direction data (12B)||Audio features (16B)||Audio data (NB)|.
[0251] Application example: In a virtual concert scenario, the lead singer's position coordinates are updated in real time as they move on the stage (e.g., from (5, 0, 1.7) to (3, 2, 1.7)). At the same time, the microphone's directivity (cardioid) is simulated by the direction vector (0, 0, 1) and the attenuation coefficient (30dB attenuation at the rear), and combined with the timbre parameters (3dB boost in the mid-frequency) to achieve a realistic sound field reproduction.
[0252] In some embodiments, the precise sound source arrangement based on rich media information is described below. For the virtual space (i.e., the virtual scene described above) mapping mechanism, the coordinate mapping algorithm establishes a coordinate transformation matrix between the virtual space and the physical space, supporting translation, rotation, and scaling transformations. For example, the center of the VR glasses' field of view corresponds to the point (0, 0, 0) in the virtual space, and the user's head rotation angle θ corresponds to the virtual space Y-axis rotation matrix. The pseudocode is shown below:
[0253]
[0254] In some embodiments, a scene graph structure is used to manage the hierarchy of audio objects, supporting parent-child relationship binding. For example, the sound of a car engine is the parent object, and the sound of wheel friction is the child object. The position is automatically updated as the car moves. Audios bound in the parent-child relationship are synthesized synchronously to ensure the realism between different sounds.
[0255] In some embodiments, for the dynamic attribute adjustment strategy, real-time volume adjustment is based on the distance attenuation formula V = V0 × (d0 / d)α (V0 is the initial volume, d0 is the reference distance, d is the current distance, and α is the attenuation index, which is 1.5 for indoor scenes and 2.0 for outdoor scenes). For example, when the user is 1 meter away from the sound source, the volume is 0.8, and when the user moves to 2 meters, it automatically attenuates to 0.4.
[0256] Dynamic tone correction: According to the Doppler effect formula f′=f×(v+vr) / (v+vs) (f is the original frequency, v is the speed of sound, vr is the receiver's speed, and vs is the speed of the sound source). When a user approaches a 440Hz sound source at a speed of 10m / s, the perceived frequency becomes 440×(340+10) / (340+0)≈453Hz.
[0257] In some embodiments, for VR scene integration implementation, a Unity engine integration example is: associating the AudioSource component with spatial data via C# script.
[0258] In some embodiments, for real-time spatial audio rendering algorithms, for head detection and data fusion, multi-sensor fusion is used: combining data from gyroscopes, accelerometers, and magnetometers, a complementary filter algorithm is employed to eliminate drift, with an update frequency exceeding 100Hz. The short-term, precise dynamic angles from the gyroscope are fused with the long-term, stable absolute angles from the accelerometer and magnetometer, complementing each other's shortcomings and outputting a more reliable attitude angle. The fusion formula is as follows:
[0259] θ filtered = (1-α)×θ gyro +α×θ accel / mag (3)
[0260] Where α is the fusion coefficient, θ filtered The fused attitude angle, θ, is used to indicate the output after complementary filtering and fusion. It is the result of fusing gyroscope and accelerometer / magnetometer data, making the attitude angle more stable and accurate. gyro θ is used to indicate the attitude angle measured / calculated by the gyroscope. accel / mag Used to indicate the attitude angle calculated by the accelerometer or magnetometer.
[0261] In some embodiments, the position detection technology supports both Inside-out (visual positioning) and Outside-in (laser positioning) modes, with a position error of less than 2cm. For example, the HTC Vive's laser positioning system calculates the VR headset's position coordinates by emitting infrared signals from a base station.
[0262] In some embodiments, for the core audio rendering algorithm, HRTF (Head-Related Transfer Function) rendering involves interpolating the left and right ear transfer functions from the HRTF database based on personalized parameters such as the user's head size and ear shape. Example HRTF data structure:
[0263]
[0264]
[0265] Ambisonics rendering: The three-dimensional sound field is decomposed into spherical harmonic functions and reconstructed using FOA (first-order Ambisonics) or HOA (higher-order Ambisonics). The FOA signal decoding formula is shown below:
[0266] L = W + Y (4)
[0267] R = WY (5)
[0268] Where W indicates the center channel, Y is the Y-axis component, L corresponds to the left ear, and R corresponds to the right ear.
[0269] In some embodiments, for real-time rendering engine implementations, the core logic of the SpatialAudioRenderer class is: responsible for managing all audio objects and user head states, and performing audio rendering. Data members include: audio_objects: stores all sound source objects, each containing attributes such as position and volume. user_head: detects the position (3D coordinates) and rotation (quaternion representation) of the user's head. hrtf_lib: HRTF database, providing filter coefficients for the left and right ears based on azimuth and elevation angles. ambisonics_order: the order of Ambisonics (higher-order sound field coding), defaulting to order 1. buffer_size and sample_rate: the size of the audio buffer and the sampling rate, determining the precision and performance of audio processing. The core rendering process (render_frame method) initializes the output buffer: creates a dual-channel (left and right ear) buffer to store the rendered audio data. Iterates through all sound sources: performs the following processing for each sound source: calculates the relative position: based on the user's head position, calculates the relative coordinates and distance of the sound source. Calculate azimuth and elevation angles: Convert the relative position to spherical coordinates (azimuth and elevation angles) for HRTF filter lookup. Obtain HRTF filters: Retrieve the left and right ear filter coefficients for the corresponding azimuth from the HRTF library to simulate the human auditory system's perception of differences in sound from different directions. Apply distance attenuation: Calculate volume attenuation based on distance to achieve the auditory effect of sounding louder when closer and softer when farther away. Doppler effect processing: Adjust the audio pitch based on the relative motion between the sound source and the user to simulate the Doppler effect in the real world. Convolution filtering: Convolve the sound source signal with the HRTF filter to achieve spatial audio effects. Superimpose onto output: Superimpose the processed left and right channel signals onto the total output buffer. Limiting processing: Limit the output signal to prevent audio signal overload and distortion.
[0270] In some embodiments, for lightweight rendering techniques, the importance of sound sources is prioritized based on factors such as the distance between the sound source and the user, the line-of-sight angle, and the audio energy. For distant sound sources (>10 meters), rendering precision is reduced, for example, by using a simplified HRTF model or Ambisonics downscaling. Audio pre-computation techniques pre-calculate the HRTF response of static sound sources (such as ambient sound) at different locations, and accelerate rendering at runtime by looking up tables, reducing real-time computation.
[0271] In some embodiments, for latency optimization schemes, network transmission optimization involves using the RTP protocol to transmit audio data and setting a jitter buffer to dynamically adjust the latency, ensuring end-to-end latency is <20ms. Rendering pipeline optimization utilizes the GPU audio processing unit to offload CPU computational load, achieving audio processing latency of less than 1ms.
[0272] In some embodiments, for cross-platform adaptation solutions, the hardware abstraction layer is designed to adapt the head detection data and audio output interfaces of different VR devices (Oculus Rift, HTC Vive, Pico, etc.) through a unified interface. The pseudocode for an example interface definition is shown below:
[0273]
[0274] In some embodiments, for technical verification and performance indicators, subjective evaluation tests, ABX test results: In a virtual conference room scenario, compared with a traditional stereo solution, 92% of testers could correctly distinguish the direction of the sound source using this solution, while only 65% of testers could do so using the traditional solution. Immersion rating: Using a 5-point rating system, this solution achieved an average score of 4.7 points, significantly higher than the traditional head-detection-based solution (3.2 points). For objective performance indicators, computational resource consumption: On a medium-configuration PC, when rendering 10 sound sources, CPU utilization was <15%, and GPU utilization was <20%. Latency indicators: End-to-end latency from head movement to audio response was <15ms, meeting the real-time requirements of VR scenarios (typically <20ms). Spatial positioning accuracy: Sound source azimuth error was <3°, and distance perception error was <5%, meeting the accuracy requirements of most VR applications.
[0275] In this way, through rich media information collection and precise sound source placement, accurate positioning and realistic performance of sound sources in virtual space are achieved. Simultaneously, combined with real-time spatial audio rendering algorithms and advanced audio processing technologies, a lifelike and immersive sound experience is provided to users.
[0276] It is understood that in the embodiments of this application, data related to physical location information is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0277] The following description continues to illustrate the exemplary structure of the audio rendering device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2As shown, the software modules stored in the audio rendering device 455 in the memory 450 may include: an acquisition module 4551, used to acquire the physical orientation information of the physical sound source in the physical scene; a conversion module 4552, used to convert the physical orientation information of the physical sound source based on the virtual objects included in the virtual scene, to obtain the virtual orientation information corresponding to the physical orientation information in the virtual scene; a creation module 4553, used to create a virtual sound source corresponding to the physical sound source in the virtual scene based on the virtual orientation information; and a rendering module 4554, used to perform audio rendering on the sound signal generated by the physical sound source based on the virtual sound source, to obtain rendered audio, and to determine the virtual audio to be sent to the virtual object based on the rendered audio.
[0278] In some embodiments, the conversion module 4552 is further configured to obtain the virtual scene coordinate system of the virtual scene, wherein the virtual scene coordinate system takes the center position of the virtual scene as the origin; based on the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system, the virtual scene coordinate system is converted to obtain the object coordinate system with the virtual object as the origin; the transformation matrix between the object coordinate system and the physical scene coordinate system of the physical scene is determined, and based on the transformation matrix, the physical orientation information of the physical sound source is converted to obtain the virtual orientation information.
[0279] In some embodiments, the conversion module 4552 is further configured to determine the difference vector between the coordinates of the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system as the offset vector between the virtual scene coordinate system and the object coordinate system; and to perform a translation transformation on the virtual scene coordinate system according to the offset vector to obtain an object coordinate system with the virtual object as the origin.
[0280] In some embodiments, the transformation matrix includes a position transformation matrix and a direction transformation matrix. The transformation module 4552 is further configured to determine the position transformation matrix based on the position coordinates of the origin of the object coordinate system in the physical scene coordinate system; determine the unit direction vector of each coordinate axis in the physical scene coordinate system for each coordinate axis of the object coordinate system; and construct a transformation matrix between the object coordinate system and the physical scene coordinate system based on the unit direction vector of each coordinate axis.
[0281] In some embodiments, the physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system and the physical direction vector of the physical sound source in the physical scene coordinate system. The transformation matrix includes a position transformation matrix and a direction transformation matrix. The virtual orientation information includes virtual position coordinates and a virtual direction vector. The transformation module 4552 is further configured to determine the virtual position coordinates by multiplying the physical position coordinates by the position transformation matrix and to determine the virtual direction vector by multiplying the physical direction vector by the direction transformation matrix.
[0282] In some embodiments, the virtual orientation information includes virtual position coordinates and virtual direction vectors. The creation module 4553 is further configured to create an initial virtual sound source corresponding to the physical sound source at the virtual position coordinates in the virtual scene; and adjust the initial virtual sound source based on the virtual direction vector to obtain the virtual sound source.
[0283] In some embodiments, the creation module 4553 is further configured to adjust the orientation of the initial sound source to a direction consistent with the virtual direction vector to obtain a candidate virtual sound source; obtain the acoustic feature parameters of the physical sound source, and adjust the acoustic feature parameters of the physical sound source based on the virtual scene to obtain adjusted acoustic feature parameters; set the acoustic feature parameters of the candidate virtual sound source to the adjusted acoustic feature parameters to obtain the virtual sound source.
[0284] In some embodiments, the number of physical sound sources is at least one, and the virtual sound sources correspond one-to-one with the physical sound sources; the rendering module 4554 is further configured to, for each virtual sound source, perform audio rendering on the sound signal generated by the corresponding physical sound source based on the virtual sound source to obtain the rendered audio corresponding to the virtual sound source; the rendering module is further configured to, when the number of virtual sound sources is multiple, synthesize the rendered audio corresponding to multiple virtual sound sources to obtain the virtual audio; when the number of virtual sound sources is one, determine the rendered audio corresponding to the virtual sound source as the virtual audio.
[0285] In some embodiments, the rendering module 4554 is further configured to cluster the virtual sound sources to obtain at least one virtual sound source group, wherein the virtual audio corresponds one-to-one with the virtual sound source group; and for each virtual sound source group, to synthesize the rendered audio corresponding to each virtual sound source in the virtual sound source group to obtain the virtual audio corresponding to the virtual sound source group.
[0286] In some embodiments, the rendering module 4554 is further configured to determine the relative orientation information between the virtual sound source and the virtual object based on the virtual orientation information corresponding to the physical orientation information in the virtual scene; and to perform audio rendering on the sound signal generated by the physical sound source based on the relative orientation information and the acoustic feature parameters of the virtual sound source to obtain the rendered audio.
[0287] In some embodiments, the rendering module 4554 is further configured to match the sound signal generated by the physical sound source with the acoustic feature parameters to obtain a matching result; when the matching result indicates that the sound signal does not match the acoustic feature parameters, adjust the acoustic features of the sound signal based on the acoustic feature parameters to obtain a candidate sound signal; and perform audio rendering on the candidate sound signal based on the relative orientation information to obtain the rendered audio.
[0288] In some embodiments, the virtual object includes multiple audio receiving units, the relative orientation information includes relative orientation sub-information between the virtual sound source and each audio receiving unit, and the rendered audio includes sub-rendered audio corresponding to each audio receiving unit; the rendering module 4554 is further configured to perform audio rendering on the candidate sound signal for each audio receiving unit based on the relative orientation sub-information to obtain sub-rendered audio corresponding to each audio receiving unit; when the number of virtual sound sources is multiple, the rendering module 4554 is further configured to synthesize the sub-rendered audio corresponding to each virtual sound source for each audio receiving unit to obtain the virtual audio of the audio receiving unit.
[0289] In some embodiments, the rendering module 4554 is further configured to acquire the acoustic characteristics of the candidate sound signal, and determine the target acoustic characteristics of the rendered audio based on the relative orientation information between the virtual sound source and the virtual object and the acoustic characteristics of the candidate sound signal; and perform audio rendering on the candidate sound signal based on the target acoustic characteristics to obtain the rendered audio.
[0290] In some embodiments, the rendering module 4554 is further configured to adjust the acoustic characteristics of the candidate sound signal based on the relative orientation information to obtain candidate rendered audio; select a target filter coefficient corresponding to the virtual sound source from a plurality of preset filter coefficients based on the relative orientation information; and perform audio rendering on the candidate rendered audio based on the target filter coefficient to obtain the rendered audio.
[0291] In some embodiments, the rendering module 4554 is further configured to: determine the relative orientation information between the virtual sound source and the virtual object based on the virtual orientation information corresponding to the physical orientation information in the virtual scene; determine the rendering precision of the sound signal based on the relative orientation information between the virtual sound source and the virtual object; and perform audio rendering on the sound signal generated by the physical sound source according to the rendering precision of the sound signal based on the virtual sound source to obtain the rendered audio.
[0292] In some embodiments, the rendering module 4554 is further configured to obtain the sound source type of the physical sound source, compare the sound source type with a preset sound source type, and obtain a comparison result; when the comparison result indicates that the sound source type is the preset sound source type, based on the virtual orientation information, search for the preset rendering audio corresponding to the virtual orientation information in the audio library of the preset sound source type, and determine the found preset rendering audio as the rendering audio; the rendering module is further configured to perform audio rendering on the sound signal generated by the physical sound source based on the virtual sound source when the comparison result indicates that the sound source type is not the preset sound source type, to obtain the rendering audio.
[0293] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. An electronic device's processor reads the computer-executable instructions or computer program from the computer-readable storage medium or the computer program, and executes the computer-executable instructions or computer program, causing the electronic device to perform the audio rendering method described above in this application.
[0294] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the audio rendering method provided in this application. For example, ... Figure 3 The audio rendering method is shown.
[0295] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of electronic devices including one or any combination of the above-mentioned memories.
[0296] In some embodiments, computer-executable instructions or computer programs may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0297] As an example, computer executable instructions or computer programs may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., files that store one or more modules, subroutines, or code sections).
[0298] As an example, computer-executable instructions or computer programs may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected by a communication network.
[0299] In summary, the embodiments of this application have the following beneficial effects:
[0300] (1) By acquiring the physical orientation information of the physical sound source in the physical scene and converting the physical orientation information based on the virtual objects included in the virtual scene, virtual orientation information is obtained. This ensures that the spatial relationships of the physical world are accurately mapped in the virtual scene, so that the position of the virtual sound source and the actual orientation of the physical sound source form a logical correspondence, laying the foundation for the spatial realism of the virtual audio. Based on the virtual orientation information corresponding to the physical orientation information in the virtual scene, a virtual sound source corresponding to the physical sound source is created in the virtual scene. Using the virtual sound source as a reference, the sound signal generated by the physical sound source is rendered to obtain the rendered audio. This allows the rendering process to fully integrate the spatial characteristics of the virtual scene (such as the distance and angle between the virtual sound source and the virtual object), so that the rendered audio naturally carries the acoustic characteristics that conform to the virtual orientation (such as volume attenuation at a distance and spectral differences at a specific angle). By determining the virtual audio to be sent to the virtual object based on the rendered audio, and ensuring that the entire processing revolves around the perception logic of the virtual object, the final generated virtual audio not only retains the original sound characteristics of the physical sound source, but also accurately restores its spatial orientation relative to the virtual object in the virtual scene. This makes the sound received by the virtual object highly consistent with the real auditory experience in terms of spatial attributes such as orientation and distance, thereby significantly improving the simulation degree of the virtual audio.
[0301] (2) The difference vector between the coordinates of the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system is determined as the offset vector. Based on this, the virtual scene coordinate system is translated to obtain the object coordinate system. Through explicit offset vector calculation and coordinate system translation, a local spatial reference system centered on the virtual object can be quickly established in the virtual scene. This simplifies the description and calculation of the relative positional relationship between the virtual object and other virtual elements (such as virtual sound sources), avoids the complex coordinate calculations that may occur when directly using the global coordinate system, significantly reduces the computational load on the server when processing spatial interactions (such as virtual sound source orientation judgment, collision detection, etc.), and improves operating efficiency. It ensures the precise association between the object coordinate system and the virtual scene coordinate system. When the virtual object moves in the virtual scene, the server only needs to update the object coordinate system in real time by updating the offset vector, ensuring the dynamic consistency between local coordinates and global coordinates. This allows the spatial logic based on the object coordinate system (such as spatial rendering of virtual audio and behavioral interaction of virtual objects) to be adjusted in real time with the position changes of the virtual object, enhancing the accuracy and real-time performance of spatial relationships in the virtual scene.
[0302] (3) The transformation matrix is decomposed into a position transformation matrix and a direction transformation matrix. The position transformation matrix is determined by the position coordinates of the origin of the object coordinate system in the physical scene coordinate system. At the same time, the overall transformation matrix is constructed based on the unit direction vectors of each coordinate axis of the object coordinate system in the physical scene coordinate system. The decomposition method decomposes the complex spatial transformation into two independent dimensions: position offset and direction rotation. This reduces the complexity of matrix calculation, enabling the server or processing system to complete the transformation between coordinate systems more efficiently and reducing the consumption of computing resources. By clearly defining the position of the origin of the object's coordinate system and the direction vectors of each coordinate axis, the spatial relationship between the two coordinate systems can be accurately captured, ensuring the accuracy of coordinate transformation and avoiding the accumulation of errors caused by the coupling of the overall transformation matrix parameters. This provides a reliable mathematical foundation for the accurate mapping between virtual and physical spaces (such as matching the position of a virtual sound source with the physical scene). It has good scalability and real-time performance. When the position of the origin or the direction of the coordinate axes of the object's coordinate system changes (such as when a virtual object moves or rotates), the overall transformation matrix can be quickly reconstructed simply by updating the position transformation matrix or the direction vector, without having to recalculate the entire matrix. This ensures the efficiency and consistency of coordinate system transformation in dynamic scenes (such as VR interaction and real-time simulation), laying a solid foundation for improving the immersion and interaction accuracy of virtual scenes.
[0303] (4) The virtual position coordinates are obtained by multiplying the physical position coordinates with the position transformation matrix, and the virtual direction vector is obtained by multiplying the physical direction vector with the direction transformation matrix. The conversion from physical spatial parameters to virtual spatial parameters is achieved through matrix operations. With the specific functions of the position transformation matrix and the direction transformation matrix, the two types of spatial transformations, position translation and direction rotation, can be accurately separated and processed, avoiding conversion errors caused by parameter confusion. This ensures that the virtual position coordinates and virtual direction vectors can accurately reflect the relative position and orientation of the physical sound source in the virtual scene, providing a reliable spatial data foundation for subsequent processing such as virtual audio rendering and virtual object interaction. The matrix-based conversion method has mathematical rigor and efficiency. When the spatial state of the physical scene or virtual object changes (such as the physical sound source moving or the virtual object rotating), only the corresponding transformation matrix needs to be updated to quickly recalculate the virtual parameters without reconstructing the entire conversion logic. This significantly improves the real-time performance and flexibility of spatial parameter mapping in dynamic scenes, thereby enhancing the immersion and interactive response speed of the virtual scene.
[0304] (5) The orientation of the initial sound source is adjusted to be consistent with the virtual direction vector to obtain candidate virtual sound sources. Then, the acoustic characteristic parameters of the physical sound source are adjusted in combination with the virtual scene and assigned to the candidate virtual sound source to obtain the final virtual sound source. By adjusting the orientation, it is ensured that the orientation of the virtual sound source in the virtual scene corresponds precisely with the actual orientation of the physical sound source in the physical scene. This ensures that the spatial directivity of the sound in the virtual environment is consistent with that in the real world, avoiding the destruction of immersion caused by directional confusion. At the same time, the targeted adjustment of acoustic characteristic parameters based on the virtual scene allows the volume, timbre, reverberation and other characteristics of the virtual sound source to adapt to the environmental characteristics of the virtual scene (such as the spaciousness of the virtual hall and the closedness of the virtual room). This preserves the essential acoustic characteristics of the physical sound source and conforms to the acoustic logic of the virtual scene. This step-by-step optimization method ultimately achieves a natural and accurate mapping of the physical sound source to the virtual scene, making the virtual sound highly integrated with the virtual environment in terms of spatial orientation and acoustic characteristics. This greatly enhances the user's auditory immersion and experience realism in the virtual scene.
[0305] (6) An initial virtual sound source corresponding to the physical sound source is created at the virtual position coordinates in the virtual scene, and adjusted based on the virtual direction vector to obtain the final virtual sound source. By creating the initial virtual sound source at the precisely calculated virtual position coordinates, it is ensured that the spatial position of the physical sound source in the physical scene and the mapped position in the virtual scene are strictly corresponding, laying the foundation for the spatial positioning of virtual sound. The adjustment based on the virtual direction vector further ensures that the orientation of the virtual sound source is consistent with the actual sound emission direction of the physical sound source in the virtual scene, so that the radiation characteristics of the sound conform to the logic of real space. This two-step construction method not only realizes the accurate coordinate mapping from physical space to virtual space, but also ensures the real restoration of the sound direction characteristics. In the end, the virtual sound source can be accurately positioned in the virtual scene and correctly transmit direction information, providing users with a virtual auditory experience consistent with the acoustic laws of the physical scene, effectively enhancing the immersion and spatial realism of the virtual scene.
[0306] (7) When there is at least one physical sound source and the virtual sound source corresponds one-to-one with the physical sound source, and audio rendering is performed on the sound signal of the corresponding physical sound source for each virtual sound source, the one-to-one correspondence ensures that the sound characteristics of each physical sound source can be independently and accurately mapped in the virtual scene, avoiding auditory confusion caused by the confusion of multiple sound source signals. For example, the voices of speech and musical instruments in different physical locations can still maintain their own spatial positioning and acoustic characteristics in the virtual scene. Rendering each virtual sound source separately can be personalized according to its specific position, orientation and virtual environment characteristics (such as distance from the audience and material of the virtual space). For example, the sound of distant sound sources can be made weaker and have more reverberation, while the sound of nearby sound sources can be made clearer and have a stronger sense of direction. This allows multiple rendered audios to form a sound field with distinct layers and conforming to the logic of real space in the virtual scene. Ultimately, users can accurately distinguish the position and characteristics of different physical sound sources in the virtual scene, which greatly improves the realism and immersion of the virtual auditory experience in multi-sound source scenarios.
[0307] (8) The distance between the virtual sound source and the virtual object is determined by the coordinate distance between the virtual and physical position coordinates, and the pitch angle between the virtual and physical direction vectors is determined by the angle between them. By directly linking the physical and virtual position coordinates and direction vectors, it is ensured that the relative orientation information can accurately reflect the real spatial relationship between the sound source and the object in the physical space, avoiding distance or angle distortion caused by complex conversion logic. For example, a sound source and an object that are physically 5 meters apart can maintain a relative distance of 5 meters in the virtual scene, and a sound source that is physically at a 30° pitch angle is also presented at a 30° pitch angle in the virtual scene. Distance and pitch angle, as the core of relative orientation information, provide clear and accurate spatial parameters for subsequent audio rendering, enabling the rendered audio to adjust the volume attenuation according to the distance and optimize the spectral characteristics according to the pitch angle, so that the sound orientation heard by the user in the virtual scene is highly consistent with the physical scene, greatly enhancing the realism of the auditory experience and the sense of spatial immersion.
[0308] (9) Match the sound signal of the physical sound source with the acoustic feature parameters. If there is a mismatch, adjust to obtain candidate sound signals. Then, combine the relative orientation information for audio rendering. Through matching and adjustment, it can be ensured that the sound signal is consistent with the virtual scene settings in terms of basic acoustic features (such as volume, timbre, noise reduction, etc.). This avoids the destruction of the acoustic logic of the virtual scene due to deviations in the original signal of the physical sound source (such as excessive volume or noise). This provides a clean audio foundation that meets the requirements for subsequent rendering. Rendering based on relative orientation information further endows the sound with spatial attributes, so that the adjusted sound can present a listening experience that conforms to spatial rules based on parameters such as distance and pitch angle (such as distant sounds being weaker and high-frequency sounds being more prominent above). This two-step processing method of first calibrating the basic features and then endowing spatial attributes not only ensures the standardization of the acoustic features of the virtual sound source, but also realizes the accurate spatial positioning of the sound in the virtual scene. In the end, the rendered audio not only fits the virtual scene settings, but also has a real sense of spatial immersion, which greatly improves the realism of the user's auditory experience in the virtual environment.
[0309] (10) When a virtual object contains multiple audio receiving units, and the final rendered audio is generated by rendering sub-rendered audio for each unit based on the corresponding relative orientation sub-information, the multiple audio receiving units simulate the characteristics of multiple receiving points (such as human binaural ears, microphone arrays) in a real auditory system. The dedicated rendering for each unit can accurately restore the spatial differences of the virtual sound source relative to different receiving points. For example, different receiving units have different distances and angles from the virtual sound source, and the corresponding sub-rendered audio will show subtle differences in volume, spectrum, time difference, etc. This difference is like the natural difference between the sounds heard in different locations in a real scene. The combination of these sub-rendered audios allows the virtual object to more accurately locate the spatial position of the virtual sound source through multi-dimensional auditory information, which greatly enhances the stereo sense and directional recognition of virtual hearing. At the same time, this unit-based processing method also provides flexible support for the differentiation of multiple sound sources and the realization of surround sound effects in complex virtual scenes, ultimately allowing users to obtain a spatial auditory experience that is highly consistent with the real world in the virtual environment and improving the overall immersion.
[0310] (11) The determination of the target acoustic characteristics fully integrates the original acoustic features of the candidate signal (such as timbre, basic volume, etc.) with the location information of the virtual space (such as attenuation caused by distance, and spectral changes caused by angle), ensuring that the target characteristics do not deviate from the essential attributes of sound, and can accurately reflect its spatial position in the virtual scene, avoiding the problem of losing the characteristics of the sound source itself due to simply emphasizing spatial attributes. By directly adjusting the candidate signal to the target characteristics, the rendering process has a clear direction and standard, and can efficiently and accurately achieve the adaptation of acoustic characteristics. For example, it can ensure that the bright timbre of the piano sound is not destroyed, and can reflect its spatial sense from 8 meters away and above by adjusting parameters such as volume and reverberation. This allows the rendered audio to retain the true characteristics of the sound source while perfectly integrating into the spatial logic of the virtual scene, ultimately presenting users with an auditory experience that is both realistic and in line with the virtual environment setting, effectively improving the naturalness and immersion of virtual audio.
[0311] (12) Adjusting the acoustic characteristics of candidate sound signals based on relative orientation information (such as volume attenuation, spectrum adaptation, and basic reverberation addition) can quickly make the sound conform to the macroscopic orientation characteristics of the virtual space, laying a foundation for subsequent processing that conforms to spatial logic. Selecting target filter coefficients from preset filter coefficients (such as filter coefficients corresponding to the orientation in the HRTF library) and utilizing pre-measured acoustic data can capture subtle differences in sound perception caused by human physiological structure (such as head and auricle) under different orientations, making up for the shortcomings of basic adjustments in detail restoration; finally, secondary rendering of candidate audio through target filter coefficients can further integrate subtle acoustic characteristics unique to the orientation (such as volume difference, time difference, and spectrum difference between the left and right ears) into the sound in addition to macroscopic spatial attributes. This ensures the overall matching of sound with the orientation of the virtual space and restores subtle differences in real hearing through professional filter coefficients, ultimately making the rendered audio have both spatial realism and auditory delicacy, greatly improving the accuracy of sound positioning in the virtual scene and the user's auditory immersion.
[0312] (13) The physical-to-virtual orientation mapping ensures that the spatial relationship of sound in the virtual scene is highly consistent with the physical world, avoiding the disconnect between virtual hearing and physical reality. This lays a realistic spatial foundation for subsequent rendering. The rendering precision is dynamically determined based on relative orientation information (e.g., high precision for close distances and adaptive precision for long distances). This not only restores the subtle differences in sound orientation in key scenes (e.g., close-range interaction) through fine processing, enhancing the sense of auditory realism, but also reduces system resource consumption by appropriately simplifying in non-critical scenes, achieving a balance between effect and efficiency. Combining the virtual sound source attributes with corresponding precision for rendering ensures that the processed audio not only fits the virtual scene settings (e.g., the virtual lecturer's voice needs to be clearly distinguishable) but also accurately reflects its spatial location. This allows users to perceive the accurate orientation of the sound in the virtual environment and obtain an auditory experience that conforms to the scene logic, ultimately significantly enhancing the immersion and efficiency of virtual hearing.
[0313] (14) By comparing the physical sound source type with the preset sound source type, high-frequency sound sources with stable characteristics (such as common office equipment sounds, ambient sounds, etc.) can be quickly identified. For these sound sources, the corresponding virtual location rendering audio in the preset audio library can be directly called, which can greatly reduce the computing power required for real-time rendering and improve the system processing efficiency. Especially in complex virtual scenes with multiple sound sources running concurrently, it can effectively avoid the latency problem caused by real-time computing overload. The preset rendering audio is finely tuned in advance for specific types and locations. Its acoustic characteristics (such as volume attenuation, reverberation adaptation, spectrum distribution, etc.) are more in line with the listening logic of the virtual scene. Compared with the parameter deviation that may occur in real-time rendering, it can ensure the consistency of the rendering effect of the same type of sound source in the same location and enhance the stability of the virtual listening experience.
[0314] (15) When the sound source type does not belong to the preset sound source type, the rendered audio is obtained by real-time audio rendering of the sound signal of the physical sound source based on the virtual sound source. For non-preset type sound sources with varied characteristics, infrequent occurrences, or those that need to accurately reproduce unique features (such as personalized human voices, special instrument sounds, etc.), real-time rendering can avoid sound distortion or insufficient adaptability to the virtual scene caused by relying on fixed preset audio, ensuring that the unique acoustic features of these sound sources (such as the unique timbre of an individual's voice, the exclusive sound quality of a special instrument) are fully preserved. At the same time, real-time processing with the virtual sound source as a reference can accurately match the spatial attributes (such as relative orientation, environmental characteristics, etc.) of the sound signal in the virtual scene, so that the rendered audio not only fits the logical settings of the virtual scene, but also truly reflects the characteristics of the sound source itself. Thus, on the basis of covering more diverse sound source types, the authenticity, uniqueness and scene adaptability of the virtual auditory experience are guaranteed, further enhancing the overall immersion.
[0315] (16) When there are multiple virtual sound sources, by synthesizing the rendered audio of each source (such as weight adjustment, time synchronization, and frequency optimization), multi-directional and multi-type sound information can be organically integrated. This preserves the spatial positioning and acoustic characteristics of each sound source (such as the clear dominance of lecture sound, the close-range details of book turning sound, and the distant background of bird chirping sound), and forms a layered and naturally coordinated overall auditory environment. This allows virtual objects to perceive rich sound information that conforms to the scene logic, avoiding auditory confusion caused by the mixing of multiple sound sources. When there is only one virtual sound source, its rendered audio can be directly determined as virtual audio. This can ensure the accuracy of the sound (such as the sense of location of a single lecture sound), while simplifying the processing flow, reducing system resource consumption, and improving efficiency.
[0316] (17) The clustering process groups multiple virtual sound sources based on spatial location, sound source type, or scene function, which greatly simplifies the complexity of multi-sound source synthesis and avoids the sound chaos caused by the direct superposition of all sound sources. This allows for more efficient processing of complex scenes (e.g., simplifying the processing of 10 independent sound sources to processing 3 groups). When synthesizing each group, exclusive rules can be formulated based on the correlation of sound sources within the group (e.g., in the same region or of the same type) (e.g., adjusting weights and optimizing frequency coordination). This ensures the synergy of sounds within the group (e.g., the coffee machine sound and human voice in the bar group blend naturally) and forms a clear hierarchical distinction through the differences in parameters between groups (e.g., the bar group is clear, and the environment group is weakened), allowing the virtual audio to present a "regionalized" and "functionalized" auditory structure. The virtual audio corresponding to multiple groups jointly constructs a logically clear and hierarchically rich virtual auditory environment, enabling virtual objects to perceive sound information from different regions or functions more naturally (e.g., distinguishing between bar service sounds and conversations at neighboring tables). This significantly improves the immersion and recognizability of virtual hearing in complex scenes while also taking into account the efficiency of processing.
[0317] (18) By processing and synthesizing sub-rendered audio corresponding to all virtual sound sources separately for each receiving unit (such as the left and right ears of a virtual human, or different microphones of a virtual device), the unique spatial relationship between each receiving unit and multiple sound sources can be accurately restored. For example, the virtual audio received by the left ear will naturally reflect the occlusion effect of the head on the right sound source, while the right ear will show the attenuation characteristics of the left sound source. This difference makes the virtual object's perception of sound closer to the auditory logic of the real physical world. Unit-based synthesis avoids the mixing of sound information from different receiving units, ensuring that the virtual audio of each unit clearly reflects its unique acoustic characteristics (such as position and sensitivity), making the sound positioning in multi-sound source scenes more accurate (such as being able to distinguish the directional difference between the voice from the right front and the footsteps from the left rear). The synergistic effect of the virtual audio of each receiving unit not only preserves the rich information of multiple sound sources, but also constructs a three-dimensional spatial auditory experience through the auditory differences between units, greatly improving the realism and immersion of sound perception in the virtual scene, while providing a reliable audio basis for virtual objects to perform precise interactions based on sound (such as turning towards the sound source).
[0318] (19) Through rich media information collection and precise sound source placement, the accurate positioning and realistic performance of sound sources in virtual space are achieved. At the same time, combined with real-time spatial audio rendering algorithms and advanced audio processing technology, users are provided with a realistic and immersive sound experience.
[0319] (20) The initial coordinate system with the center of the virtual scene as the origin provides a unified benchmark for the entire virtual space, ensuring the integrity of spatial positioning. Transforming the coordinate system into an object coordinate system with the virtual object as the origin allows the viewpoint of the virtual object to be taken as the core, which is more in line with its perception logic. This makes the subsequent orientation calculation more intuitive in reflecting the relative relationship (such as distance and angle) between the virtual sound source and the virtual object. By determining the transformation matrix between the object coordinate system and the physical scene coordinate system, a precise mathematical relationship between the physical space and the virtual space is established, avoiding orientation deviation caused by the difference in coordinate systems. This ensures that the actual orientation (such as position and direction) of the physical sound source can be accurately mapped to the virtual scene. Finally, the virtual orientation information obtained based on this transformation matrix allows the position of the virtual sound source in the virtual scene to form a logically consistent correspondence with the position of the physical sound source in the physical scene. This provides a precise spatial basis for subsequent audio rendering, making the virtual audio received by the virtual object closer to the real listening experience in terms of orientation, distance and other spatial attributes. This effectively improves the spatial realism of the virtual scene and the user's immersion.
[0320] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An audio rendering method, characterized in that, The method includes: Obtain the physical location information of the physical sound source in the physical scene; Based on the virtual objects included in the virtual scene, the physical location information of the physical sound source is converted to obtain the virtual location information in the virtual scene corresponding to the physical location information; Based on the virtual location information, a virtual sound source corresponding to the physical sound source is created in the virtual scene; Based on the virtual sound source, the sound signal generated by the physical sound source is rendered to obtain rendered audio, and based on the rendered audio, a virtual audio for sending to the virtual object is determined.
2. The method according to claim 1, characterized in that, The process of converting the physical location information of the physical sound source based on virtual objects included in the virtual scene to obtain the virtual location information corresponding to the physical location information in the virtual scene includes: Obtain the virtual scene coordinate system of the virtual scene, wherein the virtual scene coordinate system takes the center position of the virtual scene as the origin; Based on the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system, the virtual scene coordinate system is transformed to obtain an object coordinate system with the virtual object as the origin; Determine the transformation matrix between the object coordinate system and the physical scene coordinate system, and based on the transformation matrix, transform the physical orientation information of the physical sound source to obtain the virtual orientation information.
3. The method according to claim 2, characterized in that, The process of transforming the virtual scene coordinate system based on the center position of the virtual scene and the position coordinates of the virtual object in the virtual scene coordinate system to obtain an object coordinate system with the virtual object as the origin includes: The difference vector between the coordinates of the center position of the virtual scene and the position coordinates is determined as the offset vector between the virtual scene coordinate system and the object coordinate system; The virtual scene coordinate system is translated and transformed according to the offset vector to obtain an object coordinate system with the virtual object as the origin.
4. The method according to claim 2, characterized in that, The transformation matrix includes a position transformation matrix and an orientation transformation matrix. Determining the transformation matrix between the object coordinate system and the physical scene coordinate system includes: The position transformation matrix is determined based on the position coordinates of the origin of the object coordinate system in the physical scene coordinate system; For each coordinate axis of the object coordinate system, determine the unit direction vector of the coordinate axis in the physical scene coordinate system; Based on the unit direction vector of each coordinate axis, a transformation matrix is constructed between the object coordinate system and the physical scene coordinate system of the physical scene.
5. The method according to claim 2, characterized in that, The physical orientation information includes the physical position coordinates of the physical sound source in the physical scene coordinate system, and the physical direction vector of the physical sound source in the physical scene coordinate system. The transformation matrix includes a position transformation matrix and a direction transformation matrix. The virtual orientation information includes virtual position coordinates and a virtual direction vector. The process of transforming the physical location information of the physical sound source based on the transformation matrix to obtain the virtual location information includes: The virtual position coordinates are determined by multiplying the physical position coordinates by the position transformation matrix. The virtual direction vector is determined by multiplying the physical direction vector by the direction transformation matrix.
6. The method according to any one of claims 1 to 5, characterized in that, The virtual orientation information includes virtual position coordinates and a virtual direction vector. Based on the virtual orientation information, creating a virtual sound source corresponding to the physical sound source in the virtual scene includes: At the virtual location coordinates in the virtual scene, an initial virtual sound source corresponding to the physical sound source is created; Based on the virtual direction vector, the initial virtual sound source is adjusted to obtain the virtual sound source.
7. The method according to claim 6, characterized in that, The step of adjusting the initial virtual sound source based on the virtual direction vector to obtain the virtual sound source includes: The orientation of the initial sound source is adjusted to be consistent with the virtual direction vector to obtain a candidate virtual sound source; Acquire the acoustic feature parameters of the physical sound source, and adjust the acoustic feature parameters of the physical sound source based on the virtual scene to obtain the adjusted acoustic feature parameters; The acoustic feature parameters of the candidate virtual sound source are set to the adjusted acoustic feature parameters to obtain the virtual sound source.
8. The method according to any one of claims 1 to 7, characterized in that, The number of physical sound sources is at least one, and the virtual sound sources correspond one-to-one with the physical sound sources; The step of rendering audio based on the virtual sound source and the sound signal generated by the physical sound source to obtain rendered audio includes: For each virtual sound source, based on the virtual sound source, audio rendering is performed on the sound signal generated by the corresponding physical sound source to obtain the rendered audio corresponding to the virtual sound source; The step of determining the virtual audio to be sent to the virtual object based on the rendered audio includes: When there are multiple virtual sound sources, the rendered audio corresponding to the multiple virtual sound sources is synthesized to obtain the virtual audio; When there is only one virtual sound source, the rendered audio corresponding to the virtual sound source is determined as the virtual audio.
9. The method according to claim 8, characterized in that, The step of synthesizing the rendered audio corresponding to multiple virtual sound sources to obtain the virtual audio includes: The virtual sound sources are clustered to obtain at least one virtual sound source group, and the virtual audio corresponds one-to-one with the virtual sound source group; For each virtual sound source group, the rendered audio corresponding to each virtual sound source in the virtual sound source group is synthesized to obtain the virtual audio corresponding to the virtual sound source group.
10. The method according to any one of claims 1 to 9, characterized in that, The step of rendering audio based on the virtual sound source and the sound signal generated by the physical sound source to obtain rendered audio includes: Based on the physical orientation information corresponding to the virtual orientation information in the virtual scene, the relative orientation information between the virtual sound source and the virtual object is determined. Based on the relative orientation information and the acoustic feature parameters of the virtual sound source, the sound signal generated by the physical sound source is rendered to obtain the rendered audio.
11. The method according to claim 10, characterized in that, The step of rendering the sound signal generated by the physical sound source based on the relative orientation information and the acoustic feature parameters of the virtual sound source to obtain the rendered audio includes: The sound signal generated by the physical sound source is matched with the acoustic feature parameters to obtain a matching result; When the matching result indicates that the sound signal does not match the acoustic feature parameters, the acoustic features of the sound signal are adjusted based on the acoustic feature parameters to obtain a candidate sound signal; Based on the relative orientation information, the candidate sound signal is rendered to obtain the rendered audio.
12. The method according to claim 11, characterized in that, The virtual object includes multiple audio receiving units, the relative orientation information includes the relative orientation sub-information between the virtual sound source and each audio receiving unit, and the rendered audio includes the sub-rendered audio corresponding to each audio receiving unit; The step of rendering the candidate sound signal based on the relative orientation information to obtain the rendered audio includes: For each audio receiving unit, based on the relative orientation sub-information, audio rendering is performed on the candidate sound signal to obtain the sub-rendered audio corresponding to each audio receiving unit; When there are multiple virtual audio sources, determining the virtual audio to be sent to the virtual object based on the rendered audio includes: For each audio receiving unit, the sub-rendered audio corresponding to each virtual sound source of the audio receiving unit is synthesized to obtain the virtual audio of the audio receiving unit.
13. The method according to claim 11, characterized in that, The step of rendering the candidate sound signal based on the relative orientation information to obtain the rendered audio includes: Acquire the acoustic characteristics of the candidate sound signal, and determine the target acoustic characteristics of the rendered audio based on the relative orientation information between the virtual sound source and the virtual object and the acoustic characteristics of the candidate sound signal; Based on the target acoustic characteristics, the candidate sound signal is rendered to obtain the rendered audio.
14. The method according to claim 11, characterized in that, The step of rendering the candidate sound signal based on the relative orientation information to obtain the rendered audio includes: Based on the relative orientation information, the acoustic characteristics of the candidate sound signal are adjusted to obtain candidate rendered audio. Based on the relative orientation information, the target filter coefficient corresponding to the virtual sound source is selected from multiple preset filter coefficients; Based on the target filtering coefficients, the candidate audio is rendered to obtain the rendered audio.
15. The method according to any one of claims 1 to 14, characterized in that, The step of rendering audio based on the virtual sound source and the sound signal generated by the physical sound source to obtain rendered audio includes: Based on the physical orientation information corresponding to the virtual orientation information in the virtual scene, the relative orientation information between the virtual sound source and the virtual object is determined. Based on the relative orientation information between the virtual sound source and the virtual object, the rendering accuracy of the sound signal is determined; Based on the virtual sound source, the sound signal generated by the physical sound source is rendered according to the rendering precision of the sound signal to obtain the rendered audio.
16. The method according to any one of claims 1 to 15, characterized in that, Before performing audio rendering on the sound signal generated by the physical sound source based on the virtual sound source to obtain the rendered audio, the method further includes: Obtain the sound source type of the physical sound source, and compare the sound source type with a preset sound source type to obtain the comparison result; When the comparison result indicates that the sound source type is the preset sound source type, based on the virtual location information, the preset rendered audio corresponding to the virtual location information is searched in the audio library of the preset sound source type, and the found preset rendered audio is determined as the rendered audio; The step of rendering audio based on the virtual sound source and the sound signal generated by the physical sound source to obtain rendered audio includes: When the comparison result indicates that the sound source type is not the preset sound source type, the sound signal generated by the physical sound source is rendered based on the virtual sound source to obtain the rendered audio.
17. An audio rendering apparatus, characterized in that, The device includes: The acquisition module is used to acquire the physical location information of the physical sound source in the physical scene; The conversion module is used to convert the physical location information of the physical sound source based on the virtual objects included in the virtual scene, so as to obtain the virtual location information in the virtual scene corresponding to the physical location information; A creation module is used to create a virtual sound source in the virtual scene that corresponds to the physical sound source based on the virtual orientation information; The rendering module is used to perform audio rendering on the sound signal generated by the physical sound source based on the virtual sound source, to obtain rendered audio, and to determine the virtual audio to be sent to the virtual object based on the rendered audio.
18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the audio rendering method according to any one of claims 1 to 16.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the audio rendering method according to any one of claims 1 to 16 is implemented.
20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer program or computer-executable instructions are executed by a processor, the audio rendering method according to any one of claims 1 to 16 is implemented.