Acoustic ray tracing method and apparatus, electronic device, wearable device, and storage medium
By dividing the ray tracing path into two stages—listener tracking and sound source tracking—the problem of high computational load in existing acoustic ray tracing algorithms is solved, resulting in savings in computing resources and reduced latency, thus expanding the application scenarios of mobile devices.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GRAVITYXR ELECTRONICS & TECH CO LTD
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-15
AI Technical Summary
Existing acoustic ray tracing algorithms are computationally intensive, especially in virtual reality and mixed reality applications, leading to high consumption of computing resources and latency.
A two-stage tracking strategy is adopted, which divides the ray tracing path into two stages: listener tracking and sound source tracking. Listener tracking is from the listener's position to the last collision point, and sound source tracking is from the collision point to the sound source position. The two stages are modeled and calculated separately.
It effectively reduces the computational load of ray tracing algorithms, especially in multi-source scenarios, reducing the consumption of computing resources and latency, and expanding the application scenarios on mobile devices.
Smart Images

Figure CN2025132219_15052026_PF_FP_ABST
Abstract
Description
Acoustic ray tracing methods, apparatus, electronic devices, wearable devices, and storage media
[0001] This application claims priority to Chinese Patent Application No. 202411585868.2, filed on November 7, 2024, entitled "Acoustic Ray Tracing Method, Apparatus, Electronic Device, Wearable Device and Storage Medium", and to Chinese Patent Application No. 202411587497.1, filed on November 7, 2024, entitled "Acoustic Ray Tracing Method, Apparatus, Electronic Device, Wearable Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of audio signal processing technology, and more specifically, to an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium. Background Technology
[0003] Spatial audio technology is a user-centric technique that processes sound that lacks spatial characteristics to make it sound like it possesses specific spatial features, thus making the content heard by the user more realistic. In virtual reality or mixed reality applications, spatial audio technology can match the spatial characteristics of sound with visual content, resulting in a better sense of immersion.
[0004] In virtual reality or mixed reality applications, real-time acoustic modeling allows spatial audio to match the acoustic characteristics of a predefined space (virtual space or the user's real-world space). This enables the content the user hears to complement the visual content, providing a more immersive experience. For example, when the sound source changes, the content the listener hears also changes. Typically, real-time acoustic modeling uses acoustic ray tracing to model the acoustic path from the sound source to the listener based on the current spatial geometry, acoustic parameters, sound source, and listener position. The unspatialized audio stream is then processed based on the modeling results to obtain an audio stream that matches the expected acoustic characteristics.
[0005] Ray tracing requires tracking every ray emitted from the sound source location, and each ray needs to be reflected multiple times within the room. Therefore, the computational load of ray tracing algorithms is relatively large, and how to reduce the computational load of ray tracing algorithms is an urgent technical problem to be solved. Summary of the Invention
[0006] The purpose of this application is to provide an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium to effectively reduce the computational load of ray tracing algorithms.
[0007] In a first aspect, this application provides an acoustic ray tracing method, comprising:
[0008] Determine the collision information corresponding to the light emitted from the listener's location; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times;
[0009] For each sound source, each propagation path from the sound source location to the listener location is determined based on the collision information and the sound source location, and the echo map corresponding to the sound source is determined based on each propagation path.
[0010] Determine the filter corresponding to the sound source based on the echo map;
[0011] The unspatialized audio streams of each sound source are processed according to the filters corresponding to each sound source to determine the reverberation signal.
[0012] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0013] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:
[0014] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.
[0015] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0016] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;
[0017] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.
[0018] Optionally, the method further includes:
[0019] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;
[0020] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.
[0021] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.
[0022] Optionally, the rate of change of the reflective surface includes a new addition rate and a deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes:
[0023] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.
[0024] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:
[0025] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.
[0026] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.
[0027] Optionally, the method further includes:
[0028] The external reset condition is determined to be triggered when at least one of the following conditions is met:
[0029] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.
[0030] Optionally, when the listener's position does not change significantly, determining the collision information of the target number of light rays includes:
[0031] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.
[0032] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.
[0033] Optionally, selecting the target number of light rays from the N light rays emitted from the listener's location includes:
[0034] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.
[0035] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.
[0036] Optionally, determining the collision information of the target number of light rays includes:
[0037] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0038] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;
[0039] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.
[0040] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0041] Optionally, determining the direction of the reflected light based on the scattering rate of the reflecting surface includes:
[0042] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;
[0043] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.
[0044] Optionally, the collision information includes: the location of the collision point, the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including:
[0045] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.
[0046] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.
[0047] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.
[0048] Optionally, the collision information further includes: the specular reflection angle of the light beam, the cumulative propagation distance; and determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including:
[0049] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.
[0050] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.
[0051] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.
[0052] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.
[0053] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes:
[0054] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.
[0055] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;
[0056] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.
[0057] Optionally, determining the filter corresponding to the sound source based on the echo map includes:
[0058] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;
[0059] Determine the second filter corresponding to the sound source based on the echo map;
[0060] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.
[0061] Optionally, the smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including:
[0062] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.
[0063] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;
[0064] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.
[0065] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0066] The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.
[0067] Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including:
[0068] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path.
[0069] For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path.
[0070] Optionally, determining the collision information and reporting path corresponding to the light emitted from the listener's location, and constructing a reflection tree based on the reporting path, includes: determining the collision information and reporting path corresponding to the light emitted from the listener's location, and constructing a reflection tree based on the reporting path when at least one of the following conditions is met: the listener's location changes, the scene model changes, or the reflectivity of any reflective surface in the scene model changes.
[0071] Optionally, the scene model is represented using triangulation; the method further includes:
[0072] Determine the diffraction edge in the scene model; the diffraction edge is the coincident side of two non-coplanar triangles;
[0073] Determine and save diffraction edge information; the diffraction edge information corresponds to the diffraction edge; wherein, when the dynamic geometry in the scene model changes, the corresponding diffraction edge information is updated;
[0074] Accordingly, determining the collision information and reporting path corresponding to the light emitted from the listener's location includes: performing ray tracing on the light emitted from the listener's location, determining the collision information, and determining the reporting path based on the diffraction edge information.
[0075] Optionally, the reporting path includes a reflection path and a diffraction path; ray tracing is performed on the light emitted from the listener's position to determine collision information, and the reporting path is determined based on the diffraction edge information, including:
[0076] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0077] When the light ray collides with the scene model, the collision information for this collision is determined;
[0078] When it is determined that the tracking of the ray will not end, the reflection path is determined and the direction of the reflected ray is calculated to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0079] Furthermore, when it is determined that the tracking of the light ray will not end, it is determined whether diffraction will occur based on the diffraction edge information. When diffraction occurs, the diffraction path is determined, and the direction of the diffracted light ray is calculated to determine whether it will collide with another reflective surface in the scene model based on the direction of the diffracted light ray.
[0080] Optionally, determining whether diffraction will occur based on the diffraction edge information includes:
[0081] Determine the point of collision between the light ray and the reflecting surface;
[0082] When the collision point is close to the edge of the triangle corresponding to the reflecting surface, it is determined whether the edge is the saved diffraction edge; the collision point is determined to be close to the edge of the triangle corresponding to the reflecting surface based on the transformed centroid coordinates of the collision point.
[0083] If the edge is one of a plurality of diffraction edges that are preserved, then two shadow regions are determined according to the two reflecting surfaces corresponding to the diffraction edge. When the light is in either of the shadow regions, diffraction is determined to occur.
[0084] Optionally, a reflection tree is constructed based on the reported path, including:
[0085] The reported path is deduplicated to obtain a deduplicated path; the reported path includes a start node and an end node; the start node or the end node is a reflecting surface or a diffraction edge;
[0086] The listener is identified as the root node of the reflection tree, and the following steps are repeated until the maximum depth of the reflection tree reaches a preset depth:
[0087] For each node in the reflection tree, find the target node that the node can connect to from all the deduplicated paths, and determine the target node as the child node of the node.
[0088] Optionally, the method further includes:
[0089] When determining the reflection tree, if the node is the root node, the spatial position corresponding to the root node is determined as the listener's position;
[0090] When the node is a reflective surface, determine the mirror position of the parent node of the node relative to the reflective surface, and determine the mirror position as the spatial position of the node.
[0091] When the node is a diffraction edge, the center position of the diffraction edge is determined as the diffraction point or the spatial position corresponding to the node.
[0092] Optionally, the reflection path in the second propagation path is a specular reflection path; multiple second propagation paths are determined from the sound source location to the listener location based on the reflection tree, including:
[0093] For each sound source, determine the reflecting surface and diffraction edge where the light emitted from the sound source location collides with the scene model for the first time, so as to obtain a list of visible nodes of the sound source.
[0094] For each leaf node in the reflection tree, when the leaf node is in the sound source visible list, it is determined whether the propagation path corresponding to the leaf node is a valid path;
[0095] When the leaf node is not in the visible list of sound sources, the propagation path corresponding to the leaf node is determined to be an invalid path;
[0096] The second propagation path is determined based on the established legal path; the second propagation path is a path obtained by sequentially connecting the sound source location, the reflection points or diffraction points corresponding to each node in the legal path, and the listener's location.
[0097] Optionally, determining whether the propagation path corresponding to the leaf node is a valid path includes:
[0098] If every child node from the leaf node to the root node satisfies the target condition, then the path from the leaf node to the root node is determined to be a valid path; otherwise, it is an invalid path.
[0099] When the node is a reflective surface, the target condition is: the reflection point corresponding to the reflective surface is within the triangle corresponding to the reflective surface and the corresponding path is not obstructed; the position of the reflection point is related to the spatial position corresponding to the reflective surface and the spatial position of the next level node.
[0100] When the node is a diffraction edge, the target conditions are: the corresponding path is not occluded, and the preceding and following nodes of the node are located in two different shadow areas corresponding to the node.
[0101] Optionally, determining the echo map corresponding to the second propagation path includes:
[0102] For the sound source, determine the initial energy corresponding to the sound source;
[0103] For each second propagation path, the distance attenuation coefficient and air absorption coefficient are determined based on the total length of the second propagation path. The reflection attenuation coefficient of each reflecting surface in the second propagation path is determined. The reflection attenuation coefficients of each reflecting surface are multiplied to obtain the overall reflection attenuation coefficient. The diffraction coefficient is determined based on the diffraction edge information of each diffraction edge in the second propagation path. The reflection attenuation coefficient of the reflecting surface is related to the reflectivity and scattering rate of the reflecting surface.
[0104] The multiplication result of the initial energy, the distance attenuation coefficient, the air absorption coefficient, the overall reflection attenuation coefficient, and the diffraction coefficient is determined as the energy received by the listener;
[0105] The direction in which the last node points to the listener's position is defined as the receiving direction;
[0106] The energy received by the listener, the receiving direction, and the propagation time corresponding to the second propagation path are determined as the echo map corresponding to the second propagation path.
[0107] Secondly, this application provides an acoustic ray tracing device, comprising:
[0108] The listener tracking module is configured to determine the collision information corresponding to the light emitted from the listener's location; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times.
[0109] The sound source tracking module is configured to, for each sound source, determine each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path;
[0110] The filter synthesis module is configured to determine a filter corresponding to the sound source based on the echo map;
[0111] The processing module is configured to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.
[0112] Thirdly, this application provides an electronic device, including: at least one processor and a memory;
[0113] The memory stores computer-executed instructions;
[0114] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any of the first aspects.
[0115] Fourthly, this application provides a wearable device including a processing unit; the processing unit is configured to perform the method as described in any of the first aspects.
[0116] Fifthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method as described in any of the first aspects.
[0117] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the first aspects.
[0118] This application provides an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium. By determining the collision information corresponding to the ray emitted from the listener's position—the collision information being the information of each collision point obtained from multiple collisions between the ray and the scene model—for each sound source, various propagation paths from the sound source position to the listener's position are determined based on the collision information and the sound source position. Furthermore, an echo map corresponding to the sound source is determined based on each propagation path. A filter corresponding to the sound source is determined based on the echo map. The unspatialized audio stream of each sound source is processed using the corresponding filter to determine the reverberation signal. By dividing the tracing path into a path from the listener to the last collision point and a path from the last collision point to the sound source during reverse tracing, and modeling these two paths separately, the step of determining collision information is eliminated when the listener's position remains unchanged. Moreover, when there are multiple sound sources, the step of determining collision information is calculated only once, thereby effectively reducing the computational load. Attached Figure Description
[0119] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0120] Figure 1 is an application scenario diagram provided by an embodiment of this application;
[0121] Figure 2 is a flowchart illustrating an acoustic ray tracing method provided in an embodiment of this application;
[0122] Figure 3 is a schematic diagram of an overall architecture for generating spatial audio provided in an embodiment of this application;
[0123] Figure 4 is a detailed schematic diagram of an architecture for generating spatial audio provided in an embodiment of this application;
[0124] Figure 5 is a schematic diagram of the structure of a listener tracking module provided in an embodiment of this application;
[0125] Figure 6 is a schematic diagram illustrating significant and non-significant changes in the listener's position according to an embodiment of this application;
[0126] Figure 7 is a schematic diagram of a light propagation process provided in an embodiment of this application;
[0127] Figure 8 is a schematic diagram of the specific process of a listener tracking module provided in an embodiment of this application;
[0128] Figure 9 is a schematic diagram of calculating the reflection direction provided in an embodiment of this application;
[0129] Figure 10 is a schematic flowchart of a sound source tracking module provided in an embodiment of this application;
[0130] Figure 11 is a schematic diagram of a specular reflection path provided in an embodiment of this application;
[0131] Figure 12 is a schematic diagram of a scattering path provided in an embodiment of this application;
[0132] Figure 13 is a schematic diagram of an echo map provided in an embodiment of this application;
[0133] Figure 14 is a schematic diagram of a filter synthesis module provided in an embodiment of this application;
[0134] Figure 15 is an example diagram of reconstructing a frequency band impulse response from an echo map according to an embodiment of this application;
[0135] Figure 16 is a schematic diagram of outdoor diffraction provided in an embodiment of this application;
[0136] Figure 17 is a detailed schematic diagram of an architecture for generating spatial audio provided in an embodiment of this application;
[0137] Figure 18 is a schematic diagram of a diffraction edge provided in an embodiment of this application;
[0138] Figure 19 is a schematic flowchart of another listener tracking module provided in an embodiment of this application;
[0139] Figure 20 is a schematic diagram of random diffraction provided in an embodiment of this application;
[0140] Figure 21 is a top view schematic diagram of random diffraction provided in an embodiment of this application;
[0141] Figure 22 is a schematic diagram of a path report provided in an embodiment of this application;
[0142] Figure 23 is a schematic flowchart of another sound source tracking module provided in an embodiment of this application;
[0143] Figure 24 is a schematic diagram of a reflection tree provided in an embodiment of this application;
[0144] Figure 25 is a schematic diagram of a mirror position and propagation path provided in an embodiment of this application;
[0145] Figure 26 is a schematic diagram of a path search module provided in an embodiment of this application;
[0146] Figure 27 is a schematic diagram of a diffraction path modeling provided in an embodiment of this application;
[0147] Figure 28 is a schematic diagram of another filter synthesis module provided in an embodiment of this application;
[0148] Figure 29 is a schematic diagram of an acoustic ray tracing device provided in an embodiment of this application;
[0149] Figure 30 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0150] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0151] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0152] In virtual reality or mixed reality applications, spatial audio technology can be used to match the spatial characteristics of sound with visual content to achieve a better sense of immersion. Figure 1 shows an application scenario provided by an embodiment of this application. As shown in Figure 1, when there are two sound sources in a space, the listener can simultaneously receive the spatial audio of these two sound sources in the virtual or real space. That is to say, the spatial audio ultimately received by the listener is related to the location of the sound sources, the listener's location, and the geometry of the room, thereby enabling the user to match the spatial audio heard with the visual content to provide a stronger sense of realism.
[0153] A common method for determining spatial audio is to use acoustic ray tracing to model the acoustic path from the sound source to the listener based on the current spatial geometry, acoustic parameters (such as the reflectivity of reflective surfaces in a room), the location of the sound source, and the location of the listener. Here, modeling refers to determining filters with propagation characteristics from the sound source to the listener based on the energy received by the listener, so as to process the unspatialized audio stream according to the filters to obtain spatial audio that matches the expected acoustic characteristics.
[0154] In ray tracing, energy-carrying rays are randomly emitted from the sound source location, and the propagation of each ray in the room is tracked. When a ray collides with a wall in the room, its energy is attenuated, and it is then reflected and continues to propagate. This process is repeated until the condition for ending ray tracing is met, at which point ray tracing stops. For each collision point, the energy scattered to the listener can be calculated. An impulse response is generated based on the energy received by the listener, and this impulse response is then used to process the unspatialized audio stream.
[0155] The ray tracing process described above requires tracking a large number of rays, and each ray reflects multiple times within the room, resulting in a massive computational burden. To improve computational speed, one approach is to sort the rays emitted from the sound source, group them based on coherence, and assign each group of coherent rays to different CPU (Central Processing Unit) cores for parallel processing. Due to the characteristics of the ray intersection algorithm, processing coherent rays is faster than processing completely random rays, thus achieving a faster tracing speed.
[0156] However, the above method still has the following problems: only light rays emitted from the sound source can generate coherent rays, and the light rays after reflection of a set of coherent rays are no longer coherent. Therefore, this method can only utilize the faster computational property of coherent rays during the first ray collision, and cannot utilize this property for the computation of subsequent reflected rays. Furthermore, using multiple CPU cores for parallel processing only reduces the time required to run the ray tracing algorithm, without reducing power consumption, because the total computational load is not effectively reduced. In addition, as the number of sound sources in the scene increases, the same algorithm needs to be run for each sound source, and the overall time consumption will increase linearly. Therefore, the above ray tracing algorithm still suffers from a high computational load, and further reductions in the computational load of the ray tracing process are needed to expand its application scenarios on mobile devices.
[0157] Existing ray tracing methods involve ray light emanating from a sound source, colliding multiple times, and reaching the listener, forming a complete path. All rays from the sound source to the listener are then traced. Addressing these issues, this application proposes a two-stage tracing strategy. This strategy modifies the ray tracing path to a path from the listener to the sound source, splitting the complete path into two parts for separate tracing. Specifically, it splits the path into two stages: listener tracking and sound source tracking, thus decoupling the listener from the sound source. Therefore, a complete path is divided into two parts: the first part is the path from the listener to the last collision point (the last collision point where ray tracing from the listener ends), which is independent of the sound source location; the second part is the path from each collision point to the sound source, where each sound source is modeled separately. The advantages of the above method are as follows: Since the path in the first part is independent of the sound source location, for scenarios with multiple sound sources, the tracking process in the first part only needs to be run once, and the computational cost of the tracking process in the second part is linearly related to the number of sound sources. Furthermore, the sound source or the listener may move. When the sound source moves but the listener remains stationary, the tracking result in the first part will not change, thus the previous listener tracking result can be reused. The listener tracking part does not need to be run; only the sound source tracking part needs to be run. In summary, in scenarios with multiple sound sources and only sound source movement, the computational cost can be effectively reduced, thereby reducing system power consumption.
[0158] Figure 2 is a flowchart illustrating an acoustic ray tracing method provided in an embodiment of this application. The method includes steps S201 to S204:
[0159] Step S201: Determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times.
[0160] This application describes a reverse ray tracing process, where the ray is traced originating from the listener's location. By reversing the tracing process, ray tracing can be divided into two stages, thus decoupling the listener from the sound source. It should be noted that the ray emanating from the listener's location here does not refer to real light rays, but rather to simulated light rays, such as multiple light rays emanating from the listener's location represented by equations. The ray tracing here is used to simulate the sound propagation path.
[0161] Users and listeners are situated within a scene, which can be represented using a scene model. Optionally, the scene can be a room, in which case the scene model can be used to indicate the room's dimensions, the positions of objects within the room, and / or the reflective surfaces contained within the room and the reflective surfaces of each object, etc.
[0162] Optionally, multiple random rays can be emitted from the listener's position. These rays will collide with the scene model during propagation, and will randomly reflect and scatter. The collision information can be determined using the first-stage tracking process. This collision information can be obtained by tracking each ray emitted from the listener's position and recording the collision point information when it collides with the scene model. For example, the collision information can include the location of the collision point, the reflecting surface it occupies, and other information.
[0163] Optionally, the collision information may also include the order of the collision points corresponding to the light rays, or the propagation path from the listener's position to each collision point, in order to determine the propagation paths from the sound source position to the listener's position.
[0164] For example, three rays are emitted from the listener's position. For each ray, the corresponding collision information can be determined. For one of the rays, it may collide with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model in sequence. After the condition for ending ray tracing is met, the ray tracing ends. Here, collision point 3 is the last collision point corresponding to the ray, so the information of these three collision points can be determined.
[0165] Tracing the light rays emitted from the listener's position to the last collision point reveals the path of the light rays propagating from the listener's position to the last collision point. This process is independent of the sound source's location; that is, regardless of the number of sound sources or their locations, the collision information of each light ray remains unchanged. Therefore, the path of the light rays from the listener's position to the last collision point can be traced, and this tracing process can be considered as listener tracing.
[0166] Step S202: For each sound source, determine each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path.
[0167] Once the collision information is determined, ray tracing can continue based on the collision information. Ray tracing here is to trace the path from each collision point to the sound source location. This path is related to the sound source location. When the sound source location is different, the path is different. This tracing process can be regarded as sound source tracing.
[0168] For a given sound source, based on the locations of each collision point, the paths from each collision point to the sound source can be determined. Then, based on the paths from the listener's location to each collision point, multiple complete propagation paths from the listener to the sound source can be obtained, and it can be equivalently assumed that there are identical propagation paths from the sound source to the listener's location. Therefore, based on the above steps, each propagation path from the sound source to the listener's location can be determined, and thus the echo map from the sound source to the listener's location can be determined, i.e., the echo map corresponding to the sound source. The echo map represents the arrival time, direction of arrival, and energy received by the listener for each propagation path from the sound source to the listener's location.
[0169] For example, if a ray of light collides sequentially with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model, and there are no other obstructions between each collision point and the sound source, then there are three propagation paths from the sound source position to the listener position. Path 1 is the ray of light from the sound source position through collision point 1 to the listener position; path 2 is the ray of light from the sound source position through collision point 2 and collision point 1 in sequence to the listener position; and path 3 is the ray of light from the sound source position through collision point 3, collision point 2, and collision point 1 in sequence to the listener position.
[0170] Step S203: Determine the filter corresponding to the sound source based on the echo map.
[0171] Once the echo map for each sound source is determined, each source can be processed individually to generate a filter for that source based on the source tracing results (i.e., the echo map). Optionally, the filter can be in the form of an Ambisonics filter bank.
[0172] Step S204: Process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.
[0173] After generating the filter corresponding to each sound source, a fast convolution method can be used to convolve the unspatialized audio stream of the corresponding sound source based on the filter to obtain the Ambisonics reverberation signal.
[0174] Figure 3 is a schematic diagram of the overall architecture for generating spatial audio according to an embodiment of this application. As shown in Figure 3, the room geometry is considered as the scene model in this application. The method of this application can be applied to a processing device, which includes an acoustic ray tracing module and a spatial audio rendering module. The room acoustic parameters, room geometry (also known as the scene model), sound source positions, and listener positions can be input into the processing device. The acoustic ray tracing module can process the input information to obtain filters corresponding to each sound source, and input the obtained filters corresponding to each sound source, the unspatialized audio stream in the memory, and the listener orientation information provided by the inertial sensor into the spatial audio rendering module, so that the spatial audio rendering module can output a spatialized audio stream and output it to a sound playback device, such as a speaker and headphones. Steps S201 to S203 of this application are the execution content of the acoustic ray tracing module, and step S204 is the execution content of the spatial audio rendering module.
[0175] Figure 4 is a detailed schematic diagram of a spatial audio generation architecture provided by an embodiment of this application. The implementation process of steps S201 to S204 can be divided into two stages: simulation stage 110 and processing stage 120. In simulation stage 110, based on the room scene model (3D Mesh), the reflectivity of each reflective surface in the corresponding scene model, and the positions of the listener and sound sources, the acoustic characteristics of the room are modeled using a ray tracing method. The modeling result is a filter corresponding to each sound source. In processing stage 120, the estimated filters are used to process the unspatialized audio to obtain a reverberation signal. This room reverberation signal is added to the direct sound signal processed by HRTF (Head Related Transfer Functions) and then output to an audio playback device, such as a speaker or headphones.
[0176] Optionally, the simulation and processing stages can have different update frequencies depending on the actual scenario. The processing stage is typically designed to generate a real-time audio stream. For example, when the sampling rate is 48000Hz and 1024 samples are processed each time, the processing stage needs to run at least once every 21.3ms (1024 / 48000*1000) to generate the real-time audio stream. The update frequency of the simulation stage is set according to the complexity of the actual scenario and the system's computing power. For example, it can be set to update once every 100ms. This application does not limit the update frequency of the simulation and processing stages. Depending on the system architecture, the simulation and processing stages can also run on different processing devices; this application does not impose any restrictions on this.
[0177] Optionally, the simulation phase mainly includes: a listener tracking module 111, a sound source tracking module 112, and a filter synthesis module 113. The listener tracking part is only related to the listener's position; therefore, the listener tracking module 111 only needs to run when the listener's position changes. When the listener's position remains unchanged, this step can be skipped, and the output of the previous listener tracking module 111 can be reused to directly run the sound source tracking module 112. The sound source tracking module 112 runs separately for each sound source, depending on both the sound source position and the output of the listener tracking module 111. If the position and number of sound sources do not change, and the listener tracking module 111 is not running, then the sound source tracking module 112 can also be skipped. Similarly, the filter synthesis module 113 runs separately for each sound source, synthesizing the output of the corresponding sound source tracking module 112 into an Ambisonic filter bank for that sound source.
[0178] The processing stage mainly includes steps such as a fast convolution module 121, an HRTF processing module 122, an Ambisonics rotation module 123, and an Ambisonics decoding module 124. The fast convolution module 121 uses a fast convolution method to convolve the Ambisonics filter bank with the input unspatialized audio signal to obtain an Ambisonics reverberation signal. The Ambisonics rotation module 123 spatially rotates the Ambisonics reverberation signal based on the user's orientation information provided by an inertial sensor to match the user's current orientation. The Ambisonics decoding module 124 performs binaural decoding on the rotated Ambisonics reverberation signal to obtain a binaural reverberation signal. This signal is mixed with the binaural direct sound signal output by the HRTF processing module 122 to obtain the final binaural signal used for playback.
[0179] This application provides an acoustic ray tracing method that determines collision information corresponding to light rays emitted from a listener's position. The collision information comprises information about each collision point obtained from multiple collisions between the light ray and a scene model. For each sound source, based on the collision information and the sound source position, various propagation paths from the sound source position to the listener's position are determined. Furthermore, based on each propagation path, an echo map corresponding to the sound source is determined. Based on the echo map, a filter corresponding to the sound source is determined. The unspatialized audio stream of each sound source is processed using the corresponding filter to determine a reverberation signal. By dividing the tracing path into a path from the listener to the last collision point and a path from the last collision point to the sound source during reverse tracing, and modeling these two paths separately, the method eliminates the need to determine collision information when the listener's position remains unchanged. Moreover, when there are multiple sound sources, the collision information determination step is only calculated once, effectively reducing the computational load.
[0180] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0181] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:
[0182] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.
[0183] The collision information corresponding to the light emitted from the listener's position is determined by the listener tracking module 111, but this module is not executed every time an update time is reached. Optionally, the execution of this module can be determined based on whether the listener's position has changed, whether the scene model has changed, or whether the reflectivity of any reflective surface in the scene model has changed.
[0184] When the listener's position changes, the collision information when the light ray collides with the scene model will change. Even if the listener's position remains the same, the collision information will still change when the scene model changes. When the reflectivity of any reflective surface in the scene model changes, the remaining energy of the light ray after a collision can be altered, thus determining whether to continue tracking the ray; therefore, the collision information will also change.
[0185] When the update time of the listener tracking module 111 is reached and any one of the above three conditions is met, the listener tracking module 111 is run. Conversely, when none of the above three conditions are met, the listener tracking module 111 is not run, and the output of the previous run of the listener tracking module 111 can be reused directly.
[0186] By determining whether the listener's position, scene model, or the reflectivity of any reflective surface has changed, it is possible to accurately determine whether the listener tracking module is running, and in some scenarios, the listener tracking module can be disabled.
[0187] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0188] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;
[0189] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.
[0190] When determining the collision information corresponding to the light emitted from the listener's location, in some scenarios, only a portion of the light rays emitted from the listener's location can be tracked, instead of all light rays, in order to reduce computational load and power consumption.
[0191] When the listener's position changes, it is necessary to determine the collision information corresponding to the light emitted from the listener's position. In reality, the sound source and the listener are mostly located in the same space (such as the same room), and the listener's movements are mostly small, continuous movements, with a low percentage of significant movements. Significant movement can be understood as a large movement, or moving from one room to another. Insignificant movement can be understood as a small movement, or the acoustic environment in which the listener is located does not change significantly (e.g., the listener does not move from one space to another). When the listener's position changes insignificantly, a small number of light rays can be tracked and combined with the output of the previous listener tracking module 111 to obtain the current output result.
[0192] Specifically, when the listener's position changes in a non-significant way, some light rays can be selected for tracking; when the listener's position changes significantly, all light rays can be selected for tracking.
[0193] For example, when a non-significant movement of the listener is detected, the listener tracking module 111 can select only a portion of the light rays for tracking during operation.
[0194] The above methods can reduce computational load and power consumption in scenarios where the listener moves but not significantly.
[0195] Optionally, the method further includes:
[0196] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;
[0197] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.
[0198] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.
[0199] Figure 5 is a schematic diagram of the structure of a listener tracking module provided in an embodiment of this application. As shown in Figure 5, the listener tracking module 111 is divided into a listener tracking module (first reflection) 201 and a listener tracking module (subsequent multiple reflections) 206, and an adaptive adjustment strategy for the number of light rays is used to reduce the amount of computation.
[0200] The core of adaptively adjusting the number of rays lies in determining whether the listener's position has changed significantly (including significant changes in the scene in which the user and sound source are currently located, or changes in the reflectivity of the reflecting surface). If the listener's position has not changed significantly, then the simulation can use fewer rays and make incremental updates based on the previous simulation results. However, if the listener's position has changed significantly, then the previous simulation results need to be discarded, and a full simulation should be performed directly using the pre-set maximum number of rays (or all rays emitted from the listener's position).
[0201] Figure 6 is a schematic diagram illustrating significant and non-significant changes in the listener's position according to an embodiment of this application. Two rooms are connected by a door. Small, continuous movements of the listener within one room, such as moving from listener position 1 to listener position 2, are considered non-significant changes. However, moving from one room to another, such as moving from listener position 2 to listener position 3, is considered a significant change.
[0202] When determining whether the listener's position has changed significantly, the first collision information of the light rays emitted from the listener's position can be used. When N light rays are emitted from the listener's position, the first collision information includes the reflecting surfaces of the N light rays that collide for the first time. The listener's position can be determined based on the first collision information corresponding to the current simulation and historical simulations.
[0203] Historical simulations can be a single historical simulation or multiple historical simulations. Whether the listener's position has changed significantly can be determined based on the first collision information of the current simulation and the first collision information of multiple historical simulations. Specifically, if the listener's position has not changed significantly in multiple historical simulations, then the historical simulations are the simulation in which the listener's position changed significantly most recently, plus at least one simulation from each subsequent simulation. For example, if the current simulation is the 10th one, and the listener's position changed significantly in the 5th simulation, but not significantly in the 6th to 9th simulations (actually, the listener's position did not change significantly relative to the 5th simulation), then the historical simulations corresponding to the 10th simulation are one or more simulations from the 5th to 9th simulations.
[0204] Optionally, in the listener tracking module (first reflection) 201, N rays are emitted from the listener's position, each carrying initial energy E0. Their directions are randomly generated to ensure that they cover as uniformly as possible throughout the space. By performing ray intersection calculations with the scene, the reflecting surfaces S0, S1, ... S that these N rays collide with can be obtained. N-1 .
[0205] The reflective surface statistics module 202 counts all reflective surfaces S0, S1, ... S N-1 Perform deduplication statistics and maintain a list to represent all unique reflective surfaces. Optionally, this list can be implemented using a hash table (HashMap) for fast updates and lookups. This list is stored and used for the next simulation run, and is also input into the reflective surface change rate module 203.
[0206] The reflective surface change rate module 203 calculates the reflective surface change rate by comparing the reflective surface recorded in this simulation with the reflective surface recorded in historical simulations (which can be all non-repeating reflective surfaces that appear in multiple historical simulations) and outputs it to the comprehensive decision module 204.
[0207] In addition to determining whether the listener's position has changed significantly based on the collision information at the time of the first collision, it can also be determined based on whether an external reset condition is triggered. The external reset condition can be input into the comprehensive decision module 204.
[0208] The comprehensive decision module 204 can determine whether a significant change has occurred based on whether the rate of change of the reflective surface is greater than a preset value and whether an external reset condition has been triggered.
[0209] Judging whether the listener's position has moved significantly by using the first collision information of light emitted from the listener's position can be effective in scenarios with complex models, such as rooms with many objects that can reflect light. Even if the listener moves less than a preset distance compared to historical simulations, the objects around the listener may have changed significantly. This method can effectively identify such changes and improve the accuracy of judging whether the listener's position has moved significantly.
[0210] Optionally, the rate of change of the reflective surface includes a new addition rate and a deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes:
[0211] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.
[0212] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:
[0213] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.
[0214] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.
[0215] Optionally, the reflective surface change rate module 203 adjusts the value based on the number of newly added reflective surfaces n. new Delete the number of reflective surfaces n remove The total number of reflective surfaces in the historical simulation, n previous (Number of reflective surfaces after deduplication), and the total number of reflective surfaces n in this simulation. current (Number of reflective surfaces after deduplication), and calculate the rate of change of reflective surfaces, including the increase rate update_ratio. new and deletion rate update_ratio remove :
[0216] Optionally, the addition rate is the ratio of the number of newly added reflective surfaces to the total number of reflective surfaces in this simulation; the number of newly added reflective surfaces is the number of reflective surfaces added in this simulation compared to historical simulations; the deletion rate is the ratio of the number of deleted reflective surfaces to the total number of reflective surfaces in historical simulations; the number of deleted reflective surfaces is the number of reflective surfaces deleted in this simulation compared to historical simulations.
[0217] A first threshold can be set for the new addition rate and a second threshold can be set for the deletion rate. The calculated new addition rate is compared with the first threshold and the calculated deletion rate is compared with the second threshold. When the new addition rate is less than the first threshold and the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.
[0218] Conversely, a significant change in the listener's location is determined when any one of the following conditions is met: the new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, or the external reset condition is triggered.
[0219] By comparing the addition rate, deletion rate, and threshold, it is possible to determine whether the listener's position has changed significantly based on the first collision information of the light emitted from the listener's position.
[0220] Optionally, the method further includes:
[0221] The external reset condition is determined to be triggered when at least one of the following conditions is met:
[0222] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.
[0223] When the scene model changes or the reflectivity of any reflective surface in the scene model changes, the collision information corresponding to the light emitted from the listener's position will change. If the distance the listener's position moves in this simulation exceeds a preset distance compared to previous simulations, it indicates a significant change in the listener's position. Furthermore, if no significant change is triggered in multiple consecutive simulations, an external reset condition can be triggered to avoid the continuous accumulation of errors.
[0224] By determining whether the external reset condition has been triggered by the above conditions, the accuracy of judging whether the listener's position has changed significantly can be improved.
[0225] Optionally, when the listener's position does not change significantly, determining the collision information of the target number of light rays includes:
[0226] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.
[0227] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.
[0228] When the listener's position changes insignificantly, the number of targets is less than N. The specific value of the number of targets can be related to the rate of change of the reflective surface. For example, a rate of change of the reflective surface less than 0.5 is considered a non-significant change. When the rate of change of the reflective surface is 0.1 and 0.4, the number of targets will be different. When the rate of change of the reflective surface is large, it can be intuitively understood that the user's movement distance is large, and therefore the number of targets is large. In this simulation, more light rays are selected for tracking.
[0229] Optionally, the rate of change of the reflective surface includes the addition rate and the deletion rate. The larger the addition rate, the larger the number of targets; or, the larger the deletion rate, the larger the number of targets.
[0230] Optionally, curves relating the new addition rate to the target quantity and the deletion rate to the target quantity can be set to determine the target quantity 1 corresponding to the new addition rate and the target quantity 2 corresponding to the deletion rate, and then the final target quantity can be determined based on the target quantity 1 and the target quantity 2.
[0231] For example, when the rate of change of the reflective surface is 0.1, the number of targets is one-quarter of all light rays; when the rate of change of the reflective surface is 0.2, the number of targets is one-third of all light rays; when the rate of change of the reflective surface is 0.3, the number of targets is half of all light rays, and so on. Conversely, if the number of targets is always one-quarter of all light rays for different rates of change of the reflective surface, the accuracy of the calculation results will be low; if the number of targets is always half of all light rays for different rates of change of the reflective surface, the computational load will be too large.
[0232] By determining the target quantity based on the rate of change of the reflective surface, different target quantities can be set based on different reflectivities, which can reduce the amount of calculation while ensuring the accuracy of the calculation results.
[0233] Optionally, selecting the target number of light rays from the N light rays emitted from the listener's location includes:
[0234] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.
[0235] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.
[0236] When selecting the target number of rays, if there is a newly added reflective surface of the first collision, the rays that pass through the newly added reflective surface of the first collision can be selected first. If the total number of rays that pass through the newly added reflective surface of the first collision does not reach the target number, then multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.
[0237] As shown in Figure 5, the light selection module 205 can determine the number of targets based on the decision result. If there is no significant change, then the light rays of the target number are selected from the listener tracking module (first reflection) 201 and the listener tracking module (subsequent multiple reflections) 206 is executed. If there is a significant change, then all light rays are selected and the listener tracking module (subsequent multiple reflections) 206 is executed.
[0238] By prioritizing the light rays from the newly added first collision reflector, the effectiveness of ray tracing can be achieved, allowing the final collision information to reflect the listener's movement, thereby improving the accuracy of the determined collision information.
[0239] Optionally, determining the collision information of the target number of light rays includes:
[0240] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0241] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;
[0242] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.
[0243] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0244] Figure 7 is a schematic diagram of a light propagation process provided in an embodiment of this application. Multiple light rays can be randomly emitted from the listener's position, and each light ray can be tracked to obtain collision information with the scene model.
[0245] Figure 8 is a schematic flowchart of a listener tracking module provided in an embodiment of this application. When light hits a reflective surface, reflection and scattering occur. The reflectivity α of each reflective surface is used to determine the specific characteristics of the light. s The remaining energy α after reflection can be calculated. sE, where E represents the energy currently carried by the ray. The distance between the current collision point and the previous collision point is calculated and accumulated to obtain the ray's propagation distance (the distance from the listener's position to the current collision point). The number of collisions, the ray's propagation distance, and the remaining ray energy determine whether to end tracking of the current ray. If the conditions for ending tracking are not met, the direction of the reflected ray is calculated based on the scattering rate of the reflective surface, and the next collision point with the scene model is calculated. The above steps are repeated for all rays until the conditions for ending tracking are met for all rays.
[0246] When the number of collisions reaches the upper limit, the propagation distance of the light reaches the upper limit, or the remaining light energy reaches the lower limit, the conditions for ending the tracking of the current light are determined to be met.
[0247] The collision information for each collision can include: the reflecting surface where the collision point is located, the specific collision location, the angle of the incident light and the angle of the reflected light, the cumulative distance, and other information.
[0248] By tracing rays, collision information can be determined, and by determining whether to end ray tracing, ray tracing can be terminated in a timely manner when conditions are met, thereby reducing the amount of computation.
[0249] Optionally, determining the direction of the reflected light based on the scattering rate of the reflecting surface includes:
[0250] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;
[0251] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.
[0252] Figure 9 is a schematic diagram of calculating the reflection direction provided in an embodiment of this application. As shown in Figure 9, the direction of the reflected light is obtained by combining the specular reflection direction and the random scattering direction. If the incident direction is d... incident If the normal vector of the collision surface is n, then the direction of mirror reflection is: d specular =d incident -2*d incident ·n
[0253] Optionally, the random scattering direction dscattering is a randomly generated direction vector that follows a Lambert distribution relative to the normal vector n.
[0254] Based on the scattering rate αscattering of the collision surface, the direction of the reflected light is obtained and normalized, as shown in the following equation: dreflection=αscatteringdscattering+(1-αscattering)d specular
[0255] The scattering rate is related to the material of the reflecting surface. The reflection direction obtained by the above method simulates the difference in dispersion caused by the scattering rate of different acoustic materials.
[0256] Optionally, the collision information includes: the location of the collision point, the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including:
[0257] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.
[0258] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.
[0259] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.
[0260] Multiple sound source tracking modules 112 are included, which can process all collision information generated by the listener tracking module 111. Based on the sound source location, scene model, and reflectivity information of the corresponding sound source, they generate the propagation path and echo map from the sound source to the listener. Since the tracked light rays originate from the listener, the actual calculated path is the propagation path from the listener to the sound source, and it is equivalent to assuming that there is a common propagation path from the sound source location to the listener location.
[0261] Figure 10 is a schematic flowchart of a sound source tracking module provided in an embodiment of this application. For the light ray corresponding to the collision point, it can be determined whether there is an obstruction between the collision point and the sound source location. If there is no obstruction, the path between the collision point and the sound source location can be determined. If there is an obstruction, the next collision point is processed. When the path between the collision point and the sound source location is determined, the propagation path between the corresponding sound source location and the listener location can be obtained, and thus multiple paths between the sound source location and the listener location can be obtained. For the propagation path between the sound source location and the listener location, an echo map can be calculated, including: the energy received by the sound source, the arrival time and direction of the light ray reaching the sound source. When calculating the echo map, it can be determined based on the reflectivity and scattering rate of the reflecting surface where the collision point is located.
[0262] Optionally, the collision information further includes: the specular reflection angle of the light beam, the cumulative propagation distance; and determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including:
[0263] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.
[0264] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.
[0265] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.
[0266] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.
[0267] Optionally, based on the reflectivity α of the reflecting surface where the collision point is located. s The remaining energy α after reflection is obtained. s E, where E is the energy currently carried by the ray. α s E is further divided into two parts: mirror energy and scattering energy, where the scattering energy is αscatteringα. sE, αscattering is the scattering rate of the reflecting surface at the current collision point, and the remaining part is (1-αscattering)α s E is the mirror energy.
[0268] Figure 11 is a schematic diagram of a specular reflection path provided in an embodiment of this application, and Figure 12 is a schematic diagram of a scattering path provided in an embodiment of this application. As shown in Figures 11 and 12, when there are no other obstructions between the collision point and the sound source, there is a propagation path between the collision point and the sound source, and its direction comes from the position corresponding to the collision point. The specular reflection angle of the light ray (which has also been calculated in the listener tracking module 111) is judged to see if it can pass through the sound source. If so, the path is considered to be a specular reflection path, and the specular energy of the light ray is (1-αscattering)α. s If all of E is emitted to the sound source, then the path is considered a scattering path, and a portion of the scattered energy of the light is emitted to the sound source.
[0269] Optionally, assuming a scattering ratio of α, the energy emitted to the sound source is αscatteringα. s The value of E.α can be given by the following formula:
[0270] Where θ is the angle between the scattered ray (the ray pointing from the collision point to the sound source) and the normal vector of the reflecting surface (collision surface), d is the distance from the collision point to the center of the sound source, and r is the radius of the sound source.
[0271] Furthermore, the energy received by the sound source is also related to the air absorption coefficient. The total propagation distance of the light beam is obtained by adding the cumulative propagation distance of the light beam to the distance between the collision point and the sound source. Based on this total propagation distance and the exponential decay model, the air absorption coefficient α of the current path is calculated. air The energy α received by the sound source is obtained. air (1-αscattering)α s E (Mirror Path) or α air ααscatteringα s E (scattering path).
[0272] Figure 13 is a schematic diagram of an echogram provided in an embodiment of this application, which saves the energy received by the sound source, the total time for light to travel from the listener to the sound source, and the direction of light emitted from the listener into an echogram. As shown in Figure 13, after tracking all light propagation paths, the sound source tracking module 112 counts all propagation paths from the sound source to the listener and saves them as an echogram. Each point in the echogram represents the energy E carried by a possible propagation path i when it reaches the listener.i Propagation time t i And the horizontal and vertical angles θ of the direction of arrival. i ,
[0273] It should be noted that the tracked light rays originate from the listener's position; therefore, the direction of arrival here is the opposite of the emission direction of the light rays originating from the listener's position. Since the reflective surface may have different reflectivities for different frequency signals, the sound source's received energy can also be divided into multiple frequency bands.
[0274] By considering the type of path between the collision point and the sound source location, as well as the air absorption coefficient, the accuracy of the calculated sound source received energy can be improved.
[0275] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes:
[0276] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.
[0277] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;
[0278] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.
[0279] After determining the energy received by the sound source, the light energy can be updated by subtracting the energy received by the sound source from the remaining light energy after reflection. The updated light energy can be used to calculate the next collision point and can be used as the energy carried by the light corresponding to the next collision point.
[0280] Furthermore, it can be determined whether the updated ray energy has reached the lower limit, and whether to continue calculating subsequent collision points for that ray. When the updated ray energy reaches the lower limit, the calculation of subsequent collision points for that collision point is stopped.
[0281] By updating the light energy based on the energy received from the sound source, the accuracy of the calculation results can be improved to avoid calculating every collision point output by the listener tracking module 111.
[0282] Optionally, determining the filter corresponding to the sound source based on the echo map includes:
[0283] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;
[0284] Determine the second filter corresponding to the sound source based on the echo map;
[0285] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.
[0286] Once the echo map is determined, the second filter corresponding to the sound source can be identified, which is the filter obtained in this simulation.
[0287] Figure 14 is a schematic diagram of a filter synthesis module provided in an embodiment of this application. The filter synthesis module 113 synthesizes the echo map generated by the sound source tracking module 112 into an Ambisonic impulse response. The filter synthesis module 113 includes an Ambisonics encoding module 210, a filter generation module 211, a sub-band synthesis module 212, and a filter smoothing module 213. The Ambisonics encoding module 210 first performs Ambisonic encoding on the echo map. Based on the arrival direction of each propagation path, the corresponding Ambisonics coefficients can be obtained. The Ambisonics intensity is obtained by multiplying it by the corresponding sound source received energy. in, It is a spherical harmonic basis function, P l |m| It is a combined Legendre function.
[0288] Figure 15 is an example diagram of reconstructing a frequency band impulse response from an echo map according to an embodiment of this application. The filter generation module 211 processes each Ambisonics channel in the Ambisonics echo map separately. For each channel, an impulse response for the corresponding frequency band is generated based on the propagation time and intensity of each path in different frequency bands. Specifically, the propagation time and corresponding Ambisonics intensity of each point in the echo map are first read. Then, the sampling point position corresponding to the propagation time is found on the impulse response of the corresponding Ambisonics channel and frequency band. Subsequently, a pulse with a peak value equal to the corresponding intensity is inserted at the sampling point, as shown in Figure 15. Depending on actual needs, when the sampling point position corresponding to the propagation time is not an integer, interpolation can also be performed on the corresponding intensity within a certain time range before and after the position, such as using Lagrange interpolation. This application does not impose any restrictions on this.
[0289] The subband synthesis module 212 can reconstruct the full-band impulse response from the subband impulse response using an FIR filter bank or a Linkwitz-Riley filter bank based on a Biquad IIR filter, and this application does not limit this.
[0290] The filter smoothing module 213 smooths the generated filter based on the number of targets (current number of rays). When the number of targets is not equal to N, the randomness of the ray tracing results will increase, which will cause the generated filter to have a certain degree of difference between multiple simulations, thus causing jitter in the final spatialized audio signal.
[0291] To address the aforementioned issues, for each sound source, when determining the corresponding final filter, a smoothing coefficient can be set. The smoothing coefficient α represents the proportion of the previous filter (first filter) retained in the current output, and 1-α represents the proportion of the currently generated filter (second filter) adopted. The final filter is generated by weighted summation of the filters generated in the previous simulation and the current simulation based on the smoothing coefficient.
[0292] The above methods can smooth out the jitter of spatialized audio signals. On the other hand, since this simulation only uses a portion of the rays, the information contained in the ray tracing results may not be comprehensive enough. By retaining a portion of the ray tracing results from the previous simulation, the comprehensiveness of the information contained in the final ray tracing results can be improved.
[0293] Optionally, the smoothing coefficient can be related to the number of targets. When the number of targets is smaller, the smoothing coefficient is larger, indicating that the proportion of the filter obtained from the previous simulation is adopted. Conversely, when the number of targets is larger, the smoothing coefficient is smaller, indicating that the proportion of the filter obtained from the previous simulation is adopted.
[0294] Optionally, when the number of targets equals the maximum number of rays N, α is set to 0, indicating no smoothing is performed to achieve the fastest response speed to changes in the listener's position. When the number of targets is small, a larger smoothing coefficient, such as 0.7, is used to generate a smooth filter.
[0295] By setting a smoothing coefficient related to the target number, and by weighting and summing the filters generated in the previous simulation and the current simulation based on the correlation coefficient, a final filter is generated to smooth out this jitter.
[0296] Optionally, the smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including:
[0297] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.
[0298] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;
[0299] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.
[0300] The compensation coefficient can be determined based on the smoothing coefficient, and is used to compensate for the amplitude attenuation caused by adding two phase-mismatched filters. The smoothing coefficient is a value between 0 and 1, and the compensation coefficient is a value greater than 1.
[0301] The filter smoothing module 213 smooths the filter h generated in the previous simulation. t-1 (First filter) and the filter h generated this time t The second filter is weighted and summed to generate the final filter, as shown in the following equation:
[0302] Optionally, the compensation coefficient can be expressed as
[0303] By first calculating the product of the smoothing coefficient and the first filter, and then the product of the difference between the value 1 and the smoothing coefficient and the second filter, and finally adding the two products together, and then multiplying the sum with the compensation coefficient, the amplitude attenuation caused by the addition of two phase-mismatched filters can be compensated.
[0304] By calculating the compensation coefficient based on the smoothing coefficient, and then calculating the final filter based on the compensation coefficient, the accuracy of the final filter can be improved, and the amplitude attenuation problem caused by the sum of mismatched filters can be eliminated.
[0305] In reality, ray tracing treats sound propagation as equivalent to particle motion, neglecting the diffraction effect of sound waves. Diffraction refers to the phenomenon where sound waves can bypass obstacles and continue propagating. In the real world, when there is an obstruction between a sound source and a listener, sound wave diffraction is a crucial path for sound propagation between them, allowing us to "hear the sound before we see it." Without modeling diffraction, the spatial audio simulated using ray tracing algorithms will lack realism. Therefore, in some scenarios, sound waves may need to travel through a combination of diffraction and reflection to reach the listener, but the methods described above cannot support modeling such combined paths.
[0306] Figure 16 is a schematic diagram of outdoor diffraction provided in an embodiment of this application. In scenarios with weak reflection, such as the outdoor scenario shown in Figure 16, there is obstruction between the sound source and the listener. Diffraction often becomes the main path for sound wave propagation. If diffraction is not modeled, the sound emitted by the sound source may be completely inaudible in the simulated audio when there is obstruction between the sound source and the listener. However, the sound may suddenly reappear when the user or the sound source moves, which will seriously affect the continuity and immersion of the spatial audio. Therefore, modeling diffraction is crucial in the above scenarios.
[0307] One approach to diffraction modeling is to solve the wave equation to model the diffraction of sound waves more accurately. However, for virtual scenes with dynamic interaction and large scale, the computational load required by the above method is enormous, and the running time is difficult to meet the requirements of real-time interaction.
[0308] Another diffraction modeling method involves pre-analyzing the static scene, exhaustively enumerating all possible intermediate diffraction paths from one diffraction edge to another, and calculating the filter corresponding to each path, storing the results in a diffraction path provider. During the rendering phase, based on the positions of the sound source and the listener, the visible diffraction edges of each are located, and all possible diffraction paths are searched from the diffraction path provider to construct a complete diffraction path in the form of "from the sound source through any number of diffraction edges to the listener." The filter corresponding to the complete diffraction path is then calculated, thereby processing the unspatialized audio stream to obtain a simulated diffraction signal.
[0309] The above method saves a significant amount of path-finding work during the rendering process by pre-calculating and generating diffraction path providers. However, it still has some drawbacks: First, it can only pre-analyze and generate diffraction path providers for static parts of the scene. When dynamic geometry exists in the scene, it needs to be re-analyzed when the dynamic geometry moves, rotates, scales up, or shrinks. This analysis process requires a large amount of computation, thus reducing modeling speed. Second, in some scenes, sound waves may need to travel through a combination of diffraction and reflection paths to reach the listener. As shown in Figure 16, sound waves need to travel through a combination of diffraction-reflection-diffraction paths from the sound source to the listener, but the above method cannot support modeling such combined paths. Third, in some complex scenes, the number of diffraction edges may be very large, while the number of diffraction paths increases exponentially relative to the number of diffraction edges. Enumerating all possible diffraction paths is impractical. Therefore, a new diffraction modeling method is needed to solve the above problems.
[0310] Furthermore, ray tracing involves tracking a large number of rays, each of which reflects multiple times within a room, resulting in a significant computational burden. To address this, reverse ray tracing is employed, splitting a complete path into two parts for separate tracking: listener tracking and sound source tracking. This decouples the listener from the sound source. Therefore, when the sound source moves while the listener remains stationary, the tracking results of the first part remain unchanged, allowing the reuse of previous listener tracking results. The listener tracking portion can be omitted, requiring only the sound source tracking portion to run, thus effectively reducing the computational load.
[0311] To address the aforementioned issues, this application considers that modeling diffraction paths should ideally be easily integrated with ray tracing methods. This would generate modeling results that simultaneously incorporate effects such as reflection, scattering, and diffraction, thereby enabling unified processing of unspatialized audio. Therefore, based on the scattering path determined by inverse ray tracing from the sound source to the listener, this application obtains a path related to the listener's position, consisting of reflecting surfaces and diffraction edges, based on the results of inverse ray tracing. This results in a reflection tree, which can be understood as a connectivity model between all reflecting surfaces and diffraction edges related to the listener's current position. This model is independent of the sound source, thus allowing the generation of a combined path of reflection and diffraction paths based on the reflection tree and the sound source position.
[0312] By constructing a reflection tree based on the results of inverse ray tracing, a reflection tree that is only related to the listener's position can be obtained, resulting in a smaller reflection tree and lower computational cost. Furthermore, since the reflection tree includes reflective surfaces, a combined path of reflection and diffraction can be obtained. When the scene model changes, i.e., when dynamic geometry exists, inverse ray tracing is performed again, thus updating the constructed reflection tree simultaneously. In other words, constructing a reflection tree based on inverse ray tracing results does not require additional computation, as modeling scattering paths based on inverse ray tracing is an essential part of the system. Moreover, when dynamic geometry exists in the scene, since inverse ray tracing is performed again, there is no need to re-analyze and construct the reflection tree. In summary, inverse ray tracing can be used to model scattering paths, and the reflection tree constructed based on the results of inverse ray tracing can model a combined path of reflection and diffraction without analyzing the scene model to determine the diffraction path, offering advantages such as lower computational cost and higher accuracy throughout the modeling process.
[0313] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:
[0314] The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.
[0315] Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including:
[0316] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path.
[0317] For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path.
[0318] The ray tracing in this application is inverse ray tracing, which means tracing the light rays emitted from the listener's position. By performing the inverse tracing process, ray tracing can be divided into two stages, thereby decoupling the listener from the sound source. It should be noted that the light rays emitted from the listener's position here do not refer to real light rays, but can be simulated light rays, such as multiple light rays emitted from the listener's position represented by equations. The ray tracing here is used to simulate the propagation of sound.
[0319] The scene in which the user and listener are situated can be represented by a scene model. Optionally, the scene can be a room, in which case the scene model can be used to indicate the size of the room, the position of each object in the room, and / or the reflective surfaces contained in the room and the reflective surfaces of each object, etc.
[0320] Optionally, multiple random light rays can be emitted from the listener's position. These rays will collide with the scene model during propagation, randomly reflecting and scattering. The collision information can be determined using the first-stage tracking process (i.e., listener tracking). This collision information can involve tracking each light ray emitted from the listener's position and recording the collision point when it collides with the scene model. For example, the collision information can include the location of the collision point and the reflective surface it occupies.
[0321] For example, three rays are emitted from the listener's position. For each ray, the corresponding collision information can be determined. For one of the rays, it may collide with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model in sequence. After the condition for ending ray tracing is met, the ray tracing ends. Here, collision point 3 is the last collision point corresponding to the ray, so the information of these three collision points can be determined.
[0322] Tracing the light rays emitted from the listener's position to the last collision point reveals the propagation path from the listener's position to the last collision point, independent of the sound source's location. In other words, regardless of the number or location of the sound sources, the collision information of each light ray remains unchanged. Therefore, the path of the light rays from the listener's position to the last collision point can be traced; this tracing process can be considered listener tracking.
[0323] During the listener tracking process, when each ray is tracked, if the ray collides with the reflecting surface, it is determined whether there is a diffraction edge based on the collision result. The reporting path is then determined based on the reflecting surface or diffraction edge corresponding to the collision point. A reflection tree is then constructed based on the reporting path. The path from the leaf node to the root node in the reflection tree is a possible propagation path of the sound.
[0324] After determining the collision information, ray tracing can continue based on the collision information. Ray tracing here is to trace the path from each collision point to the sound source location. This path is the first propagation path, which is related to the sound source location. When the sound source location is different, the first propagation path is different. This tracing process can be regarded as sound source tracing.
[0325] For a given sound source, based on the location of each collision point, the path from each collision point to the sound source can be determined. Then, based on the path from the listener's location to each collision point, multiple complete propagation paths from the listener to the sound source can be obtained, and it can be equivalently assumed that there is a common propagation path from the sound source location to the listener's location.
[0326] For example, when a ray of light collides sequentially with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model, there are three propagation paths from the sound source position to the listener position. Path 1 is the ray of light from the sound source position through collision point 3 to the listener position; path 2 is the ray of light from the sound source position through collision point 3 and collision point 2 to the listener position; and path 3 is the ray of light from the sound source position through collision point 3, collision point 2, and collision point 1 to the listener position.
[0327] The paths determined above based on the locations of various collision points and sound sources are mostly scattering paths, meaning that the specular reflection of the incident light rays fails to pass through the sound source. Therefore, a ray of light scattered to the sound source is calculated. However, the specular reflection path is a real-world path, and the energy of the specular reflection light rays is relatively high, significantly impacting the listener's perception. Therefore, based on the reflection tree and the sound source location, a combined path of specular reflection and diffraction can be determined from the sound source to the listener, which is the second propagation path.
[0328] For each sound source, the echo map corresponding to the first propagation path and the echo map corresponding to the second propagation path can be determined. The echo map represents the arrival time, direction of arrival, and energy received by the listener for each propagation path from the sound source location to the listener location. For example, if there are three propagation paths from the sound source to the listener, the echo map corresponding to each path can be determined.
[0329] Based on the echo maps corresponding to each sound source, the filter corresponding to that sound source can be determined. Optionally, the filter can be in the form of an Ambisonics filter bank. After generating the filter corresponding to each sound source, a fast convolution method can be used to convolve the unspatialized audio stream of the corresponding sound source based on the filter to obtain the Ambisonics reverberation signal.
[0330] This application proposes an acoustic ray tracing method. First, random diffraction is incorporated into the inverse ray tracing method, and the results of random ray tracing are statistically analyzed. A reflection tree is constructed based on all traced reflection and diffraction nodes. All legal combinations of reflection and diffraction paths are found based on the reflection tree. Finally, all found paths are modeled to obtain a filter. This approach has the following advantages: it can model combinations of reflection and diffraction paths; it does not require prior calculation and analysis of diffraction paths, thus enabling real-time modeling of scenes containing dynamic geometry; it does not require exhaustively enumerating all diffraction edges and paths in the scene, but instead utilizes the ray tracing results to construct and traverse only the traced portions of the reflection tree, reducing computational and storage requirements for complex scenes; the diffraction modeling results are merged with the filters generated by random ray tracing, allowing for unified processing of unspatialized audio streams during the processing stage.
[0331] As shown in Figure 3, the room geometry is considered the scene model in this application. The method of this application can be applied to a processing device, which includes an acoustic ray tracing module and a spatial audio rendering module. Room acoustic parameters, room geometry (also known as the scene model), sound source locations, and listener locations can be input into the processing device. The acoustic ray tracing module processes the input information to obtain filters corresponding to each sound source (including sound diffraction, scattering, and reflection effects). The obtained filters corresponding to each sound source, the unspatialized audio stream in memory, and the listener orientation information provided by the inertial sensor are input into the spatial audio rendering module, enabling the spatial audio rendering module to output a spatialized audio stream and output it to a sound playback device, such as a speaker or headphones. The above steps and the process of determining filters based on the echo map constitute the execution content of the acoustic ray tracing module, while determining the reverberation signal based on the filters constitutes the execution content of the spatial audio rendering module.
[0332] The constructed reflection tree is a listener-related reflection tree, which has the advantages of low computational cost and high accuracy. This enables the modeling of the combined path of reflection and diffraction in the sound propagation process to obtain an accurate reverberation signal.
[0333] This application proposes an acoustic ray tracing method. First, random diffraction is incorporated into the inverse ray tracing method, and the results of random ray tracing are statistically analyzed. A reflection tree is constructed based on all traced reflection and diffraction nodes. All legal combinations of reflection and diffraction paths are found based on the reflection tree. Finally, all found paths are modeled to obtain a filter. This approach has the following advantages: it can model combinations of reflection and diffraction paths; it does not require prior calculation and analysis of diffraction paths, thus enabling real-time modeling of scenes containing dynamic geometry; it does not require exhaustively enumerating all diffraction edges and paths in the scene, but instead utilizes the ray tracing results to construct and traverse only the traced portions of the reflection tree, reducing computational and storage requirements for complex scenes; the diffraction modeling results are merged with the filters generated by random ray tracing, allowing for unified processing of unspatialized audio streams during the processing stage.
[0334] The reflection tree constructed in this application is a listener-related reflection tree, which has the advantages of low computational cost and high accuracy, thereby enabling the modeling of the combined path of reflection and diffraction in the sound propagation process to obtain an accurate reverberation signal.
[0335] Optionally, determine the collision information and reporting path corresponding to the light emitted from the listener's location, and construct a reflection tree based on the reporting path, including:
[0336] When at least one of the following conditions is met, the collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path:
[0337] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.
[0338] The collision information corresponding to the light emitted from the listener's position is determined by the listener tracking process; however, listener tracking and reflection tree construction are not performed every time an update occurs. Optionally, listener tracking and reflection tree construction can be determined based on whether the listener's position changes, whether the scene model changes, or whether the reflectivity of any reflective surface in the scene model changes.
[0339] When the listener's position changes, the collision information when the light ray collides with the scene model will change. Even if the listener's position remains the same, the collision information will still change when the scene model changes. When the reflectivity of any reflective surface in the scene model changes, the remaining energy of the light ray after a collision can be altered, thus determining whether to continue tracking the ray; therefore, the collision information will also change.
[0340] If any one of the three conditions above is met, the listener tracking process and reflection tree construction are executed. Conversely, if none of the three conditions are met, the listener tracking process and reflection tree construction are not run, and the output results of the previous listener tracking and reflection tree construction can be reused directly.
[0341] By determining whether the listener's position, scene model, or the reflectivity of any reflective surface has changed, it is possible to accurately determine whether to perform the listener tracking process and reflectance tree construction. In some scenarios, listener tracking and reflectance tree construction can be omitted, thereby reducing the computational load.
[0342] Figure 17 is a detailed architectural diagram of a spatial audio generation method provided in an embodiment of this application. As shown in Figure 17, the entire implementation scheme is divided into three stages: preprocessing stage 110, simulation stage 120, and processing stage 130. Preprocessing stage 110 runs during system initialization. Its main purpose is to find all diffraction edges in the scene and save the diffraction edge information for use in the simulation stage. Simulation stage 120, based on the room scene model (3D Mesh), the diffraction edge information generated in the preprocessing stage, the acoustic reflectivity data of each reflective surface in the corresponding scene model, and the listener's and sound source's positions, models the acoustic characteristics of the room using ray tracing. The modeling result is the Ambisonics filter bank corresponding to each sound source. Simultaneously, in the simulation stage, a reflection tree is constructed using the ray tracing results. Then, for each sound source, the reflection tree is traversed to search for all legal diffraction and reflection combination paths. The searched paths are modeled and accumulated into the Ambisonics filter bank generated by the ray tracing method. In processing stage 130, the estimated Ambisonics filter bank is used to process the unspatialized audio to obtain an Ambisonics reverberation signal. This room reverberation signal is then added to the direct sound signal processed by HRTF (Head Related Transfer Function) and output to an audio playback device, such as a speaker or headphones.
[0343] Depending on the specific scenario, the simulation phase 120 and the processing phase 130 can have different update frequencies. The processing phase 130 is typically designed to generate a real-time audio stream. For example, when the sampling rate is 48000Hz and 1024 samples are processed each time, the processing phase needs to run at least once every 21.3ms (1024 / 48000*1000) to generate the real-time audio stream. The update frequency of the simulation phase 120 is set according to the complexity of the actual scenario and the system's computing power; for example, it can be set to update once every 100ms. Depending on the system architecture, the simulation phase and the processing phase can also run on different processing devices; this application does not impose any restrictions on this.
[0344] The simulation phase 120 mainly comprises six steps: listener tracking module 121, reflection tree construction module 122, sound source tracking module 123, path search module 124, path modeling module 125, and filter synthesis module 126. Among these, the listener tracking module 121 and reflection tree construction module 122 are only related to the listener's position; therefore, they only need to run when the listener's position changes, the scene model changes, or the reflectivity of the reflecting surface changes. Otherwise, this step can be skipped, and the output of the previous listener tracking module 121 and reflection tree construction module 122 can be reused to directly run the sound source tracking module 123 and path search module 124. The sound source tracking module 123, path search module 124, and path modeling module 125 run separately for each sound source, depending on both the sound source's position and the output of the listener tracking module and reflection tree construction module. Similarly, the filter synthesis module 126 runs separately for each sound source, synthesizing the output of the corresponding sound source tracking module 123 and path modeling module 125 into an Ambisonic filter bank for that sound source.
[0345] Processing stage 130 mainly includes steps such as a fast convolution module 131, an HRTF processing module 132, an Ambisonics rotation module 133, and an Ambisonics decoding module 134. The fast convolution module 131 uses a fast convolution method to convolve the Ambisonics filter bank with the input unspatialized audio signal to obtain an Ambisonics reverberation signal. The Ambisonics rotation module 133 spatially rotates the Ambisonics reverberation signal based on user orientation information provided by an inertial sensor to match the user's current orientation. The Ambisonics decoding module 134 performs binaural decoding on the rotated Ambisonics reverberation signal to obtain a binaural reverberation signal. This signal is mixed with the binaural direct sound signal output by the HRTF processing module 132 to obtain the final binaural signal used for playback.
[0346] The detailed execution process of each of the above stages is explained below.
[0347] Optionally, the scene model is represented using triangulation; the method further includes:
[0348] Determine the diffraction edge in the scene model; the diffraction edge is the coincident side of two non-coplanar triangles;
[0349] Determine and save diffraction edge information; the diffraction edge information corresponds to the diffraction edge; wherein, when the dynamic geometry in the scene model changes, the corresponding diffraction edge information is updated;
[0350] Accordingly, the collision information and reporting path corresponding to the light emitted from the listener's location are determined, including:
[0351] Ray tracing is performed on the light emitted from the listener's location to determine collision information, and the reporting path is determined based on the diffraction edge information.
[0352] When constructing a reflection tree based on the listener tracking results, diffraction edge information is also needed. Diffraction edge information is the result of analyzing static or dynamic geometry in the scene. The information corresponding to the possible diffraction edges in the scene model is stored to accurately determine the reporting path.
[0353] Optionally, the diffraction edge finding module 111 analyzes the input scene model (3D Mesh) represented by triangulation, finds and records all possible diffraction edges, and saves them in the diffraction edge data module 112. Objects in the scene model can include static and dynamic geometry. Since the movement, rotation, scaling, or reduction of dynamic geometry does not affect the existence of diffraction edges, a unified search strategy can be used for both static and dynamic geometry in the scene.
[0354] Figure 18 is a schematic diagram of a diffraction edge provided in an embodiment of this application. A diffraction edge exists when any two triangles in a scene satisfy the condition that they have exactly one overlapping edge and are not coplanar. The determined diffraction edge is this overlapping edge. As shown in Figure 18, triangle 1 and triangle 2 represent reflecting surfaces; in this case, a diffraction edge exists.
[0355] The data recorded in the diffraction edge data module 112 includes diffraction edge information corresponding to each diffraction edge. Optionally, the diffraction edge information may include vertex index, corresponding triangle index, triangle normal vector, diffraction surface vector, diffraction edge length, diffraction angle, etc.
[0356] For static geometry, the diffraction edge information remains unchanged throughout the system's operation. However, for dynamic geometry, when it rotates, translates, enlarges, or shrinks, its triangle normal vector, diffraction surface vector, diffraction edge length, and diffraction angle need to be updated accordingly.
[0357] As shown in Figure 17, the determined diffraction edge information can be transmitted to the listener tracking module 121 and the reflection tree construction module 122 (actually transmitted to the path modeling module 125 for use in determining the diffraction coefficients). The listener tracking module 121 needs to use the diffraction edge information when determining the reporting path to determine whether diffraction has occurred, thereby determining the reporting path.
[0358] By determining the diffraction edge information in the scene model during system initialization, the listener tracking module can accurately determine the reporting path, thereby improving the accuracy of the subsequently determined diffraction path.
[0359] Optionally, the reporting path includes a reflection path and a diffraction path; ray tracing is performed on the light emitted from the listener's position to determine collision information, and the reporting path is determined based on the diffraction edge information, including:
[0360] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0361] When the light ray collides with the scene model, the collision information for this collision is determined;
[0362] When it is determined that the tracking of the ray will not end, the reflection path is determined and the direction of the reflected ray is calculated to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0363] Furthermore, when it is determined that the tracking of the light ray will not end, it is determined whether diffraction will occur based on the diffraction edge information. When diffraction occurs, the diffraction path is determined, and the direction of the diffracted light ray is calculated to determine whether it will collide with another reflective surface in the scene model based on the direction of the diffracted light ray.
[0364] Figure 19 is a schematic diagram of another listener tracking module provided in an embodiment of this application. As shown in Figure 19, during ray tracing, the light emitted from the listener's position is tracked. The light emitted from the listener's position can collide with the scene model. When the light hits a reflective surface, it will be reflected and scattered. According to the reflectivity α of the reflective surface... s The remaining energy α after reflection can be calculated. s E represents the energy currently carried by the ray. The distance between the current collision point and the previous collision point is calculated and summed to obtain the ray's propagation distance (the distance from the listener's position to the current collision point). The number of collisions, the ray's propagation distance, and the remaining ray energy determine whether to terminate the tracking of the current ray.
[0365] If the tracing termination condition is not met, determine the reflection path and calculate the direction of the reflected ray based on the scattering rate of the reflecting surface, then continue calculating the next collision point between the reflected ray and the scene model. If the tracing termination condition is still not met, it is also necessary to determine whether diffraction will occur. If diffraction is determined, the diffraction path can be determined and the direction after diffraction calculated. Repeat the above steps for all rays until all rays meet the tracing termination condition.
[0366] When the number of collisions reaches the upper limit, the propagation distance of the light reaches the upper limit, or the remaining light energy reaches the lower limit, the conditions for ending the tracking of the current light are determined to be met.
[0367] The collision information for each collision can include: the reflecting surface where the collision point is located, the specific collision location, the angle of the incident ray and the angle of the reflected ray, the accumulated distance, etc.; when diffraction occurs after the collision, the collision information can also include the diffraction edge and diffraction point, the angle of the diffracted ray, etc. The above collision information can be transmitted to the sound source tracking module 123.
[0368] As shown in Figure 9, the direction of the reflected light is obtained by combining the direction of specular reflection and the direction of random scattering. If the incident direction is d... incident If the normal vector of the collision surface is n, then the reflection direction of the mirror is d. specular For: d specular =d incident -2*d incident ·n
[0369] Optionally, the random scattering direction dscattering is a randomly generated direction vector that follows a Lambert distribution relative to the normal vector n.
[0370] Based on the scattering rate αscattering of the collision surface, the direction of the reflected light is obtained and normalized, as shown in the following equation: dreflection=αscatteringdscattering+(1-αscattering)d specular
[0371] When diffraction occurs, the direction of the diffracted light rays can be any random direction that meets certain conditions. These conditions are: the shadow region of the diffracted light ray is different from the shadow region of the incident light ray.
[0372] By determining whether diffraction will occur during listener tracking, the diffraction path can be accurately obtained, thereby enabling the accurate construction of the reflection tree.
[0373] Optionally, determining whether diffraction will occur based on the diffraction edge information includes:
[0374] Determine the point of collision between the light ray and the reflecting surface;
[0375] When the collision point is close to the edge of the triangle corresponding to the reflecting surface, it is determined whether the edge is the saved diffraction edge; the collision point is determined to be close to the edge of the triangle corresponding to the reflecting surface based on the transformed centroid coordinates of the collision point.
[0376] If the edge is one of a plurality of diffraction edges that are preserved, then two shadow regions are determined according to the two reflecting surfaces corresponding to the diffraction edge. When the light is in either of the shadow regions, diffraction is determined to occur.
[0377] To enable the light to track possible diffraction paths, a diffraction determination is incorporated at each collision. Figure 20 illustrates a random diffraction pattern provided in an embodiment of this application. As shown in Figure 20, the intersection point of the incident light ray and the reflecting surface, i.e., the collision point, is first located. It is then determined whether this collision point is close to the edge of the triangle (i.e., the reflecting surface). When the collision point is close to the edge of the triangle, it is further determined whether the edge is a diffraction edge. This can be determined by reading pre-stored diffraction edge information. When the collision point is close to a diffraction edge, a diffraction point is created at the position closest to the intersection point on that diffraction edge, and a new diffracted light ray is emitted from this diffraction point, thus continuing to track the diffracted light ray.
[0378] When determining whether a collision point is close to the edge of a triangle, the coordinates of the collision point can be converted to the centroid coordinates of the triangle. The proximity of the collision point to the triangle's edge is then determined by comparing the centroid coordinates with preset values. Optionally, a coordinate system can be established with one side of the triangle as the u-axis and another side as the v-axis. The centroid coordinates can be represented as (u, v). The representation of the centroid coordinates depends on the established triangle's coordinate system. When u is less than a first preset value (actually, u is close to 0), the collision point is close to the v-axis; when v is less than a second preset value (actually, v is close to 0), the collision point is close to the u-axis; when the sum of u and v is close to 1, the collision point is close to the third side.
[0379] A second condition can be used to determine whether diffraction will occur. Figure 21 is a top view of a random diffraction pattern provided in an embodiment of this application. As shown in Figure 21, an example of diffracted light generation is illustrated from another angle. For each diffraction edge, the two triangles that make up it are extended outward from the diffraction edge to form two shaded regions as shown in Figure 21. To simplify the calculation, only diffraction between the two shaded regions can be considered. For example, light can only diffract from shaded region 1 to shaded region 2, or vice versa. If the incident light is in other regions, no new diffracted light is generated. Generating diffracted light based on this rule allows limited computational power to be allocated as much as possible to diffraction paths that have a significant impact on the listener's hearing, i.e., paths from one shaded region to another. For example, when the incident light is in the blank area in the lower right corner, reflection will occur, and the reflected light can propagate to shaded region 1. Even without considering diffracted light, it will not have a significant impact on the listener's hearing.
[0380] By determining whether diffraction occurs as described above, the accuracy of determining whether diffraction will occur during a collision can be improved, thereby making the determined diffraction path more accurate.
[0381] Figure 22 is a schematic diagram of a path reporting method provided in an embodiment of this application. As shown in Figure 22, the reported path includes a reflection path and a diffraction path. As shown in Figure 17, in the listener tracking module 121, each time a light ray collides and diffracts, the current path is reported to the reflection tree construction module 122. The reported path is in the form of {start node, end node}, indicating that these two nodes can affect the information received by the listener and that they are interconnected. For example, when a light ray starts from the listener's position and collides with reflecting surface A, the reported path is {listener, reflecting surface A}. If the light ray also diffracts at the diffraction edge B, an additional path {listener, diffraction edge B} is reported. If a light ray emitted from reflecting surface A continues to propagate, collides with reflecting surface C, and diffracts at the diffraction edge D, the reported paths are {reflecting surface A, reflecting surface C}, {reflecting surface A, diffraction edge D}. The light emitted from the diffraction edge B continues to propagate and collides with the reflecting surface E, and diffracts at the diffraction edge F. The reported path is {diffraction edge B, reflecting surface E}, {diffraction edge B, diffraction edge F}.
[0382] The ray diffraction in the listener tracking module only provides basic path possibilities. That is, each ray can be transmitted based on each reported path, but the specific combination path of diffraction and reflection paths from the sound source location to the listener location (that is, the sound wave diffraction behavior) needs to be determined based on the reflection tree.
[0383] The reflection tree construction module 122 abstracts all reflecting surfaces and diffraction edges into nodes, and several nodes connected together can form a sound wave transmission path. When the scene model is large, processing and storing all nodes is impractical because the total number of paths they form increases exponentially. It is also unnecessary, as only a small portion of the nodes may affect the current listener's location. This is why this application uses the calculation results of the listener tracking module 121 to track which nodes the light rays originating from the listener pass through, and constructs the reflection tree only for these nodes.
[0384] As shown in Figure 17, the output of the listener tracking module 121 is also input to the sound source tracking module 123, so that the sound source tracking module 123 can determine multiple first propagation paths from the sound source position to the listener position based on the collision information and the sound source position. Optionally, the sound source tracking module 123 can generate the propagation path and echo map from the sound source to the listener based on the sound source position, scene model, and reflectivity information. The tracked light rays originate from the listener position, so what is actually calculated is the propagation path from the listener position to the sound source position, and it is equivalent to assuming that there is a common propagation path from the sound source position to the listener position.
[0385] Specifically, a set of light rays is randomly emitted from the listener's position, each carrying the same initial energy E0. However, the sound source tracking module 123 no longer needs to calculate the intersection between the light rays and the scene model, because the output of the listener tracking module 121 already contains this collision information. For each collision point, it is only necessary to calculate the propagation path between it and the sound source position.
[0386] Figure 23 is a schematic flowchart of another sound source tracking module provided in this application embodiment. The execution steps in Figure 23 determine the echo map corresponding to the first propagation path. For the light ray corresponding to the collision point, it can be determined whether diffraction occurs at the collision point. If diffraction occurs, the light energy is updated, i.e., the energy of the diffracted light ray. If no diffraction occurs, it can also be determined whether there is an obstruction between the collision point and the sound source location. If there is no obstruction, the path between the collision point and the sound source location can be determined. If there is an obstruction, the next collision point is processed. When the path between the collision point and the sound source location is determined, the propagation path between the corresponding sound source location and the listener location can be obtained, and thus multiple paths between the sound source location and the listener location can be obtained. For the propagation path between the sound source location and the listener location, the echo map can be calculated, including: the energy received by the sound source, the arrival time and direction of the light ray reaching the sound source. When calculating the echo map, it can be determined based on the reflectivity and scattering rate of the reflecting surface where the collision point is located.
[0387] As shown in Figures 11 and 12, specifically, firstly, based on the reflectivity α of the reflecting surface... s The remaining energy α after reflection is obtained. s E, where E is the energy currently carried by the ray. α s E is further divided into two parts: mirror energy and scattering energy, where the scattering energy is αscatteringα. s E, αscattering is the scattering coefficient of the current collision surface, and the remaining part is (1-αscattering)α sE is the mirror energy. When there are no other obstructions between the collision point and the sound source, there exists a propagation path between them, and its direction originates from the location corresponding to the collision point. The mirror reflection angle of this ray is determined to see if the mirrored ray can pass through the sound source. If so, the path is considered a mirror reflection path, and the mirror energy of this ray is (1-αscattering)α. s If all energy E is emitted to the sound source, then the path is considered a scattering path, and a portion of the scattered energy of the light ray is emitted to the sound source. Assuming the scattering proportion is α, then the energy emitted to the sound source is αscatteringα. s The value of E.α can be given by the following formula:
[0388] Where θ is the angle between the scattered ray (the ray pointing from the collision point to the sound source) and the normal vector of the reflecting surface, d is the distance from the collision point to the center of the sound source, and r is the radius of the sound source.
[0389] The cumulative propagation distance of the current light ray is added to the distance between the collision point and the sound source to obtain the total distance the light ray travels from the listener to the sound source. Based on this distance and the exponential attenuation model, the air absorption coefficient α of the current path is calculated. air This energy is then multiplied by the energy emitted from the collision point to the sound source to obtain the energy α received by the sound source. air (1-αscattering)α s E (Mirror Path) or α air ααscatteringα s E (scattering path).
[0390] The energy received by the sound source, the total time it takes for the light ray to travel from the listener to the sound source, and the direction in which the light ray is emitted from the listener are saved to the echo map (since it is a reverse ray tracing, the direction of arrival is the opposite of the direction in which the light ray is emitted from the listener).
[0391] When light rays collide, if diffraction occurs simultaneously, it is necessary to continue tracking the reflected and diffracted rays. After calculating the energy received by the sound source, the energy received by the sound source, the energy of the diffracted rays, and the energy absorbed by the reflecting surface are subtracted from the energy carried by the light rays to obtain the updated light energy, which is the energy of the reflected rays after the collision. This updated light energy is used to determine if it reaches a lower limit. If it does not reach the lower limit, this energy is used for calculation at the next collision point, thus continuing to track the reflected rays. On the other hand, if diffraction occurs, the energy of the diffracted rays can also be obtained as the updated light energy, which is the energy of the diffracted rays after the collision. This updated light energy is also used to determine if it reaches a lower limit. If it does not reach the lower limit, this energy is used for calculation at the next collision point, thus continuing to track the diffracted rays. Optionally, the energy of the diffracted rays can be the product of the energy carried by the light rays and a fixed coefficient, which is not specifically limited here.
[0392] Determine if the updated ray energy has reached the lower limit and decide whether to continue calculating subsequent collision points for that ray. There's no need to calculate the intersection of the new ray with the scene model here, as the reflection direction and the next collision point have already been calculated in the sound source tracking module. You can directly begin calculating the propagation path between the next collision point and the sound source.
[0393] Through the sound source tracing process described above, the first propagation path from the sound source to the listener can be obtained. This first propagation path is mostly a scattering path, and the probability of it being a specular reflection path is relatively small. Furthermore, the above process does not determine the diffraction path. Therefore, a combined path of specular reflection and diffraction can be modeled based on a reflection tree. This part will be explained in detail below.
[0394] Optionally, a reflection tree is constructed based on the reported path, including:
[0395] The reported path is deduplicated to obtain a deduplicated path; the reported path includes a start node and an end node; the start node or the end node is a reflecting surface or a diffraction edge;
[0396] The listener is identified as the root node of the reflection tree, and the following steps are repeated until the maximum depth of the reflection tree reaches a preset depth:
[0397] For each node in the reflection tree, find the target node that the node can connect to from all the deduplicated paths, and determine the target node as the child node of the node.
[0398] When constructing a reflection tree based on the reported paths, all unique paths can be counted from the reported paths, i.e., deduplication can be performed, and then the reflection tree can be constructed. Figure 22 shows the form of the reported paths, including the start node and the end node.
[0399] Figure 24 is a schematic diagram of a reflection tree provided in an embodiment of this application. The reflection tree is a tree structure, with the root node representing the listener. Each node has several child nodes, representing a node in the sound propagation path. This node can be a reflecting surface or a diffraction edge. For each leaf node in the reflection tree, tracing back the path to the root node corresponds to a possible sound propagation path in the scene model. Specifically, when constructing the reflection tree, the listener is first determined as the root node. Then, a target node that the root node can connect to is found from the deduplicated paths, and this target node is determined as a child node of the root node. The above steps are repeated for the child nodes of the root node to determine the reflection tree. That is, when constructing the reflection tree, each node is found to connect to from all deduplicated paths, and a corresponding child node is added to each node. This step is repeated for each child node until the maximum depth of the reflection tree reaches a preset depth. The value of the preset depth is not specifically limited here.
[0400] By determining the reflection tree based on the reporting nodes, each sound propagation path can be intuitively identified, which facilitates the subsequent determination of a combined path containing reflection and diffraction paths based on the reflection tree.
[0401] Optionally, the method further includes:
[0402] When determining the reflection tree, if the node is the root node, the spatial position corresponding to the root node is determined as the listener's position;
[0403] When the node is a reflective surface, determine the mirror position of the parent node of the node relative to the reflective surface, and determine the mirror position as the spatial position of the node.
[0404] When the node is a diffraction edge, the center position of the diffraction edge is determined as the diffraction point or the spatial position corresponding to the node.
[0405] While constructing the reflection tree, the spatial position of each node can be calculated, thus determining the reflection point position based on the spatial position of each node. The spatial position of the root node is the listener's position. For other nodes, when a node is a reflecting surface, its spatial position is the mirror image of its parent node relative to that reflecting surface. For nodes corresponding to diffraction edges, the center of the diffraction edge is determined as the spatial position of that node.
[0406] Figure 25 is a schematic diagram of a mirror position and propagation path provided in an embodiment of this application. As shown in Figure 25, sound propagates from the N+2 level node to the N+1 level node. The N+1 level node is the diffraction edge. The diffracted light propagates to the N level node, which is a reflecting surface. After reflection by the reflecting surface, the light propagates to the N-1 level node, which is the parent node of the N level node. For the N level node, since it is a reflecting surface, the spatial position of this node is the mirror position of the spatial position of its parent node relative to the reflecting surface. This mirror position is the spatial position of the N level node.
[0407] In this way, when performing path modeling, the actual path of sound propagation from the child node to the parent node after reflection by the current reflecting surface can be equivalent to the propagation path from the child node position to the current reflecting node position (the parent node's mirror image). For nodes on other diffraction edges, the center position of the diffraction edge is directly used as the spatial position of the current node.
[0408] By determining the spatial location of each node, it becomes easier to determine the accurate propagation path later, thus eliminating the need for further calculations during the path search phase.
[0409] Optionally, the reflection path in the second propagation path is a specular reflection path; multiple second propagation paths are determined from the sound source location to the listener location based on the reflection tree, including:
[0410] For each sound source, determine the reflecting surface and diffraction edge where the light emitted from the sound source location collides with the scene model for the first time, so as to obtain a list of visible nodes of the sound source.
[0411] For each leaf node in the reflection tree, when the leaf node is in the sound source visible list, it is determined whether the propagation path corresponding to the leaf node is a valid path;
[0412] When the leaf node is not in the visible list of sound sources, the propagation path corresponding to the leaf node is determined to be an invalid path;
[0413] The second propagation path is determined based on the established legal path; the second propagation path is a path obtained by sequentially connecting the sound source location, the reflection points or diffraction points corresponding to each node in the legal path, and the listener's location.
[0414] After constructing the reflection tree, for each sound source, the path search module 124 searches for all legal paths. Figure 26 is a schematic diagram of the specific process of a path search module provided in an embodiment of this application. As shown in Figure 26, firstly, random light rays are emitted from the sound source, and the reflecting surfaces and diffraction edges where the light rays collide with the scene model for the first time are counted, resulting in a unique list of visible sound source nodes. The nodes in this list are the visible reflecting surfaces or diffraction edges of the sound source. Subsequently, for each leaf node in the reflection tree, it is first determined whether the leaf node is in the list of visible sound source nodes. If it is not in the list, the sound propagation path corresponding to the leaf node is an invalid path. If it is in the list, it is further determined whether the path corresponding to the leaf node is a valid path.
[0415] By determining a list of visible sound source nodes, and then identifying the leaf nodes within that list, we can further determine whether the sound propagation path corresponding to that leaf node is a valid path. This approach can improve the accuracy of the determined second propagation path while reducing computational complexity.
[0416] Optionally, determining whether the propagation path corresponding to the leaf node is a valid path includes:
[0417] If every child node from the leaf node to the root node satisfies the target condition, then the path from the leaf node to the root node is determined to be a valid path; otherwise, it is an invalid path.
[0418] When the node is a reflective surface, the target condition is: the reflection point corresponding to the reflective surface is within the triangle corresponding to the reflective surface and the corresponding path is not obstructed; the position of the reflection point is related to the spatial position corresponding to the reflective surface and the spatial position of the next level node.
[0419] When the node is a diffraction edge, the target conditions are: the corresponding path is not occluded, and the preceding and following nodes of the node are located in two different shadow areas corresponding to the node.
[0420] When determining whether the propagation path corresponding to a leaf node is a valid path, the leaf node can be designated as the current node. If the current node is a reflection node (reflecting surface), the reflection point position is determined, as shown in Figure 25. The reflection point position is the intersection of the line connecting the spatial position of the next-level node and the spatial position of the current node with the current reflecting surface. This reflection point position is the intersection of the sound propagation path and the reflecting surface. After determining the reflection point, it is then judged that it must be within the triangle corresponding to the reflecting surface to be a valid path. In addition, the path that can be formed by the spatial position of the next-level node and the reflection point position can also be judged whether the path is occluded by other objects in the scene model. If it is not occluded, the above steps are continued to be performed on the parent node of the current node until all child nodes from the leaf node to the root node satisfy the condition that the reflection point is within the triangle corresponding to the reflecting surface and the corresponding path is not occluded by other objects. Then, the propagation path corresponding to the leaf node is determined to be a valid path. In other words, a complete propagation path is formed by connecting the reflection points and diffraction points corresponding to all propagation nodes in the propagation path. Only when the propagation path is not blocked can it be determined as a legal path.
[0421] When a node is a diffraction edge, it is necessary not only to determine that the corresponding path is not occluded, but also to determine that the node's predecessor and successor nodes are located in two different shadow regions corresponding to the current node. The two different shadow regions corresponding to the current node refer to the two different shadow regions corresponding to the diffraction edge. By restricting this condition, we can focus only on diffraction paths that have a significant impact on the auditory experience, and not on diffraction paths that have a smaller impact on the auditory experience, thereby reducing the amount of computation.
[0422] By determining whether the reflection point is located within the triangle corresponding to the reflecting surface, and whether there is any obstruction along the corresponding path, the accuracy of determining whether the propagation path is a legitimate path can be improved.
[0423] Optionally, determining the echo map corresponding to the second propagation path includes:
[0424] For the sound source, determine the initial energy corresponding to the sound source;
[0425] For each second propagation path, the distance attenuation coefficient and air absorption coefficient are determined based on the total length of the second propagation path. The reflection attenuation coefficient of each reflecting surface in the second propagation path is determined. The reflection attenuation coefficients of each reflecting surface are multiplied to obtain the overall reflection attenuation coefficient. The diffraction coefficient is determined based on the diffraction edge information of each diffraction edge in the second propagation path. The reflection attenuation coefficient of the reflecting surface is related to the reflectivity and scattering rate of the reflecting surface.
[0426] The multiplication result of the initial energy, the distance attenuation coefficient, the air absorption coefficient, the overall reflection attenuation coefficient, and the diffraction coefficient is determined as the energy received by the listener;
[0427] The direction in which the last node points to the listener's position is defined as the receiving direction;
[0428] The energy received by the listener, the receiving direction, and the propagation time corresponding to the second propagation path are determined as the echo map corresponding to the second propagation path.
[0429] The path modeling module 125 models all the second propagation paths searched by the path search module 124 and obtains the echo map corresponding to each second propagation path. This result is merged with the echo map obtained by the sound source tracing module 123 and used for subsequent filter synthesis. For each propagation path, the listener's received energy, arrival time, and direction of arrival are saved to the echo map.
[0430] Figure 27 is a schematic diagram of a diffraction path modeling method provided in an embodiment of this application. As shown in Figure 27, the left half is the second propagation path searched in the reflection tree, and the right half is the actual sound propagation path corresponding to the second propagation path. The path modeling module 125 is used to calculate the arrival time, arrival direction, and listener received energy for each second propagation path.
[0431] Modeling the energy received by the listener mainly consists of several parts: distance attenuation, air absorption, reflection attenuation, and diffraction modeling. Among these, the distance attenuation coefficient α... d The total length d of the propagation path can be obtained using an inverse distance model. The air absorption coefficient α air It can be calculated using a distance and exponential attenuation model. The reflection attenuation coefficient of each reflecting surface can be calculated from the reflectivity α of that reflecting surface. s The scattering rate α is obtained by the following equation: (1-αscattering)α sThe overall reflection attenuation coefficient αreflection is obtained by multiplying the reflection attenuation coefficients of all reflecting surfaces along the second propagation path. Modeling diffraction can be done using a unified theory of diffraction (UTD) or a biot–Tolstoy model (BTM), and this application is not limited to either. The UTD or BTM method requires diffraction edge information (stored in the diffraction edge data 112), such as triangle normal vectors, diffraction surface vectors, diffraction edge lengths, and diffraction angles. Based on the diffraction edge information corresponding to each diffraction edge in the second propagation path, the diffraction coefficient αdiffraction corresponding to that second propagation path can be determined, representing the ratio of diffracted energy to incident energy. Furthermore, the direction from the last node to the listener's position can be determined as the receiving direction, i.e., the direction of arrival.
[0432] The energy ultimately received by the listener can be represented as α. d α air αreflectionαdiffractionE1, where E1 represents the initial energy corresponding to the sound source. The ratio of the initial energy E1 corresponding to the sound source to E0 in the sound source tracing module 123 reflects the proportion of the two modeling methods, reflection tree and random ray tracing, in the final result. This ratio can be adjusted according to the actual situation. In a specific embodiment, the calculation of the above attenuation coefficient can also be divided into different frequency bands to obtain more accurate modeling results.
[0433] As shown in Figure 13, the echoogram generated by the sound source tracing module 123 and the path modeling module 125 records all reflection, scattering, or diffraction paths from the sound source to the listener. Each point in the echoogram represents the energy E carried by a possible propagation path i when it reaches the listener. i Propagation time t i And the horizontal and vertical angles θ of the direction of arrival. i , Because the reflectivity of the reflecting surface may vary for signals of different frequencies, this energy value can also be divided into multiple frequency bands.
[0434] By taking into account factors such as distance attenuation coefficient, air absorption coefficient, overall reflection attenuation coefficient, and diffraction coefficient, the accuracy of the determined listener's received energy is relatively high.
[0435] Figure 28 is a schematic diagram of another filter synthesis module provided in the embodiment of this application. As shown in Figure 28, the filter synthesis module 126 synthesizes the echo map into an Ambisonic impulse response, which is divided into several steps: Ambisonics encoding 210, filter generation 211, and subband synthesis 212.
[0436] The Ambisonics encoding module 210 first performs Ambisonic encoding on the echo map. Based on the arrival direction of each propagation path, the corresponding Ambisonics coefficients can be obtained. Multiplying this by the corresponding energy yields the Ambisonics strength:
[0437] in, These are spherical harmonic basis functions:
[0438] Among them, P l |m| It is a combined Legendre function.
[0439] As shown in Figure 15, the filter generation module 211 processes each Ambisonics channel in the Ambisonics echo map separately. For each channel, based on the propagation time and intensity of each path in different frequency bands, an impulse response for the corresponding frequency band is generated. Specifically, the propagation time and corresponding Ambisonics intensity of each point in the echo map are first read. Then, the sampling point position corresponding to the propagation time is found in the impulse response of the corresponding Ambisonics channel and frequency band. Subsequently, a pulse with a peak value equal to the corresponding intensity is inserted at this sampling point. Depending on actual needs, when the sampling point position corresponding to the propagation time is not an integer, interpolation can also be performed on the corresponding intensity within a certain time range before and after that position, such as using Lagrange interpolation. This application does not impose any restrictions on this.
[0440] The subband synthesis module 212 can reconstruct the full-band impulse response from the subband impulse response using an FIR filter bank or a Linkwitz-Riley filter bank based on a Biquad IIR filter, and this application does not limit this.
[0441] Figure 29 is a schematic diagram of an acoustic ray tracing device provided in an embodiment of this application. The device 160 includes:
[0442] The listener tracking module 1601 is configured to determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times.
[0443] The sound source tracking module 1602 is configured to, for each sound source, determine each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path.
[0444] The filter synthesis module 1603 is configured to determine a filter corresponding to the sound source based on the echo map;
[0445] Processing module 1604 is configured to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.
[0446] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically configured as follows:
[0447] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:
[0448] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.
[0449] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically configured as follows:
[0450] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;
[0451] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.
[0452] Optionally, the device further includes: a determination module, configured to:
[0453] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;
[0454] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.
[0455] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.
[0456] Optionally, the rate of change of the reflective surface includes the addition rate and the deletion rate; when the determination module determines whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions, it is specifically configured as follows:
[0457] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.
[0458] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:
[0459] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.
[0460] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.
[0461] Optionally, the device further includes: a trigger module configured to:
[0462] The external reset condition is determined to be triggered when at least one of the following conditions is met:
[0463] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.
[0464] Optionally, when the listener's position does not change significantly, the listener tracking module 1601 is specifically configured to determine the collision information of the target number of light rays as follows:
[0465] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.
[0466] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.
[0467] Optionally, when the listener tracking module 1601 selects the target number of light rays from the N light rays emitted from the listener's location, it is specifically configured as follows:
[0468] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.
[0469] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.
[0470] Optionally, when determining the collision information of the target number of light rays, the listener tracking module 1601 is specifically configured as follows:
[0471] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0472] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;
[0473] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.
[0474] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0475] Optionally, when determining the direction of the reflected light based on the scattering rate of the reflective surface, the listener tracking module 1601 is specifically configured as follows:
[0476] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;
[0477] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.
[0478] Optionally, the collision information includes: the location of the collision point and the reflecting surface where the collision point is located; when the sound source tracking module 1602 determines each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determines the echo map corresponding to the sound source based on each propagation path, it is specifically configured as follows:
[0479] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.
[0480] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.
[0481] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.
[0482] Optionally, the collision information further includes: the specular reflection angle of the light and the cumulative propagation distance; when the sound source tracking module 1602 determines the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, it is specifically configured as follows:
[0483] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.
[0484] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.
[0485] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.
[0486] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.
[0487] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the sound source tracking module 1602 is further configured to:
[0488] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.
[0489] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;
[0490] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.
[0491] Optionally, when determining the filter corresponding to the sound source based on the echo map, the filter synthesis module 1603 is specifically configured as follows:
[0492] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;
[0493] Determine the second filter corresponding to the sound source based on the echo map;
[0494] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.
[0495] Optionally, the smoothing coefficient is a value between 0 and 1; when the filter synthesis module 1603 determines the final filter generated in this simulation based on the smoothing coefficient, the first filter, and the second filter, it is specifically configured as follows:
[0496] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.
[0497] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;
[0498] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.
[0499] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically configured as follows:
[0500] Determine the collision information and reporting path corresponding to the light emitted from the listener's location;
[0501] The device further includes a reflection tree construction module, configured to:
[0502] A reflection tree is constructed based on the reported path; the reported path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.
[0503] The sound source tracking module 1602 is specifically configured to, for each sound source, determine various propagation paths from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path, as follows:
[0504] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and the echo map corresponding to the first propagation path is determined.
[0505] The device further includes a path search module, configured to:
[0506] Based on the reflection tree, multiple second propagation paths are determined from the sound source location to the listener location; the second propagation path is a combination of a reflection path and a diffraction path.
[0507] The device further includes a path modeling module configured to determine the echo map corresponding to the second propagation path;
[0508] Optionally, when the listener tracking module 1601 determines the collision information and reporting path corresponding to the light emitted from the listener's location, and when the reflection tree construction module constructs a reflection tree based on the reported path, it is specifically configured as follows:
[0509] When at least one of the following conditions is met, the listener tracking module 1601 performs ray tracing on the light rays emitted from the listener's location to determine collision information and a reporting path, and the reflection tree construction module constructs a reflection tree based on the reported path:
[0510] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.
[0511] Optionally, the scene model is represented using triangulation; the device further includes a preprocessing module configured to:
[0512] Determine the diffraction edge in the scene model; the diffraction edge is the coincident side of two non-coplanar triangles;
[0513] Determine and save diffraction edge information; the diffraction edge information corresponds to the diffraction edge; wherein, when the dynamic geometry in the scene model changes, the corresponding diffraction edge information is updated;
[0514] Accordingly, when the listener tracking module 1601 determines the collision information and reporting path corresponding to the light emitted from the listener's location, it is specifically configured as follows:
[0515] Ray tracing is performed on the light emitted from the listener's location to determine collision information, and the reporting path is determined based on the diffraction edge information.
[0516] Optionally, the reporting path includes a reflection path and a diffraction path; when the listener tracking module 1601 performs ray tracing on the light emitted from the listener's position, determines collision information, and determines the reporting path based on the diffraction edge information, it is specifically configured as follows:
[0517] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:
[0518] When the light ray collides with the scene model, the collision information for this collision is determined;
[0519] When it is determined that the tracking of the ray will not end, the reflection path is determined and the direction of the reflected ray is calculated to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.
[0520] Furthermore, when it is determined that the tracking of the light ray will not end, it is determined whether diffraction will occur based on the diffraction edge information. When diffraction occurs, the diffraction path is determined, and the direction of the diffracted light ray is calculated to determine whether it will collide with another reflective surface in the scene model based on the direction of the diffracted light ray.
[0521] Optionally, when determining whether diffraction will occur based on the diffraction edge information, the listener tracking module 1601 is specifically configured as follows:
[0522] Determine the point of collision between the light ray and the reflecting surface;
[0523] When the collision point is close to the edge of the triangle corresponding to the reflecting surface, it is determined whether the edge is the saved diffraction edge; the collision point is determined to be close to the edge of the triangle corresponding to the reflecting surface based on the transformed centroid coordinates of the collision point.
[0524] If the edge is one of a plurality of diffraction edges that are preserved, then two shadow regions are determined according to the two reflecting surfaces corresponding to the diffraction edge. When the light is in either of the shadow regions, diffraction is determined to occur.
[0525] Optionally, the reflection tree construction module is specifically configured to construct the reflection tree based on the reported path as follows:
[0526] The reported path is deduplicated to obtain a deduplicated path; the reported path includes a start node and an end node; the start node or the end node is a reflecting surface or a diffraction edge;
[0527] The listener is identified as the root node of the reflection tree, and the following steps are repeated until the maximum depth of the reflection tree reaches a preset depth:
[0528] For each node in the reflection tree, find the target node that the node can connect to from all the deduplicated paths, and determine the target node as the child node of the node.
[0529] Optionally, the device further includes: a spatial location determination module, configured to:
[0530] When determining the reflection tree, if the node is the root node, the spatial position corresponding to the root node is determined as the listener's position;
[0531] When the node is a reflective surface, determine the mirror position of the parent node of the node relative to the reflective surface, and determine the mirror position as the spatial position of the node.
[0532] When the node is a diffraction edge, the center position of the diffraction edge is determined as the diffraction point or the spatial position corresponding to the node.
[0533] Optionally, the reflection path in the second propagation path is a specular reflection path; when the path search module determines multiple second propagation paths from the sound source location to the listener location based on the reflection tree, it is specifically configured as follows:
[0534] For each sound source, determine the reflecting surface and diffraction edge where the light emitted from the sound source location collides with the scene model for the first time, so as to obtain a list of visible nodes of the sound source.
[0535] For each leaf node in the reflection tree, when the leaf node is in the sound source visible list, it is determined whether the propagation path corresponding to the leaf node is a valid path;
[0536] When the leaf node is not in the visible list of sound sources, the propagation path corresponding to the leaf node is determined to be an invalid path;
[0537] The second propagation path is determined based on the established legal path; the second propagation path is a path obtained by sequentially connecting the sound source location, the reflection points or diffraction points corresponding to each node in the legal path, and the listener's location.
[0538] Optionally, when determining whether the propagation path corresponding to the leaf node is a valid path, the path search module is specifically configured as follows:
[0539] If every child node from the leaf node to the root node satisfies the target condition, then the path from the leaf node to the root node is determined to be a valid path; otherwise, it is an invalid path.
[0540] When the node is a reflective surface, the target condition is: the reflection point corresponding to the reflective surface is within the triangle corresponding to the reflective surface and the corresponding path is not obstructed; the position of the reflection point is related to the spatial position corresponding to the reflective surface and the spatial position of the next level node.
[0541] When the node is a diffraction edge, the target conditions are: the corresponding path is not occluded, and the preceding and following nodes of the node are located in two different shadow areas corresponding to the node.
[0542] Optionally, the path modeling module is specifically configured to determine the echo map corresponding to the second propagation path as follows:
[0543] For the sound source, determine the initial energy corresponding to the sound source;
[0544] For each second propagation path, the distance attenuation coefficient and air absorption coefficient are determined based on the total length of the second propagation path. The reflection attenuation coefficient of each reflecting surface in the second propagation path is determined. The reflection attenuation coefficients of each reflecting surface are multiplied to obtain the overall reflection attenuation coefficient. The diffraction coefficient is determined based on the diffraction edge information of each diffraction edge in the second propagation path. The reflection attenuation coefficient of the reflecting surface is related to the reflectivity and scattering rate of the reflecting surface.
[0545] The multiplication result of the initial energy, the distance attenuation coefficient, the air absorption coefficient, the overall reflection attenuation coefficient, and the diffraction coefficient is determined as the energy received by the listener;
[0546] The direction in which the last node points to the listener's position is defined as the receiving direction;
[0547] The energy received by the listener, the receiving direction, and the propagation time corresponding to the second propagation path are determined as the echo map corresponding to the second propagation path.
[0548] The acoustic ray tracing device 160 provided in this application embodiment can implement the acoustic ray tracing method shown in FIG2 above. Its implementation principle and technical effect are similar, and will not be described again here.
[0549] Figure 30 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. As shown in Figure 30, the electronic device provided in this embodiment includes at least one processor 1701 and a memory 1702. The processor 1701 and the memory 1702 are connected via a bus 1703.
[0550] In a specific implementation, at least one processor 1701 executes computer execution instructions stored in memory 1702, causing at least one processor 1701 to execute the method in the above method embodiment.
[0551] The specific implementation process of processor 1701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0552] Optionally, the electronic device can be a head-mounted display device.
[0553] In the embodiment shown in Figure 30 above, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0554] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.
[0555] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0556] This application also provides a wearable device, including a processing unit; the processing unit is used to implement the method of the above method embodiments.
[0557] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in the above-described method embodiments.
[0558] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the above method embodiments.
[0559] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0560] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0561] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0562] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0563] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0564] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
An acoustic ray tracing method, characterized in that, include: Determine the collision information corresponding to the light emitted from the listener's location; The collision information is the information of each collision point obtained by the light ray colliding with the scene model multiple times; For each sound source, each propagation path from the sound source location to the listener location is determined based on the collision information and the sound source location, and the echo map corresponding to the sound source is determined based on each propagation path. Determine the filter corresponding to the sound source based on the echo map; The unspatialized audio stream of each sound source is processed according to the filter corresponding to each sound source to determine the reverberation signal. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met: The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N; Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N. The method according to claim 3, characterized in that, The method further includes: Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing; The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations. The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions. The method according to claim 4, characterized in that, The rate of change of the reflective surface includes the addition rate and the deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes: When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly. And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met: The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered. The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations. The method according to claim 4 or 5, characterized in that, The method further includes: The external reset condition is determined to be triggered when at least one of the following conditions is met: The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations. The method according to claim 5, characterized in that, When the listener's position undergoes a non-significant change, the collision information for determining the target number of light rays includes: The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface. Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays. The method according to claim 7, characterized in that, Selecting the target number of light rays from the N light rays emitted from the listener's location includes: When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained. When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays. The method according to claim 7 or 8, characterized in that, Determining the collision information of the target number of light rays includes: For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met: When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined; The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions. When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray. The method according to claim 9, characterized in that, Determining the direction of the reflected light based on the scattering rate of the reflecting surface includes: The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface; Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction. The method according to any one of claims 1-10, characterized in that, The collision information includes: the location of the collision point and the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including: For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener. The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point. For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path. The method according to claim 11, characterized in that, The collision information also includes: the specular reflection angle of the light beam and the cumulative propagation distance; determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including: The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light. The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location. When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light. When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source. The method according to claim 11 or 12 is characterized in that, After determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes: The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source. When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map; When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped. The method according to any one of claims 3-10, characterized in that, Determining the filter corresponding to the sound source based on the echo map includes: The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation; Determine the second filter corresponding to the sound source based on the echo map; For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number. The method according to claim 14, characterized in that, The smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including: The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient. Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter; The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point. Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including: For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path. For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path. The method according to claim 16, characterized in that, Determine the collision information and reporting path corresponding to the light emitted from the listener's location, and construct a reflection tree based on the reporting path, including: When at least one of the following conditions is met, the collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path: The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes. The method according to claim 16, characterized in that, The scene model is represented using triangulation; the method further includes: Determine the diffraction edge in the scene model; the diffraction edge is the coincident side of two non-coplanar triangles; Determine and save diffraction edge information; the diffraction edge information corresponds to the diffraction edge; wherein, when the dynamic geometry in the scene model changes, the corresponding diffraction edge information is updated; Accordingly, the collision information and reporting path corresponding to the light emitted from the listener's location are determined, including: Ray tracing is performed on the light emitted from the listener's location to determine collision information, and the reporting path is determined based on the diffraction edge information. The method according to claim 18, characterized in that, The reporting path includes the reflection path and the diffraction path; Ray tracing is performed on the light emitted from the listener's location to determine collision information, and the reporting path is determined based on the diffraction edge information, including: For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met: When the light ray collides with the scene model, the collision information for this collision is determined; When it is determined that the tracking of the ray will not end, the reflection path is determined and the direction of the reflected ray is calculated to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray. Furthermore, when it is determined that the tracking of the light ray will not end, it is determined whether diffraction will occur based on the diffraction edge information. When diffraction occurs, the diffraction path is determined, and the direction of the diffracted light ray is calculated to determine whether it will collide with another reflective surface in the scene model based on the direction of the diffracted light ray. The method according to claim 19, characterized in that, Determining whether diffraction will occur based on the diffraction edge information includes: Determine the point of collision between the light ray and the reflecting surface; When the collision point is close to the edge of the triangle corresponding to the reflecting surface, it is determined whether the edge is the saved diffraction edge; the collision point is determined to be close to the edge of the triangle corresponding to the reflecting surface based on the transformed centroid coordinates of the collision point. If the edge is one of a plurality of diffraction edges that are preserved, then two shadow regions are determined according to the two reflecting surfaces corresponding to the diffraction edge. When the light is in either of the shadow regions, diffraction is determined to occur. The method according to any one of claims 18-20, characterized in that, Constructing a reflection tree based on the reported path includes: The reported path is deduplicated to obtain a deduplicated path; the reported path includes a start node and an end node; the start node or the end node is a reflecting surface or a diffraction edge; The listener is identified as the root node of the reflection tree, and the following steps are repeated until the maximum depth of the reflection tree reaches a preset depth: For each node in the reflection tree, find the target node that the node can connect to from all the deduplicated paths, and determine the target node as the child node of the node. The method according to claim 21, characterized in that, The method further includes: When determining the reflection tree, if the node is the root node, the spatial position corresponding to the root node is determined as the listener's position; When the node is a reflective surface, determine the mirror position of the spatial position of the parent node of the node relative to the reflective surface, and determine the mirror position as the spatial position of the node. When the node is a diffraction edge, the center position of the diffraction edge is determined as the diffraction point or the spatial position corresponding to the node. The method according to claim 22, characterized in that, The reflection path in the second propagation path is a specular reflection path; multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree, including: For each sound source, determine the reflecting surface and diffraction edge where the light emitted from the sound source location collides with the scene model for the first time, so as to obtain a list of visible nodes of the sound source. For each leaf node in the reflection tree, when the leaf node is in the sound source visible list, it is determined whether the propagation path corresponding to the leaf node is a valid path; When the leaf node is not in the visible list of sound sources, the propagation path corresponding to the leaf node is determined to be an invalid path; The second propagation path is determined based on the established legal path; the second propagation path is a path obtained by sequentially connecting the sound source location, the reflection points or diffraction points corresponding to each node in the legal path, and the listener's location. The method according to claim 23, characterized in that, Determining whether the propagation path corresponding to the leaf node is a valid path includes: If every child node from the leaf node to the root node satisfies the target condition, then the path from the leaf node to the root node is determined to be a valid path; otherwise, it is an invalid path. When the node is a reflective surface, the target condition is: the reflection point corresponding to the reflective surface is within the triangle corresponding to the reflective surface and the corresponding path is not obstructed; the position of the reflection point is related to the spatial position corresponding to the reflective surface and the spatial position of the next level node. When the node is a diffraction edge, the target conditions are: the corresponding path is not occluded, and the preceding and following nodes of the node are located in two different shadow areas corresponding to the node. The method according to claim 24, characterized in that, Determining the echo map corresponding to the second propagation path includes: For the sound source, determine the initial energy corresponding to the sound source; For each second propagation path, the distance attenuation coefficient and air absorption coefficient are determined based on the total length of the second propagation path. The reflection attenuation coefficient of each reflecting surface in the second propagation path is determined. The reflection attenuation coefficients of each reflecting surface are multiplied to obtain the overall reflection attenuation coefficient. The diffraction coefficient is determined based on the diffraction edge information of each diffraction edge in the second propagation path. The reflection attenuation coefficient of the reflecting surface is related to the reflectivity and scattering rate of the reflecting surface. The multiplication result of the initial energy, the distance attenuation coefficient, the air absorption coefficient, the overall reflection attenuation coefficient, and the diffraction coefficient is determined as the energy received by the listener; The direction in which the last node points to the listener's position is defined as the receiving direction; The energy received by the listener, the receiving direction, and the propagation time corresponding to the second propagation path are determined as the echo map corresponding to the second propagation path. An acoustic ray tracing device, characterized in that, include: The listener tracking module is configured to determine the collision information corresponding to the light emitted from the listener's location; The collision information is the information of each collision point obtained by the light ray colliding with the scene model multiple times; The sound source tracking module is configured to, for each sound source, determine each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path; The filter synthesis module is configured to determine a filter corresponding to the sound source based on the echo map; The processing module is configured to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal. An electronic device, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1 to 25. A wearable device, characterized in that, It includes a processing unit; the processing unit is used to perform the method as described in any one of claims 1 to 25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 25.