Acoustic ray tracing method and device, electronic equipment, wearable equipment and storage medium

By dividing the ray tracing path into two stages—listener tracking and sound source tracking—the problem of high computational load in existing acoustic ray tracing algorithms is solved, thus optimizing computing resources and reducing system power consumption, and expanding its application on mobile devices.

CN122002207APending Publication Date: 2026-05-08GRAVITYXR ELECTRONICS & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GRAVITYXR ELECTRONICS & TECH CO LTD
Filing Date
2024-11-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing acoustic ray tracing algorithms are computationally intensive, resulting in excessive consumption of computing resources in virtual reality and mixed reality applications, which limits their application on mobile devices.

Method used

A two-stage tracking strategy is adopted, which divides the ray tracing path into a path from the listener to the last collision point and a path from the collision point to the sound source, and performs listener tracking and sound source tracking separately to reduce redundant calculations.

Benefits of technology

It effectively reduces the computational load of ray tracing algorithms, especially in multi-source scenarios, thereby reducing system power consumption and expanding application scenarios on mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002207A_ABST
    Figure CN122002207A_ABST
Patent Text Reader

Abstract

According to the acoustic light tracking method and device, the electronic equipment, the wearable equipment and the storage medium provided by the invention, the collision information corresponding to the light emitted by the position of the listener is determined, and each propagation path from the position of the sound source to the position of the listener is determined according to the collision information and the position of the sound source for each sound source; determining an echo graph corresponding to the sound source according to each propagation path, determining a filter corresponding to the sound source according to the echo graph, and processing the non-spatialized audio stream of the corresponding sound source according to the filter corresponding to each sound source to determine a reverberation signal, the tracking path is divided into the path from the listener to the last collision point and the path from the last collision point to the sound sources during reverse tracking, and the two paths are modeled respectively, so that when the position of the listener is not changed, the step of determining the collision information does not need to be executed, and when the number of the sound sources is multiple, the step of determining the collision information does not need to be executed. The step of determining the collision information is only calculated once, so that the calculation amount is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and in particular to an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium. Background Technology

[0002] Spatial audio technology is a user-centric technique that processes sound that lacks spatial characteristics to make it sound like it possesses specific spatial features, thus making the content heard by the user more realistic. In virtual reality or mixed reality applications, spatial audio technology can match the spatial characteristics of sound with visual content, resulting in a better sense of immersion.

[0003] In virtual reality or mixed reality applications, real-time acoustic modeling allows spatial audio to match the acoustic characteristics of a predefined space (virtual space or the user's real-world space). This enables the content the user hears to complement the visual content, providing a more immersive experience. For example, when the sound source changes, the content the listener hears also changes. Typically, real-time acoustic modeling uses acoustic ray tracing to model the acoustic path from the sound source to the listener based on the current spatial geometry, acoustic parameters, sound source, and listener position. The unspatialized audio stream is then processed based on the modeling results to obtain an audio stream that matches the expected acoustic characteristics.

[0004] Ray tracing requires tracking every ray emitted from the sound source location, and each ray needs to be reflected multiple times within the room. Therefore, the computational load of ray tracing algorithms is relatively large, and how to reduce the computational load of ray tracing algorithms is an urgent technical problem to be solved. Summary of the Invention

[0005] This invention provides an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium to effectively reduce the computational load of ray tracing algorithms.

[0006] In a first aspect, the present invention provides an acoustic ray tracing method, comprising:

[0007] Determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times;

[0008] For each sound source, each propagation path from the sound source location to the listener location is determined based on the collision information and the sound source location, and the echo map corresponding to the sound source is determined based on each propagation path.

[0009] Determine the filter corresponding to the sound source based on the echo map;

[0010] The unspatialized audio stream of each sound source is processed according to the filter corresponding to each sound source to determine the reverberation signal.

[0011] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0012] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:

[0013] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.

[0014] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0015] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;

[0016] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.

[0017] Optionally, the method further includes:

[0018] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;

[0019] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.

[0020] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.

[0021] Optionally, the rate of change of the reflective surface includes a new addition rate and a deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes:

[0022] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.

[0023] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:

[0024] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.

[0025] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.

[0026] Optionally, the method further includes:

[0027] The external reset condition is determined to be triggered when at least one of the following conditions is met:

[0028] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.

[0029] Optionally, when the listener's position does not change significantly, determining the collision information of the target number of light rays includes:

[0030] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.

[0031] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.

[0032] Optionally, selecting the target number of light rays from the N light rays emitted from the listener's location includes:

[0033] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.

[0034] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.

[0035] Optionally, determining the collision information of the target number of light rays includes:

[0036] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:

[0037] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;

[0038] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.

[0039] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.

[0040] Optionally, determining the direction of the reflected light based on the scattering rate of the reflecting surface includes:

[0041] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;

[0042] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.

[0043] Optionally, the collision information includes: the location of the collision point, the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including:

[0044] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.

[0045] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.

[0046] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.

[0047] Optionally, the collision information further includes: the specular reflection angle of the light beam, the cumulative propagation distance; and determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including:

[0048] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.

[0049] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.

[0050] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.

[0051] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.

[0052] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes:

[0053] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.

[0054] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;

[0055] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.

[0056] Optionally, determining the filter corresponding to the sound source based on the echo map includes:

[0057] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;

[0058] Determine the second filter corresponding to the sound source based on the echo map;

[0059] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.

[0060] Optionally, the smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including:

[0061] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.

[0062] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;

[0063] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.

[0064] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0065] The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.

[0066] Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including:

[0067] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path.

[0068] For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path.

[0069] In a second aspect, the present invention provides an acoustic ray tracing device, comprising:

[0070] The listener tracking module is used to determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times;

[0071] The sound source tracking module is used to determine, for each sound source, various propagation paths from the sound source location to the listener location based on the collision information and the sound source location, and to determine the echo map corresponding to the sound source based on each propagation path.

[0072] A filter synthesis module is used to determine the filter corresponding to the sound source based on the echo map;

[0073] The processing module is used to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.

[0074] Thirdly, the present invention provides an electronic device, comprising: at least one processor and a memory;

[0075] The memory stores computer-executed instructions;

[0076] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any of the first aspects.

[0077] Fourthly, the present invention provides a wearable device including a processing unit; the processing unit is configured to perform the method as described in any of the first aspects.

[0078] Fifthly, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any one of the first aspects.

[0079] In a sixth aspect, the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the method described in any of the first aspects.

[0080] This invention provides an acoustic ray tracing method, apparatus, electronic device, wearable device, and storage medium. By determining the collision information corresponding to the ray emitted from the listener's position—the collision information being the information of each collision point obtained from multiple collisions between the ray and the scene model—for each sound source, based on the collision information and the sound source position, various propagation paths from the sound source position to the listener's position are determined. Furthermore, based on each propagation path, an echo map corresponding to the sound source is determined. Based on the echo map, a filter corresponding to the sound source is determined. Based on the filter corresponding to each sound source, the unspatialized audio stream of the corresponding sound source is processed to determine the reverberation signal. By dividing the tracing path into a path from the listener to the last collision point and a path from the last collision point to the sound source during reverse tracing, and modeling these two paths separately, the step of determining collision information is eliminated when the listener's position remains unchanged. Moreover, when there are multiple sound sources, the step of determining collision information is calculated only once, thereby effectively reducing the computational load. Attached Figure Description

[0081] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0082] Figure 1 An application scenario diagram provided by an embodiment of the present invention;

[0083] Figure 2 A schematic flowchart of an acoustic ray tracing method provided in an embodiment of the present invention;

[0084] Figure 3 This is a schematic diagram of an overall architecture for generating spatial audio provided in an embodiment of the present invention;

[0085] Figure 4 A detailed schematic diagram of an architecture for generating spatial audio is provided for an embodiment of the present invention;

[0086] Figure 5 This is a schematic diagram of the structure of a listener tracking module 111 provided in an embodiment of the present invention;

[0087] Figure 6 A schematic diagram illustrating significant and non-significant changes in the listener's position, provided as an embodiment of the present invention;

[0088] Figure 7 A schematic diagram of the light propagation process provided in an embodiment of the present invention;

[0089] Figure 8 This is a schematic diagram illustrating the specific process of a listener tracking module 206 provided in an embodiment of the present invention;

[0090] Figure 9 A schematic diagram illustrating the calculation of reflection direction provided in an embodiment of the present invention;

[0091] Figure 10 This is a schematic flowchart of a sound source tracking module 112 provided in an embodiment of the present invention;

[0092] Figure 11 This is a schematic diagram of a specular reflection path provided in an embodiment of the present invention;

[0093] Figure 12 A schematic diagram of a scattering path provided in an embodiment of the present invention;

[0094] Figure 13 A schematic diagram of an echo map provided in an embodiment of the present invention;

[0095] Figure 14 This is a schematic diagram of the structure of a filter synthesis module 113 provided in an embodiment of the present invention;

[0096] Figure 15 An example diagram for reconstructing a frequency band impulse response from an echo map, provided as an embodiment of the present invention;

[0097] Figure 16 This is a schematic diagram of the structure of an acoustic ray tracing device provided in an embodiment of the present invention;

[0098] Figure 17 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention.

[0099] The accompanying drawings have illustrated specific embodiments of the invention, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the invention in any way, but rather to illustrate the concept of the invention to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0100] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention.

[0101] In virtual reality or mixed reality applications, spatial audio technology can be used to match the spatial characteristics of sound with visual content to achieve a better sense of immersion. Figure 1 An application scenario diagram provided by an embodiment of the present invention, such as... Figure 1As shown, when there are two sound sources in a space, the listener can simultaneously receive the spatial audio of these two sound sources in the virtual or real space. In other words, the spatial audio ultimately received by the listener is related to the location of the sound source, the listener's location, and the geometry of the room, thus enabling the user to combine the spatial audio they hear with the visual content to provide a stronger sense of realism.

[0102] A common method for determining spatial audio is to use acoustic ray tracing to model the acoustic path from the sound source to the listener based on the current spatial geometry, acoustic parameters (such as the reflectivity of reflective surfaces in a room), the location of the sound source, and the location of the listener. Here, modeling refers to determining filters with propagation characteristics from the sound source to the listener based on the energy received by the listener, so as to process the unspatialized audio stream according to the filters to obtain spatial audio that matches the expected acoustic characteristics.

[0103] In ray tracing, energy-carrying rays are randomly emitted from the sound source location, and the propagation of each ray in the room is tracked. When a ray collides with a wall in the room, its energy is attenuated, and it is then reflected and continues to propagate. This process is repeated until the condition for ending ray tracing is met, at which point ray tracing stops. For each collision point, the energy scattered to the listener can be calculated. An impulse response is generated based on the energy received by the listener, and this impulse response is then used to process the unspatialized audio stream.

[0104] The ray tracing process described above requires tracking a large number of rays, and each ray reflects multiple times within the room, resulting in a massive computational burden. To improve computational speed, one approach is to sort the rays emitted from the sound source, group them based on coherence, and assign each group of coherent rays to different CPU (Central Processing Unit) cores for parallel processing. Due to the characteristics of the ray intersection algorithm, processing coherent rays is faster than processing completely random rays, thus achieving a faster tracing speed.

[0105] However, the above method still has the following problems: only light rays emitted from the sound source can generate coherent rays, and the light rays after reflection of a set of coherent rays are no longer coherent. Therefore, this method can only utilize the faster computational property of coherent rays during the first ray collision, and cannot utilize this property for the computation of subsequent reflected rays. Furthermore, using multiple CPU cores for parallel processing only reduces the time required to run the ray tracing algorithm, without reducing power consumption, because the total computational load is not effectively reduced. In addition, as the number of sound sources in the scene increases, the same algorithm needs to be run for each sound source, and the overall time consumption will increase linearly. Therefore, the above ray tracing algorithm still suffers from a high computational load, and further reductions in the computational load of the ray tracing process are needed to expand its application scenarios on mobile devices.

[0106] Existing ray tracing methods involve ray light emanating from a sound source, colliding multiple times, and reaching the listener, forming a complete path. All rays from the sound source to the listener are then traced. Addressing these issues, this application proposes a two-stage tracing strategy. This strategy modifies the ray tracing path to a path from the listener to the sound source, splitting the complete path into two parts for separate tracing. Specifically, it splits the path into two stages: listener tracking and sound source tracking, thus decoupling the listener from the sound source. Therefore, a complete path is divided into two parts: the first part is the path from the listener to the last collision point (the last collision point where ray tracing from the listener ends), which is independent of the sound source location; the second part is the path from each collision point to the sound source, where each sound source is modeled separately. The advantages of the above method are as follows: Since the path in the first part is independent of the sound source location, for scenarios with multiple sound sources, the tracking process in the first part only needs to be run once, and the computational cost of the tracking process in the second part is linearly related to the number of sound sources. Furthermore, the sound source or the listener may move. When the sound source moves but the listener remains stationary, the tracking result in the first part will not change, thus the previous listener tracking result can be reused. The listener tracking part does not need to be run; only the sound source tracking part needs to be run. In summary, in scenarios with multiple sound sources and only sound source movement, the computational cost can be effectively reduced, thereby reducing system power consumption.

[0107] Figure 2 This is a flowchart illustrating an acoustic ray tracing method provided in an embodiment of the present invention. The method includes steps S201 to S204:

[0108] Step S201: Determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times.

[0109] This application describes a reverse ray tracing process, where the ray is traced originating from the listener's location. By reversing the tracing process, ray tracing can be divided into two stages, thus decoupling the listener from the sound source. It should be noted that the ray emanating from the listener's location here does not refer to real light rays, but rather to simulated light rays, such as multiple light rays emanating from the listener's location represented by equations. The ray tracing here is used to simulate the sound propagation path.

[0110] Users and listeners are situated within a scene, which can be represented using a scene model. Optionally, the scene can be a room, in which case the scene model can be used to indicate the room's dimensions, the positions of objects within the room, and / or the reflective surfaces present in the room and the reflective surfaces of each object, etc.

[0111] Optionally, multiple random rays can be emitted from the listener's position. These rays will collide with the scene model during propagation, and will randomly reflect and scatter. The collision information can be determined using the first-stage tracking process. This collision information can be obtained by tracking each ray emitted from the listener's position and recording the collision point information when it collides with the scene model. For example, the collision information can include the location of the collision point, the reflecting surface it occupies, and other information.

[0112] Optionally, the collision information may also include the order of the collision points corresponding to the light rays, or the propagation path from the listener's position to each collision point, in order to determine the propagation paths from the sound source position to the listener's position.

[0113] For example, three rays are emitted from the listener's position. For each ray, the corresponding collision information can be determined. For one of the rays, it may collide with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model in sequence. After the condition for ending ray tracing is met, the ray tracing ends. Here, collision point 3 is the last collision point corresponding to the ray, so the information of these three collision points can be determined.

[0114] Tracing the light rays emitted from the listener's position to the last collision point reveals the path of the light rays propagating from the listener's position to the last collision point. This process is independent of the sound source's location; that is, regardless of the number of sound sources or their locations, the collision information of each light ray remains unchanged. Therefore, the path of the light rays from the listener's position to the last collision point can be traced, and this tracing process can be considered as listener tracing.

[0115] Step S202: For each sound source, determine each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determine the echo map corresponding to the sound source based on each propagation path.

[0116] Once the collision information is determined, ray tracing can continue based on the collision information. Ray tracing here is to trace the path from each collision point to the sound source location. This path is related to the sound source location. When the sound source location is different, the path is different. This tracing process can be regarded as sound source tracing.

[0117] For a given sound source, based on the locations of each collision point, the paths from each collision point to the sound source can be determined. Then, based on the paths from the listener's location to each collision point, multiple complete propagation paths from the listener to the sound source can be obtained, and it can be equivalently assumed that there are identical propagation paths from the sound source to the listener's location. Therefore, based on the above steps, each propagation path from the sound source to the listener's location can be determined, and thus the echo map from the sound source to the listener's location can be determined, i.e., the echo map corresponding to the sound source. The echo map represents the arrival time, direction of arrival, and energy received by the listener for each propagation path from the sound source to the listener's location.

[0118] For example, if a ray of light collides sequentially with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model, and there are no other obstructions between each collision point and the sound source, then there are three propagation paths from the sound source position to the listener position. Path 1 is the ray of light from the sound source position through collision point 1 to the listener position; path 2 is the ray of light from the sound source position through collision point 2 and collision point 1 in sequence to the listener position; and path 3 is the ray of light from the sound source position through collision point 3, collision point 2, and collision point 1 in sequence to the listener position.

[0119] Step S203: Determine the filter corresponding to the sound source based on the echo map.

[0120] Once the echo map for each sound source is determined, each source can be processed individually to generate a filter for that source based on the source tracing results (i.e., the echo map). Optionally, the filter can be in the form of an Ambisonics filter bank.

[0121] Step S204: Process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.

[0122] After generating the filter corresponding to each sound source, a fast convolution method can be used to convolve the unspatialized audio stream of the corresponding sound source based on the filter to obtain the Ambisonics reverberation signal.

[0123] Figure 3 This is a schematic diagram of an overall architecture for generating spatial audio provided in an embodiment of the present invention, such as... Figure 3 As shown, the room geometry is considered the scene model in this application. The method of this application can be applied to a processing device, which includes an acoustic ray tracing module and a spatial audio rendering module. The room acoustic parameters, room geometry (also known as the scene model), sound source positions, and listener positions can be input into the processing device. The acoustic ray tracing module can process the input information to obtain filters corresponding to each sound source, and input the obtained filters corresponding to each sound source, the unspatialized audio stream in the memory, and the listener orientation information provided by the inertial sensor into the spatial audio rendering module. This allows the spatial audio rendering module to output a spatialized audio stream and output it to a sound playback device, such as a speaker or headphones. Steps S201 to S203 of this application are the execution content of the acoustic ray tracing module, and step S204 is the execution content of the spatial audio rendering module.

[0124] Figure 4 This is a detailed architectural diagram illustrating a method for generating spatial audio according to an embodiment of the present invention. The implementation process of steps S201 to S204 can be divided into two stages: a simulation stage 110 and a processing stage 120. In the simulation stage 110, based on a room scene model (3D Mesh), the reflectivity of each reflective surface in the scene model, and the positions of the listener and sound sources, the acoustic characteristics of the room are modeled using ray tracing. The modeling result is a filter corresponding to each sound source. In the processing stage 120, the estimated filters are used to process the unspatialized audio to obtain a reverberation signal. This room reverberation signal is added to the direct sound signal processed by HRTF (Head Related Transfer Functions) and then output to an audio playback device, such as a speaker or headphones.

[0125] Optionally, the simulation and processing stages can have different update frequencies depending on the actual scenario. The processing stage is typically designed to generate a real-time audio stream. For example, when the sampling rate is 48000Hz and 1024 samples are processed each time, the processing stage needs to run at least once every 21.3ms (1024 / 48000*1000) to generate the real-time audio stream. The update frequency of the simulation stage is set according to the complexity of the actual scenario and the system's computing power. For example, it can be set to update once every 100ms. This application does not limit the update frequency of the simulation and processing stages. Depending on different system architectures, the simulation and processing stages can also run on different processing devices, and this invention does not limit this.

[0126] Optionally, the simulation phase mainly includes: a listener tracking module 111, a sound source tracking module 112, and a filter synthesis module 113. The listener tracking part is only related to the listener's position; therefore, the listener tracking module 111 only needs to run when the listener's position changes. When the listener's position remains unchanged, this step can be skipped, and the output of the previous listener tracking module 111 can be reused to directly run the sound source tracking module 112. The sound source tracking module 112 runs separately for each sound source, depending on both the sound source position and the output of the listener tracking module 111. If the position and number of sound sources do not change, and the listener tracking module 111 is not running, then the sound source tracking module 112 can also be skipped. Similarly, the filter synthesis module 113 runs separately for each sound source, synthesizing the output of the corresponding sound source tracking module 112 into an Ambisonic filter bank for that sound source.

[0127] The processing stage mainly includes steps such as a fast convolution module 121, an HRTF processing module 122, an Ambisonics rotation module 123, and an Ambisonics decoding module 124. The fast convolution module 121 uses a fast convolution method to convolve the Ambisonics filter bank with the input unspatialized audio signal to obtain an Ambisonics reverberation signal. The Ambisonics rotation module 123 spatially rotates the Ambisonics reverberation signal based on the user's orientation information provided by an inertial sensor to match the user's current orientation. The Ambisonics decoding module 124 performs binaural decoding on the rotated Ambisonics reverberation signal to obtain a binaural reverberation signal. This signal is mixed with the binaural direct sound signal output by the HRTF processing module 122 to obtain the final binaural signal used for playback.

[0128] This invention provides an acoustic ray tracing method. By determining the collision information corresponding to the ray emitted from the listener's position—the collision information being the information of each collision point obtained from multiple collisions between the ray and the scene model—for each sound source, based on the collision information and the sound source position, various propagation paths from the sound source position to the listener's position are determined. Furthermore, based on each propagation path, an echo map corresponding to the sound source is determined. Based on the echo map, a filter corresponding to the sound source is determined. The unspatialized audio stream of each sound source is processed using the corresponding filter to determine the reverberation signal. By dividing the tracing path into a path from the listener to the last collision point and a path from the last collision point to the sound source during reverse tracing, and modeling these two paths separately, the method eliminates the need to determine collision information when the listener's position remains unchanged, and when there are multiple sound sources, the collision information determination step is only calculated once, thereby effectively reducing the computational load.

[0129] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0130] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:

[0131] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.

[0132] The collision information corresponding to the light emitted from the listener's position is determined by the listener tracking module 111, but this module is not executed every time an update time is reached. Optionally, the execution of this module can be determined based on whether the listener's position has changed, whether the scene model has changed, or whether the reflectivity of any reflective surface in the scene model has changed.

[0133] When the listener's position changes, the collision information when the light ray collides with the scene model will change. Even if the listener's position remains the same, the collision information will still change when the scene model changes. When the reflectivity of any reflective surface in the scene model changes, the remaining energy of the light ray after a collision can be altered, thus determining whether to continue tracking the ray; therefore, the collision information will also change.

[0134] When the update time of the listener tracking module 111 is reached and any one of the above three conditions is met, the listener tracking module 111 is run. Conversely, when none of the above three conditions are met, the listener tracking module 111 is not run, and the output of the previous run of the listener tracking module 111 can be reused directly.

[0135] By determining whether the listener's position, scene model, or the reflectivity of any reflective surface has changed, it is possible to accurately determine whether the listener tracking module is running, and in some scenarios, the listener tracking module can be disabled.

[0136] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0137] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;

[0138] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.

[0139] When determining the collision information corresponding to the light emitted from the listener's location, in some scenarios, only a portion of the light rays emitted from the listener's location can be tracked, instead of all light rays, in order to reduce computational load and power consumption.

[0140] When the listener's position changes, it is necessary to determine the collision information corresponding to the light emitted from the listener's position. In reality, the sound source and the listener are mostly located in the same space (such as the same room), and the listener's movements are mostly small, continuous movements, with a low percentage of significant movements. Significant movement can be understood as a large movement, or moving from one room to another. Insignificant movement can be understood as a small movement, or the acoustic environment in which the listener is located does not change significantly (e.g., the listener does not move from one space to another). When the listener's position changes insignificantly, a small number of light rays can be tracked and combined with the output of the previous listener tracking module 111 to obtain the current output.

[0141] Specifically, when the listener's position changes in a non-significant way, some light rays can be selected for tracking; when the listener's position changes significantly, all light rays can be selected for tracking.

[0142] For example, when a non-significant movement of the listener is detected, the listener tracking module 111 can select only a portion of the light rays for tracking during operation.

[0143] The above methods can reduce computational load and power consumption in scenarios where the listener moves but not significantly.

[0144] Optionally, the method further includes:

[0145] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;

[0146] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.

[0147] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.

[0148] Figure 5 This is a schematic diagram of the structure of a listener tracking module 111 provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the listener tracking module 111 is split into a listener tracking module (first reflection) 201 and a listener tracking module (subsequent multiple reflections) 206, and uses an adaptive adjustment strategy for the number of rays to reduce computation.

[0149] The core of adaptively adjusting the number of rays lies in determining whether the listener's position has changed significantly (including significant changes in the scene in which the user and sound source are currently located, or changes in the reflectivity of the reflecting surface). If the listener's position has not changed significantly, then the simulation can use fewer rays and make incremental updates based on the previous simulation results. However, if the listener's position has changed significantly, then the previous simulation results need to be discarded, and a full simulation should be performed directly using the pre-set maximum number of rays (or all rays emitted from the listener's position).

[0150] Figure 6 This is a schematic diagram illustrating significant and insignificant changes in the listener's position, provided as an embodiment of the present invention. Two rooms are connected by a door. Small, continuous movements of the listener within one room, such as moving from listener position 1 to listener position 2, are considered insignificant changes. However, moving from one room to another, such as moving from listener position 2 to listener position 3, is considered a significant change.

[0151] When determining whether the listener's position has changed significantly, the first collision information of the light rays emitted from the listener's position can be used. When N light rays are emitted from the listener's position, the first collision information includes the reflecting surfaces of the N light rays that collide for the first time. The listener's position can be determined based on the first collision information corresponding to the current simulation and historical simulations.

[0152] Historical simulations can be a single historical simulation or multiple historical simulations. Whether the listener's position has changed significantly can be determined based on the first collision information of the current simulation and the first collision information of multiple historical simulations. Specifically, if the listener's position has not changed significantly in multiple historical simulations, then the historical simulations are the simulation in which the listener's position changed significantly most recently, plus at least one simulation from each subsequent simulation. For example, if the current simulation is the 10th one, and the listener's position changed significantly in the 5th simulation, but not significantly in the 6th to 9th simulations (actually, the listener's position did not change significantly relative to the 5th simulation), then the historical simulations corresponding to the 10th simulation are one or more simulations from the 5th to 9th simulations.

[0153] Optionally, in the listener tracking module (first reflection) 201, N rays are emitted from the listener's position, each carrying initial energy E0. Their directions are randomly generated to ensure that they cover as uniformly as possible throughout the space. By performing ray intersection calculations with the scene, the reflecting surfaces S0, S1, ... S that these N rays collide with can be obtained. N-1 .

[0154] The reflective surface statistics module 202 counts all reflective surfaces S0, S1, ... SN-1 Perform deduplication statistics and maintain a list to represent all unique reflective surfaces. Optionally, this list can be implemented using a hash table (HashMap) for fast updates and lookups. This list is stored and used for the next simulation run, and is also input into the reflective surface change rate module 203.

[0155] The reflective surface change rate module 203 calculates the reflective surface change rate by comparing the reflective surface recorded in this simulation with the reflective surface recorded in historical simulations (which can be all non-repeating reflective surfaces that appear in multiple historical simulations) and outputs it to the comprehensive decision module 204.

[0156] In addition to determining whether the listener's position has changed significantly based on the collision information at the time of the first collision, it can also be determined based on whether an external reset condition is triggered. The external reset condition can be input into the comprehensive decision module 204.

[0157] The comprehensive decision module 204 can determine whether a significant change has occurred based on whether the rate of change of the reflective surface is greater than a preset value and whether an external reset condition has been triggered.

[0158] Judging whether the listener's position has moved significantly by using the first collision information of light emitted from the listener's position can be effective in scenarios with complex models, such as rooms with many objects that can reflect light. Even if the listener moves less than a preset distance compared to historical simulations, the objects around the listener may have changed significantly. This method can effectively identify such changes and improve the accuracy of judging whether the listener's position has moved significantly.

[0159] Optionally, the rate of change of the reflective surface includes a new addition rate and a deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes:

[0160] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.

[0161] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:

[0162] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.

[0163] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.

[0164] Optionally, the reflective surface change rate module 203 adjusts the value based on the number of newly added reflective surfaces n. new Delete the number of reflective surfaces n remove The total number of reflective surfaces in the historical simulation, n previous (Number of reflective surfaces after deduplication), and the total number of reflective surfaces n in this simulation. current (Number of reflective surfaces after deduplication), and calculate the rate of change of reflective surfaces, including the increase rate update_ratio. new and deletion rate update_ratio remove :

[0165]

[0166]

[0167] Optionally, the addition rate is the ratio of the number of newly added reflective surfaces to the total number of reflective surfaces in this simulation; the number of newly added reflective surfaces is the number of reflective surfaces added in this simulation compared to historical simulations; the deletion rate is the ratio of the number of deleted reflective surfaces to the total number of reflective surfaces in historical simulations; the number of deleted reflective surfaces is the number of reflective surfaces deleted in this simulation compared to historical simulations.

[0168] A first threshold can be set for the new addition rate and a second threshold can be set for the deletion rate. The calculated new addition rate is compared with the first threshold and the calculated deletion rate is compared with the second threshold. When the new addition rate is less than the first threshold and the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.

[0169] Conversely, a significant change in the listener's location is determined when any one of the following conditions is met: the new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, or the external reset condition is triggered.

[0170] By comparing the addition rate, deletion rate, and threshold, it is possible to determine whether the listener's position has changed significantly based on the first collision information of the light emitted from the listener's position.

[0171] Optionally, the method further includes:

[0172] The external reset condition is determined to be triggered when at least one of the following conditions is met:

[0173] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.

[0174] When the scene model changes or the reflectivity of any reflective surface in the scene model changes, the collision information corresponding to the light emitted from the listener's position will change. If the distance the listener's position moves in this simulation exceeds a preset distance compared to previous simulations, it indicates a significant change in the listener's position. Furthermore, if no significant change is triggered in multiple consecutive simulations, an external reset condition can be triggered to avoid the continuous accumulation of errors.

[0175] By determining whether the external reset condition has been triggered by the above conditions, the accuracy of judging whether the listener's position has changed significantly can be improved.

[0176] Optionally, when the listener's position does not change significantly, determining the collision information of the target number of light rays includes:

[0177] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.

[0178] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.

[0179] When the listener's position changes insignificantly, the number of targets is less than N. The specific value of the number of targets can be related to the rate of change of the reflective surface. For example, a rate of change of the reflective surface less than 0.5 is considered a non-significant change. When the rate of change of the reflective surface is 0.1 and 0.4, the number of targets will be different. When the rate of change of the reflective surface is large, it can be intuitively understood that the user's movement distance is large, and therefore the number of targets is large. In this simulation, more light rays are selected for tracking.

[0180] Optionally, the rate of change of the reflective surface includes the addition rate and the deletion rate. The larger the addition rate, the larger the number of targets; or, the larger the deletion rate, the larger the number of targets.

[0181] Optionally, curves relating the new addition rate to the target quantity and the deletion rate to the target quantity can be set to determine the target quantity 1 corresponding to the new addition rate and the target quantity 2 corresponding to the deletion rate, and then the final target quantity can be determined based on the target quantity 1 and the target quantity 2.

[0182] For example, when the rate of change of the reflective surface is 0.1, the number of targets is one-quarter of all light rays; when the rate of change of the reflective surface is 0.2, the number of targets is one-third of all light rays; when the rate of change of the reflective surface is 0.3, the number of targets is half of all light rays, and so on. Conversely, if the number of targets is always one-quarter of all light rays for different rates of change of the reflective surface, the accuracy of the calculation results will be low; if the number of targets is always half of all light rays for different rates of change of the reflective surface, the computational load will be too large.

[0183] By determining the target quantity based on the rate of change of the reflective surface, different target quantities can be set based on different reflectivities, which can reduce the amount of calculation while ensuring the accuracy of the calculation results.

[0184] Optionally, selecting the target number of light rays from the N light rays emitted from the listener's location includes:

[0185] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.

[0186] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.

[0187] When selecting the target number of rays, if there is a newly added reflective surface of the first collision, the rays that pass through the newly added reflective surface of the first collision can be selected first. If the total number of rays that pass through the newly added reflective surface of the first collision does not reach the target number, then multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.

[0188] like Figure 5 As shown, the light selection module 205 can determine the number of targets based on the decision result. If there is no significant change, then the light selection module (subsequent multiple reflections) 206 is executed, selecting the target number of light rays from the listener tracking module (first reflection) 201. If there is a significant change, then all light rays are selected and the listener tracking module (subsequent multiple reflections) 206 is executed.

[0189] By prioritizing the light rays from the newly added first collision reflector, the effectiveness of ray tracing can be achieved, allowing the final collision information to reflect the listener's movement, thereby improving the accuracy of the determined collision information.

[0190] Optionally, determining the collision information of the target number of light rays includes:

[0191] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:

[0192] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;

[0193] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.

[0194] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.

[0195] Figure 7 This is a schematic diagram of a light propagation process provided in an embodiment of the present invention. Multiple light rays can be randomly emitted from the listener's position, and each light ray is tracked to obtain collision information with the scene model.

[0196] Figure 8 This is a schematic flowchart of a listener tracking module 206 provided in an embodiment of the present invention. When light hits a reflective surface, reflection and scattering occur. The reflectivity α of each reflective surface is used as a reference. s The remaining energy α after reflection can be calculated. s E, where E represents the energy currently carried by the ray. The distance between the current collision point and the previous collision point is calculated and accumulated to obtain the ray's propagation distance (the distance from the listener's position to the current collision point). The number of collisions, the ray's propagation distance, and the remaining ray energy determine whether to end tracking of the current ray. If the conditions for ending tracking are not met, the direction of the reflected ray is calculated based on the scattering rate of the reflective surface, and the next collision point with the scene model is calculated. The above steps are repeated for all rays until the conditions for ending tracking are met for all rays.

[0197] When the number of collisions reaches the upper limit, the propagation distance of the light reaches the upper limit, or the remaining light energy reaches the lower limit, the conditions for ending the tracking of the current light are determined to be met.

[0198] The collision information for each collision can include: the reflecting surface where the collision point is located, the specific collision location, the angle of the incident light and the angle of the reflected light, the cumulative distance, and other information.

[0199] By tracing rays, collision information can be determined, and by determining whether to end ray tracing, ray tracing can be terminated in a timely manner when conditions are met, thereby reducing the amount of computation.

[0200] Optionally, determining the direction of the reflected light based on the scattering rate of the reflecting surface includes:

[0201] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;

[0202] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.

[0203] Figure 9 A schematic diagram for calculating the reflection direction is provided in an embodiment of the present invention, as shown below. Figure 9 As shown, the direction of the reflected light is obtained by combining the specular reflection direction and the random scattering direction. If the incident direction is d... incident If the normal vector of the collision surface is n, then the direction of mirror reflection is:

[0204] d specular =d incident -2*d incident ·n

[0205] Optional, random scattering direction d scattering Let n be a randomly generated direction vector that follows a Lambert distribution relative to the normal vector n.

[0206] According to the scattering rate α of the collision surface scattering The direction of the reflected light is obtained and normalized, as shown in the following formula:

[0207] α reflection =α scattering d scattering +(1-α scattering )d specular

[0208]

[0209] The scattering rate is related to the material of the reflecting surface. The reflection direction obtained by the above method simulates the difference in dispersion caused by the scattering rate of different acoustic materials.

[0210] Optionally, the collision information includes: the location of the collision point, the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including:

[0211] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.

[0212] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.

[0213] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.

[0214] Multiple sound source tracking modules 112 are included, which can process all collision information generated by the listener tracking module 111. Based on the sound source location, scene model, and reflectivity information of the corresponding sound source, they generate the propagation path and echo map from the sound source to the listener. Since the tracked light rays originate from the listener, the actual calculated path is the propagation path from the listener to the sound source, and it is equivalent to assuming that there is a common propagation path from the sound source location to the listener location.

[0215] Figure 10 This is a schematic flowchart of a sound source tracking module 112 provided in an embodiment of the present invention. For the light ray corresponding to the collision point, it can be determined whether there is an obstruction between the collision point and the sound source location. If there is no obstruction, the path between the collision point and the sound source location can be determined. If there is an obstruction, the next collision point is processed. When the path between the collision point and the sound source location is determined, the propagation path between the corresponding sound source location and the listener location can be obtained, and thus multiple paths between the sound source location and the listener location can be obtained. For the propagation path between the sound source location and the listener location, an echo map can be calculated, including: the energy received by the sound source, the arrival time and direction of the light ray reaching the sound source. The echo map can be determined based on the reflectivity and scattering rate of the reflecting surface where the collision point is located.

[0216] Optionally, the collision information further includes: the specular reflection angle of the light beam, the cumulative propagation distance; and determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including:

[0217] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.

[0218] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.

[0219] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.

[0220] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.

[0221] Optionally, based on the reflectivity α of the reflecting surface where the collision point is located. s The remaining energy α after reflection is obtained. s E, where E is the energy currently carried by the ray. α s E is further divided into two parts: mirror energy and scattering energy, where the scattering energy is α. scattering α s E, α scattering It is the scattering rate of the reflecting surface at the current collision point, and the remaining part (1-α) scattering )α s E is the mirror energy.

[0222] Figure 11 This is a schematic diagram of a specular reflection path provided in an embodiment of the present invention. Figure 12 A schematic diagram of a scattering path provided in an embodiment of the present invention, as shown below. Figure 11 and Figure 12 As shown, when there are no other obstructions between the collision point and the sound source, there exists a propagation path between them, and its direction originates from the location corresponding to the collision point. The specular reflection angle of this light ray (which has also been calculated in the listener tracking module 111) is used to determine if it can pass through the sound source. If so, the path is considered a specular reflection path, and the specular energy of the light ray (1-α) is... scattering )α s If all of E is emitted to the sound source, then the path is considered a scattering path, and a portion of the scattered energy of the light is emitted to the sound source.

[0223] Optionally, assuming a scattering ratio of α, the energy emitted to the sound source is αα. scattering α s The value of E.α can be given by the following formula:

[0224]

[0225] Where θ is the angle between the scattered ray (the ray pointing from the collision point to the sound source) and the normal vector of the reflecting surface (collision surface), d is the distance from the collision point to the center of the sound source, and r is the radius of the sound source.

[0226] Furthermore, the energy received by the sound source is also related to the air absorption coefficient. The total propagation distance of the light beam is obtained by adding the cumulative propagation distance of the light beam to the distance between the collision point and the sound source. Based on this total propagation distance and the exponential decay model, the air absorption coefficient α of the current path is calculated. air The energy α received by the sound source is obtained. air (1-α scattering )α s E (Mirror Path) or α air αα scattering α s E (scattering path).

[0227] By saving the energy received by the sound source, the total time it takes for light to travel from the listener to the sound source, and the direction from which the light is emitted from the listener into an echograph, Figure 13 This is a schematic diagram of an echo map provided in an embodiment of the present invention, such as... Figure 13 As shown, after tracking all light propagation paths, the sound source tracking module 112 statistically analyzes all propagation paths from the sound source to the listener and saves them as an echoogram. Each point in the echoogram represents the energy E carried by a possible propagation path i when it reaches the listener. i Propagation time t i And the horizontal and vertical angles θ of the direction of arrival. i ,

[0228] It should be noted that the tracked light rays originate from the listener's position; therefore, the direction of arrival here is the opposite of the emission direction of the light rays originating from the listener's position. Since the reflective surface may have different reflectivities for different frequency signals, the sound source's received energy can also be divided into multiple frequency bands.

[0229] By considering the type of path between the collision point and the sound source location, as well as the air absorption coefficient, the accuracy of the calculated sound source received energy can be improved.

[0230] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes:

[0231] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.

[0232] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;

[0233] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.

[0234] After determining the energy received by the sound source, the light energy can be updated by subtracting the energy received by the sound source from the remaining light energy after reflection. The updated light energy can be used to calculate the next collision point and can be used as the energy carried by the light corresponding to the next collision point.

[0235] Furthermore, it can be determined whether the updated ray energy has reached the lower limit, and whether to continue calculating subsequent collision points for that ray. When the updated ray energy reaches the lower limit, the calculation of subsequent collision points for that collision point is stopped.

[0236] By updating the light energy based on the energy received from the sound source, the accuracy of the calculation results can be improved to avoid calculating every collision point output by the listener tracking module 111.

[0237] Optionally, determining the filter corresponding to the sound source based on the echo map includes:

[0238] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;

[0239] Determine the second filter corresponding to the sound source based on the echo map;

[0240] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.

[0241] Once the echo map is determined, the second filter corresponding to the sound source can be identified, which is the filter obtained in this simulation.

[0242] Figure 14This is a schematic diagram of a filter synthesis module 113 provided in an embodiment of the present invention. The filter synthesis module 113 synthesizes the echo map generated by the sound source tracking module 112 into an Ambisonic impulse response. The filter synthesis module 113 includes an Ambisonics encoding module 210, a filter generation module 211, a sub-band synthesis module 212, and a filter smoothing module 213. The Ambisonics encoding module 210 first performs Ambisonic encoding on the echo map. Based on the arrival direction of each propagation path, the corresponding Ambisonics coefficients can be obtained. The Ambisonics intensity is obtained by multiplying it by the corresponding sound source received energy. in, It is a spherical harmonic basis function, P l |m| It is a combined Legendre function.

[0243]

[0244]

[0245]

[0246] Figure 15 This is an example diagram illustrating the reconstruction of a frequency band impulse response from an echo map, provided by an embodiment of the present invention. The filter generation module 211 processes each Ambisonics channel in the Ambisonics echo map separately. For each channel, based on the propagation time and intensity of each path in different frequency bands, an impulse response for the corresponding frequency band is generated. Specifically, firstly, the propagation time and corresponding Ambisonics intensity of each point in the echo map are read. Then, the sampling point position corresponding to the propagation time is found in the impulse response of the corresponding Ambisonics channel and frequency band. Subsequently, a pulse with a peak value equal to the corresponding intensity is inserted at this sampling point, such as... Figure 15 As shown. Depending on actual needs, when the sampling point position corresponding to the propagation time is not an integer, interpolation can also be performed on the corresponding intensity within a certain period before and after that position, such as using Lagrange interpolation. This invention does not limit this.

[0247] The subband synthesis module 212 can reconstruct the full-band impulse response from the subband impulse response using an FIR filter bank or a Linkwitz-Riley filter bank based on a Biquad IIR filter, and the present invention does not limit this.

[0248] The filter smoothing module 213 smooths the generated filter based on the number of targets (current number of rays). When the number of targets is not equal to N, the randomness of the ray tracing results will increase, which will cause the generated filter to have a certain degree of difference between multiple simulations, thus causing jitter in the final spatialized audio signal.

[0249] To address the aforementioned issues, for each sound source, when determining the corresponding final filter, a smoothing coefficient can be set. The smoothing coefficient α represents the proportion of the previous filter (first filter) retained in the current output, and 1-α represents the proportion of the currently generated filter (second filter) adopted. The final filter is generated by weighted summation of the filters generated in the previous simulation and the current simulation based on the smoothing coefficient.

[0250] The above methods can smooth out the jitter of spatialized audio signals. On the other hand, since this simulation only uses a portion of the rays, the information contained in the ray tracing results may not be comprehensive enough. By retaining a portion of the ray tracing results from the previous simulation, the comprehensiveness of the information contained in the final ray tracing results can be improved.

[0251] Optionally, the smoothing coefficient can be related to the number of targets. When the number of targets is smaller, the smoothing coefficient is larger, indicating that the proportion of the filter obtained from the previous simulation is adopted. Conversely, when the number of targets is larger, the smoothing coefficient is smaller, indicating that the proportion of the filter obtained from the previous simulation is adopted.

[0252] Optionally, when the number of targets equals the maximum number of rays N, α is set to 0, indicating no smoothing is performed to achieve the fastest response speed to changes in the listener's position. When the number of targets is small, a larger smoothing coefficient, such as 0.7, is used to generate a smooth filter.

[0253] By setting a smoothing coefficient related to the target number, and by weighting and summing the filters generated in the previous simulation and the current simulation based on the correlation coefficient, a final filter is generated to smooth out this jitter.

[0254] Optionally, the smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including:

[0255] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.

[0256] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;

[0257] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.

[0258] The compensation coefficient can be determined based on the smoothing coefficient, and is used to compensate for the amplitude attenuation caused by adding two phase-mismatched filters. The smoothing coefficient is a value between 0 and 1, and the compensation coefficient is a value greater than 1.

[0259] The filter smoothing module 213 smooths the filter h generated in the previous simulation. t-1 (First filter) and the filter h generated this time t The second filter is weighted and summed to generate the final filter, as shown in the following equation:

[0260]

[0261] Optionally, the compensation coefficient can be expressed as

[0262] By first calculating the product of the smoothing coefficient and the first filter, and then the product of the difference between the value 1 and the smoothing coefficient and the second filter, and finally adding the two products together, and then multiplying the sum with the compensation coefficient, the amplitude attenuation caused by the addition of two phase-mismatched filters can be compensated.

[0263] By calculating the compensation coefficient based on the smoothing coefficient, and then calculating the final filter based on the compensation coefficient, the accuracy of the final filter can be improved, and the amplitude attenuation problem caused by the sum of mismatched filters can be eliminated.

[0264] In reality, ray tracing treats sound propagation as equivalent to particle motion, neglecting the diffraction effect of sound waves. Diffraction refers to the phenomenon where sound waves can bypass obstacles and continue propagating. In the real world, when there is an obstruction between a sound source and a listener, sound wave diffraction is a crucial path for sound propagation between them, allowing us to "hear the sound before we see it." Without modeling diffraction, the spatial audio simulated using ray tracing algorithms will lack realism. Therefore, in some scenarios, sound waves may need to travel through a combination of diffraction and reflection to reach the listener, but the methods described above cannot support modeling such combined paths.

[0265] Optionally, determine the collision information corresponding to the light emitted from the listener's location, including:

[0266] The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.

[0267] Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including:

[0268] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path.

[0269] For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path.

[0270] The ray tracing in this application is inverse ray tracing, which means tracing the light rays emitted from the listener's position. By performing the inverse tracing process, ray tracing can be divided into two stages, thereby decoupling the listener from the sound source. It should be noted that the light rays emitted from the listener's position here do not refer to real light rays, but can be simulated light rays, such as multiple light rays emitted from the listener's position represented by equations. The ray tracing here is used to simulate the propagation of sound.

[0271] The scene in which the user and listener are situated can be represented by a scene model. Optionally, the scene can be a room, in which case the scene model can be used to indicate the size of the room, the position of each object in the room, and / or the reflective surfaces contained in the room and the reflective surfaces of each object, etc.

[0272] Optionally, multiple random light rays can be emitted from the listener's position. These rays will collide with the scene model during propagation, randomly reflecting and scattering. The collision information can be determined using the first-stage tracking process (i.e., listener tracking). This collision information can involve tracking each light ray emitted from the listener's position and recording the collision point when it collides with the scene model. For example, the collision information can include the location of the collision point and the reflective surface it occupies.

[0273] For example, three rays are emitted from the listener's position. For each ray, the corresponding collision information can be determined. For one of the rays, it may collide with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model in sequence. After the condition for ending ray tracing is met, the ray tracing ends. Here, collision point 3 is the last collision point corresponding to the ray, so the information of these three collision points can be determined.

[0274] Tracing the light rays emitted from the listener's position to the last collision point reveals the propagation path from the listener's position to the last collision point, independent of the sound source's location. In other words, regardless of the number or location of the sound sources, the collision information of each light ray remains unchanged. Therefore, the path of the light rays from the listener's position to the last collision point can be traced; this tracing process can be considered listener tracking.

[0275] During the listener tracking process, when each ray is tracked, if the ray collides with the reflecting surface, it is determined whether there is a diffraction edge based on the collision result. The reporting path is then determined based on the reflecting surface or diffraction edge corresponding to the collision point. A reflection tree is then constructed based on the reporting path. The path from the leaf node to the root node in the reflection tree is a possible propagation path of the sound.

[0276] After determining the collision information, ray tracing can continue based on the collision information. Ray tracing here is to trace the path from each collision point to the sound source location. This path is the first propagation path, which is related to the sound source location. When the sound source location is different, the first propagation path is different. This tracing process can be regarded as sound source tracing.

[0277] For a given sound source, based on the location of each collision point, the path from each collision point to the sound source can be determined. Then, based on the path from the listener's location to each collision point, multiple complete propagation paths from the listener to the sound source can be obtained, and it can be equivalently assumed that there is a common propagation path from the sound source location to the listener's location.

[0278] For example, when a ray of light collides sequentially with reflective surface 1 (creating collision point 1), reflective surface 3 (creating collision point 2), and reflective surface 2 (creating collision point 3) in the scene model, there are three propagation paths from the sound source position to the listener position. Path 1 is the ray of light from the sound source position through collision point 3 to the listener position; path 2 is the ray of light from the sound source position through collision point 3 and collision point 2 to the listener position; and path 3 is the ray of light from the sound source position through collision point 3, collision point 2, and collision point 1 to the listener position.

[0279] The paths determined above based on the locations of various collision points and sound sources are mostly scattering paths, meaning that the specular reflection of the incident light rays fails to pass through the sound source. Therefore, a ray of light scattered to the sound source is calculated. However, the specular reflection path is a real-world path, and the energy of the specular reflection light rays is relatively high, significantly impacting the listener's perception. Therefore, based on the reflection tree and the sound source location, a combined path of specular reflection and diffraction can be determined from the sound source to the listener, which is the second propagation path.

[0280] For each sound source, the echo map corresponding to the first propagation path and the echo map corresponding to the second propagation path can be determined. The echo map represents the arrival time, direction of arrival, and energy received by the listener for each propagation path from the sound source location to the listener location. For example, if there are three propagation paths from the sound source to the listener, the echo map corresponding to each path can be determined.

[0281] Based on the echo maps corresponding to each sound source, the filter corresponding to that sound source can be determined. Optionally, the filter can be in the form of an Ambisonics filter bank. After generating the filter corresponding to each sound source, a fast convolution method can be used to convolve the unspatialized audio stream of the corresponding sound source based on the filter to obtain the Ambisonics reverberation signal.

[0282] This invention proposes an acoustic ray tracing method. First, random diffraction is incorporated into the inverse ray tracing method, and the results of random ray tracing are statistically analyzed. A reflection tree is constructed based on all traced reflection and diffraction nodes. All legal combinations of reflection and diffraction paths are found based on the reflection tree. Finally, all found paths are modeled to obtain a filter. This approach has the following advantages: it can model combinations of reflection and diffraction paths; it does not require prior calculation and analysis of diffraction paths, thus enabling real-time modeling of scenes containing dynamic geometry; it does not require exhaustively enumerating all diffraction edges and paths in the scene, but instead utilizes the ray tracing results to construct and traverse only the traced portions of the reflection tree, reducing computational and storage requirements for complex scenes; the diffraction modeling results are merged with the filters generated by random ray tracing, allowing for unified processing of unspatialized audio streams during the processing stage.

[0283] like Figure 3As shown, the room geometry is considered the scene model in this application. The method of this application can be applied to a processing device, which includes an acoustic ray tracing module and a spatial audio rendering module. Room acoustic parameters, room geometry (also known as the scene model), sound source locations, and listener locations can be input into the processing device. The acoustic ray tracing module processes the input information to obtain filters corresponding to each sound source (including sound diffraction, scattering, and reflection effects, etc.). The obtained filters corresponding to each sound source, the unspatialized audio stream in memory, and the listener orientation information provided by the inertial sensor are input into the spatial audio rendering module, enabling the spatial audio rendering module to output a spatialized audio stream and output it to a sound playback device, such as a speaker or headphones. The above steps and the process of determining filters based on echo maps constitute the execution content of the acoustic ray tracing module, while determining the reverberation signal based on the filters constitutes the execution content of the spatial audio rendering module.

[0284] The constructed reflection tree is a listener-related reflection tree, which has the advantages of low computational cost and high accuracy. This enables the modeling of the combined path of reflection and diffraction in the sound propagation process to obtain an accurate reverberation signal.

[0285] Figure 16 This is a schematic diagram of an acoustic ray tracing device provided in an embodiment of the present invention. The device 160 includes:

[0286] The listener tracking module 1601 is used to determine the collision information corresponding to the light emitted from the listener's position; the collision information is the information of each collision point obtained by the light colliding with the scene model multiple times;

[0287] The sound source tracking module 1602 is used to determine, for each sound source, various propagation paths from the sound source location to the listener location based on the collision information and the sound source location, and to determine the echo map corresponding to the sound source based on each propagation path.

[0288] The filter synthesis module 1603 is used to determine the filter corresponding to the sound source based on the echo map;

[0289] The processing module 1604 is used to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.

[0290] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically used for:

[0291] Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met:

[0292] The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.

[0293] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically used for:

[0294] Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N;

[0295] Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.

[0296] Optionally, the device further includes: a judging module, used for:

[0297] Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing;

[0298] The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations.

[0299] The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.

[0300] Optionally, the rate of change of the reflective surface includes the addition rate and the deletion rate; when determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions, the judgment module is specifically used for:

[0301] When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly.

[0302] And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met:

[0303] The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered.

[0304] The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.

[0305] Optionally, the device further includes: a trigger module, used for:

[0306] The external reset condition is determined to be triggered when at least one of the following conditions is met:

[0307] The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.

[0308] Optionally, when the listener's position does not change significantly, the listener tracking module 1601, when determining the collision information of the target number of light rays, is specifically used for:

[0309] The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface.

[0310] Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.

[0311] Optionally, when selecting the target number of light rays from the N light rays emitted from the listener's location, the listener tracking module 1601 is specifically used for:

[0312] When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained.

[0313] When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.

[0314] Optionally, when determining the collision information of the target number of light rays, the listener tracking module 1601 is specifically used for:

[0315] For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met:

[0316] When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined;

[0317] The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions.

[0318] When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.

[0319] Optionally, when determining the direction of the reflected light based on the scattering rate of the reflective surface, the listener tracking module 1601 is specifically used for:

[0320] The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface;

[0321] Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.

[0322] Optionally, the collision information includes: the location of the collision point and the reflecting surface where the collision point is located; when the sound source tracking module 1602 determines each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determines the echo map corresponding to the sound source based on each propagation path, it is specifically used for:

[0323] For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener.

[0324] The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point.

[0325] For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.

[0326] Optionally, the collision information further includes: the specular reflection angle of the light and the cumulative propagation distance; when the sound source tracking module 1602 determines the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, it is specifically used for:

[0327] The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light.

[0328] The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location.

[0329] When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light.

[0330] When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.

[0331] Optionally, after determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the sound source tracking module 1602 is further used for:

[0332] The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source.

[0333] When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map;

[0334] When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.

[0335] Optionally, when determining the filter corresponding to the sound source based on the echo map, the filter synthesis module 1603 is specifically used for:

[0336] The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation;

[0337] Determine the second filter corresponding to the sound source based on the echo map;

[0338] For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.

[0339] Optionally, the smoothing coefficient is a value between 0 and 1; when the filter synthesis module 1603 determines the final filter generated in this simulation based on the smoothing coefficient, the first filter, and the second filter, it is specifically used for:

[0340] The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient.

[0341] Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter;

[0342] The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.

[0343] Optionally, when determining the collision information corresponding to the light emitted from the listener's location, the listener tracking module 1601 is specifically used for:

[0344] Determine the collision information and reporting path corresponding to the light emitted from the listener's location;

[0345] The device further includes: a reflection tree construction module, used for:

[0346] A reflection tree is constructed based on the reported path; the reported path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point.

[0347] The sound source tracking module 1602, when determining various propagation paths from the sound source location to the listener location based on the collision information and the sound source location for each sound source, and determining the echo map corresponding to the sound source based on each propagation path, is specifically used for:

[0348] For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and the echo map corresponding to the first propagation path is determined.

[0349] The device further includes: a path search module, used for:

[0350] Based on the reflection tree, multiple second propagation paths are determined from the sound source location to the listener location; the second propagation path is a combination of a reflection path and a diffraction path.

[0351] The device further includes: a path modeling module, used to: determine the echo map corresponding to the second propagation path;

[0352] The acoustic ray tracing device 160 provided in this embodiment of the invention can achieve the above-mentioned... Figure 2 The acoustic ray tracing method shown in the embodiment has a similar implementation principle and technical effect, and will not be described again here.

[0353] Figure 17This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 17 As shown, the electronic device provided in this embodiment includes at least one processor 1701 and a memory 1702. The processor 1701 and the memory 1702 are connected via a bus 1703.

[0354] In a specific implementation, at least one processor 1701 executes computer execution instructions stored in memory 1702, causing at least one processor 1701 to execute the method in the above method embodiment.

[0355] The specific implementation process of processor 1701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0356] Optionally, the electronic device can be a head-mounted display device.

[0357] In the above Figure 17 In the illustrated embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0358] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.

[0359] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0360] This invention also provides a wearable device, including a processing unit; the processing unit is used to implement the method described in the above method embodiments.

[0361] This invention also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in the above embodiments.

[0362] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the above method embodiments.

[0363] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0364] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0365] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0366] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0367] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0368] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An acoustic ray tracing method, characterized in that, include: Determine the collision information corresponding to the light emitted from the listener's location; The collision information is the information of each collision point obtained by the light ray colliding with the scene model multiple times; For each sound source, each propagation path from the sound source location to the listener location is determined based on the collision information and the sound source location, and the echo map corresponding to the sound source is determined based on each propagation path. Determine the filter corresponding to the sound source based on the echo map; The unspatialized audio stream of each sound source is processed according to the filter corresponding to each sound source to determine the reverberation signal.

2. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: Collision information corresponding to the light emitted to determine the listener's location is executed when at least one of the following conditions is met: The listener's position changes, the scene model changes, and the reflectivity of any reflective surface in the scene model changes.

3. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: Determine the collision information of the target number of light rays; the number of light rays emitted from the listener's position is N; Wherein, when the listener's position changes in a non-significant manner, the number of targets is less than N; and / or, when the listener's position changes significantly, the number of targets is equal to N.

4. The method according to claim 3, characterized in that, The method further includes: Determine the first collision information of N light rays emitted from the listener's position; the first collision information includes the reflecting surface where the N light rays collide for the first time; the reflecting surface where the first collision occurs is a reflecting surface after deduplication processing; The rate of change of the reflective surface is determined based on the reflective surface of the first collision in this simulation and the reflective surface of the first collision in historical simulations. The listener's position is determined to have undergone a non-significant change based on the rate of change of the reflective surface and the external reset conditions.

5. The method according to claim 4, characterized in that, The rate of change of the reflective surface includes the addition rate and the deletion rate; determining whether the listener's position has undergone a non-significant change based on the rate of change of the reflective surface and external reset conditions includes: When the new addition rate is less than the first threshold, the deletion rate is less than the second threshold, and the external reset condition is not triggered, it is determined that the listener's position has not changed significantly. And / or, the listener's position is determined to have changed significantly when at least one of the following conditions is met: The new addition rate is greater than or equal to the first threshold, the deletion rate is greater than or equal to the second threshold, and the external reset condition is triggered. The addition rate is related to the number of reflective surfaces added in the first collision compared to the current simulation and historical simulations; the deletion rate is related to the number of reflective surfaces deleted in the first collision compared to the current simulation and historical simulations.

6. The method according to claim 4, characterized in that, The method further includes: The external reset condition is determined to be triggered when at least one of the following conditions is met: The scene model changes or the reflectivity of any reflective surface in the scene model changes; the distance the listener's position moves exceeds a preset distance compared to previous simulations; or no significant changes are triggered in multiple consecutive simulations.

7. The method according to claim 5, characterized in that, When the listener's position undergoes a non-significant change, the collision information for determining the target number of light rays includes: The number of targets is determined based on the rate of change of the reflective surface; the number of targets is positively correlated with the rate of change of the reflective surface. Select the target number of light rays from the N light rays emitted from the listener's position, and determine the collision information of the target number of light rays.

8. The method according to claim 7, characterized in that, Selecting the target number of light rays from the N light rays emitted from the listener's location includes: When there is a new reflective surface for the first collision compared to the previous simulation, all rays that pass through the new reflective surface for the first collision out of the N rays will be retained. When the number of retained rays is less than the target number, multiple rays are selected from the remaining rays in the N rays to obtain the target number of rays.

9. The method according to claim 7, characterized in that, Determining the collision information of the target number of light rays includes: For each ray of light, repeat the following steps to determine the collision information each time the ray collides, until the condition for ending the tracking of the ray is met: When the light ray collides with a reflective surface in the scene model, the collision information for this collision is determined; The remaining light energy is determined based on the reflectivity of the reflective surface, and the distance between the current collision point and the previous collision point is determined to obtain the propagation distance of the light. The tracking of the light is then terminated based on the remaining light energy, the propagation distance of the light, and the number of collisions. When it is determined that the tracking of the ray will not end, the direction of the reflected ray is determined based on the scattering rate of the reflective surface, so as to determine whether a collision will occur with another reflective surface in the scene model based on the direction of the reflected ray.

10. The method according to claim 9, characterized in that, Determining the direction of the reflected light based on the scattering rate of the reflecting surface includes: The direction of mirror reflection is determined based on the incident direction of the light and the normal vector of the reflecting surface; Determine the random scattering direction, and then determine the direction of the reflected light based on the random scattering direction, the scattering rate of the reflecting surface, and the specular reflection direction.

11. The method according to any one of claims 1-10, characterized in that, The collision information includes: the location of the collision point and the reflecting surface where the collision point is located; determining each propagation path from the sound source location to the listener location based on the collision information and the sound source location, and determining the echo map corresponding to the sound source based on each propagation path, including: For each light ray corresponding to a collision point, when it is determined from the scene model that there is no obstruction between the position of the collision point and the position of the sound source, the path between the position of the collision point and the position of the sound source is determined, so as to determine each propagation path from the position of the sound source to the position of the listener. The echo map corresponding to the propagation path is determined based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located; the echo map includes: the energy received by the sound source, the arrival time and direction of the light rays reaching the sound source; the direction of arrival is the opposite direction of the light rays emitted from the listener's position corresponding to the collision point. For any sound source, the echo map corresponding to the sound source is determined based on the echo maps corresponding to each propagation path.

12. The method according to claim 11, characterized in that, The collision information also includes: the specular reflection angle of the light beam and the cumulative propagation distance; determining the sound source received energy corresponding to the propagation path based on the propagation path, the reflectivity and scattering rate of the reflecting surface where the collision point is located, including: The type of propagation path between the location of the collision point and the location of the sound source is determined based on the specular reflection angle of the light. The cumulative propagation distance and the distance between the collision point and the sound source location are added together to determine the total propagation distance of light from the sound source location to the listener location. The air absorption coefficient is then determined based on the total propagation distance. The air absorption coefficient is related to the total propagation path from the sound source location to the listener location. When the propagation path is a specular reflection path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, and the energy currently carried by the light. When the propagation path is a scattering path, the sound source receives energy based on the air absorption coefficient, the reflectivity and scattering rate of the reflecting surface where the collision point is located, the scattering ratio, and the energy currently carried by the light. The scattering ratio is related to the angle between the light from the collision point to the sound source and the normal vector of the reflecting surface, the distance from the location of the collision point to the center of the sound source, and the radius of the sound source.

13. The method according to claim 11, characterized in that, After determining the echo map corresponding to the propagation path based on the propagation path, the reflectivity and scattering of the reflecting surface where the collision point is located, the method further includes: The light energy is updated based on the remaining light energy after reflection and the energy received by the sound source. When the updated light energy does not reach the lower limit of light energy, if there is a next collision point after the collision point, determine the path between the position of the next collision point and the position of the sound source and calculate the echo map; When the updated light energy reaches the lower limit of light energy, the calculation of subsequent collision points for the aforementioned collision point is stopped.

14. The method according to claim 3, characterized in that, Determining the filter corresponding to the sound source based on the echo map includes: The smoothing coefficient is determined based on the target quantity; the smoothing coefficient represents the proportion of the first filter that is retained in this simulation; the first filter represents the filter of the sound source generated in the previous simulation; Determine the second filter corresponding to the sound source based on the echo map; For each sound source, the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter; the second filter represents the filter determined in this simulation based on tracing the target number of rays; the smoothing coefficient is negatively correlated with the target number.

15. The method according to claim 14, characterized in that, The smoothing coefficient is a value between 0 and 1; the final filter generated in this simulation is determined based on the smoothing coefficient, the first filter, and the second filter, including: The compensation parameter is determined based on the smoothing coefficient; the compensation parameter is a value greater than 1, and the compensation parameter is related to the square of the smoothing coefficient. Calculate the first product of the smoothing coefficient and the first filter, calculate the difference between the value 1 and the smoothing coefficient, and calculate the second product of the difference and the second filter; The sum of the first product and the second product is calculated, and the product of the compensation parameter and the sum is determined as the final filter generated in this simulation.

16. The method according to claim 1, characterized in that, Determine the collision information corresponding to the light emitted from the listener's location, including: The collision information and reporting path corresponding to the light emitted from the listener's location are determined, and a reflection tree is constructed based on the reporting path; the reporting path is a path composed of the reflecting surface or diffraction edge corresponding to the collision point. Accordingly, for each sound source, based on the collision information and the sound source location, various propagation paths from the sound source location to the listener location are determined, and based on each propagation path, the echo map corresponding to the sound source is determined, including: For each sound source, multiple first propagation paths from the sound source location to the listener location are determined based on the collision information and the sound source location, and multiple second propagation paths from the sound source location to the listener location are determined based on the reflection tree; the second propagation path is a combination of a reflection path and a diffraction path. For each sound source, determine the echo map corresponding to the first propagation path, and determine the echo map corresponding to the second propagation path.

17. An acoustic ray tracing device, characterized in that, include: The listener tracking module is used to determine the collision information corresponding to the light emitted from the listener's location; The collision information is the information of each collision point obtained by the light ray colliding with the scene model multiple times; The sound source tracking module is used to determine, for each sound source, various propagation paths from the sound source location to the listener location based on the collision information and the sound source location, and to determine the echo map corresponding to the sound source based on each propagation path. A filter synthesis module is used to determine the filter corresponding to the sound source based on the echo map; The processing module is used to process the unspatialized audio stream of each sound source according to the filter corresponding to each sound source to determine the reverberation signal.

18. An electronic device, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1 to 16.

19. A wearable device, characterized in that, It includes a processing unit; the processing unit is used to perform the method as described in any one of claims 1 to 16.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 16.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 16.