Method and system for generating acoustic pulse response of 3D room model using wave-based and geoacoustic

By hybridizing wave-based and geometric acoustic solvers, the shortcomings of existing acoustic simulation methods in speed and accuracy are addressed, and efficient and accurate acoustic simulations within different acoustic frequency ranges are achieved, adapting to the requirements of modern computer processing capabilities.

CN120660089APending Publication Date: 2025-09-16TREBLE TECHNOLOGIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380092654.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-17
Filing Date
2023-11-28
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing acoustic simulation methods cannot effectively combine speed and accuracy on computer processing systems, resulting in low fidelity during fast simulations or excessively long simulation times during high-fidelity simulations, which cannot meet the requirements of modern computer processing capabilities.

Method used

A hybrid solver is used, combining a wave-based solver and a geometric acoustic solver, to perform simulations in different acoustic frequency ranges. The wave-based solver is used for high-precision simulation, and the geometric acoustic solver is used for high-speed simulation. The final result is output by merging the two.

Benefits of technology

It achieves efficient and accurate acoustic simulation in different acoustic frequency ranges, improves simulation speed and accuracy, adapts to modern computer processing capabilities, and supports sound wave simulation in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660089A_ABST
    Figure CN120660089A_ABST
Patent Text Reader

Abstract

A computer-implemented method for generating an impulse response of a listening point in a room is disclosed, wherein the method comprises the steps of: receiving a 3D model of the room, a location of at least one sound source in the 3D model of the room, and an acoustic property of at least one boundary in the 3D model of the room; determining, using a wave-based solver, a wave-based impulse response of the wave-based propagation of pulses emitted at the at least one sound source in the 3D model of the room and received at the listening point within the first acoustic frequency range; determining, using a geometric acoustic-based solver, a geometric impulse response of a ray-based propagation of pulses emitted at the at least one sound source in the 3D model of the room and received at the listening point within a second acoustic frequency range; and generating an impulse response by combining the wave impulse response and the geometric impulse response. Methods and systems using the disclosed computer-implemented methods and features thereof are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A computer-implemented method for generating an impulse response of a listening point in a room, wherein: The method comprises the following steps: (a) receiving a 3D model of a room, a position of at least one sound source in the 3D model of the room, and an acoustic property of at least one boundary in the 3D model of the room; (b) determining, using a wave-based solver, a wave-based impulse response based on wave propagation of an impulse emitted at at least one sound source in the 3D model of the room and received at a listening point within a first acoustic frequency range; (c) using a geometric acoustics based solver to determine a ray-based propagation geometric impulse response of an impulse emitted at at least one sound source in the 3D model of the room and received at the listening point within a second acoustic frequency range; and (d) generating the impulse response by combining the wave impulse response and the geometric impulse response.

2. The computer-implemented method of claim 1 , wherein: The first acoustic frequency range and the second acoustic frequency range partially overlap.

3. The computer-implemented method of claim 1 or 2, wherein: The virtual domain comprises at least one directional sound source for emitting sound in a defined direction.

4. A computer-implemented method according to any one of the preceding claims, wherein: The computer-implemented method further includes performing a mesh model of the 3D model of the room, wherein the 3D mesh model is a 3D curved mesh model.

5. The computer-implemented method of claim 1 , 2 or 3 , wherein: The wave-based solver applies the discontinuous Galerkin finite element method (DGFEM) or the spectral element method (SEM).

6. The computer-implemented method of any one of claims 1 to 5, wherein: The method further comprises a calibration step, wherein the power level of the at least one sound source is adjusted so that the sound level received at a predetermined distance from the at least one sound source is the same in the wave-based solver and the geometric acoustic solver.

7. The computer-implemented method of any one of claims 1 to 6, wherein: The higher frequencies of the first acoustic frequency range and the lower frequencies of the second frequency range overlap at a transition frequency.

8. The computer-implemented method of claim 7, wherein: The step of merging the wave-based impulse response and the geometric impulse response comprises applying a low-pass filter to the wave-based impulse response.

9. The computer-implemented method of claim 7 or 8, wherein: The step of merging the wave-based impulse response and the geometric impulse response comprises applying a high-pass filter to the geometric impulse response.

10. The computer-implemented method of claim 8 or 9, wherein: The low-pass filter and / or the high-pass filter comprises a cut-off frequency at the transition frequency.

11. A computer-implemented method according to any one of the preceding claims, wherein: The wave-based solver and / or the geometric acoustic solver comprises extracting one or more wave-based impulse responses and / or one or more geometric impulse responses based on a simulation of sound propagation, and wherein the wave-based impulse responses are spatial impulse responses.

12. The computer-implemented method of claim 11, wherein: The spatial impulse response includes a plurality of single-channel impulse responses, wherein each of the plurality of single-channel impulse responses records a wave impulse response from a specific direction or angle at a same listening point.

13. A computer-implemented method according to any one of the preceding claims, wherein: A spherical receiver array is arranged around the listening point, wherein the spherical receiver array includes a plurality of receivers.

14. The computer-implemented method of claim 13, wherein: The spherical receiver array is an open spherical array of cardioid receivers.

15. The computer-implemented method of claim 13 or 14, wherein: The spherical receiver array comprises at least 2 receivers, preferably at least 4 receivers, more preferably at least 8 receivers, even more preferably at least 16 receivers, most preferably at least 32 receivers, even more preferably at least 64 receivers.

16. The computer-implemented method of any one of claims 13-15, wherein: The number of receivers is determined based on the maximum truncation order N so that the number of receivers is higher than or equal to (N+1) 2 .

17. A computer-implemented method according to any one of the preceding claims, wherein: The computer-implemented method further includes convolving the generated impulse response with a base audio signal such that a convolved audio signal is generated.

18. A method according to any one of the preceding claims, wherein The computer-implemented method also includes rendering the base audio signal by convolving the base audio signal with the generated impulse response, thereby producing a rendered audio signal.

19. The computer-implemented method of claim 18, wherein: The basic audio signal is speech, music, ambient sound, impulse sound or any combination thereof.

20. The computer-implemented method of any one of claims 17-19, wherein: The rendered audio signal provides an audio rendering of the base audio signal in a 3D model of the room at the listening point.

21. A system for generating an impulse response of a listening point in a room, the system comprising: A computer system having a processor coupled to a memory, the processor configured to: receiving a 3D model of a room, a location of at least one sound source in the 3D model of the room, and an acoustic property of at least one boundary in the 3D model of the room; determining, using a wave-based solver, a wave-based impulse response based on wave propagation of an impulse emitted at at least one sound source in the 3D model of the room and received at a listening point within a first acoustic frequency range; determining, using a geometric acoustics-based solver, a ray-based propagation geometric impulse response of an impulse emitted at at least one sound source in the 3D model of the room and received at a listening point within a second acoustic frequency range; as well as The impulse response is generated by combining a wave impulse response and the geometric impulse response.

22. A computer-implemented method for training a machine learning model for audio compensation, wherein: The method comprises the following steps: - receiving a plurality of 3D models of rooms, each of said 3D models comprising at least one sound source and at least one acoustic property, - receiving a plurality of impulse responses at a listening position in each of the plurality of rooms, - using at least the plurality of impulse responses as input to train the machine learning model for audio compensation.

23. The computer-implemented method of claim 22, wherein: The plurality of impulse responses are preprocessed to generate a plurality of modified impulse responses for training the machine learning model.

24. The computer-implemented method of claim 23, wherein: The pre-processing includes applying a filter to enhance the speech range in the impulse response.

25. The computer-implemented method of claim 24, wherein: The speech range is between 3kHz-17kHz or between 350Hz-17kHz.

26. The computer-implemented method of any one of claims 22-25, wherein: The method also includes providing a plurality of reverberant audio signals using the plurality of reverberant audio signals as input for training the machine learning model.

27. The computer-implemented method of claim 26, wherein: The reverberant audio signal is provided by convolving each of the plurality of impulse responses with a base audio signal.

28. The computer-implemented method of claim 26, wherein: The reverberant audio signal is provided by recording an audio signal received at at least one loudspeaker of the audio device.

29. The computer-implemented method of any one of claims 22 to 28, wherein: The method also includes using the 3D models of the plurality of rooms as input for training the machine learning model.

30. The computer-implemented method of any one of claims 22-29, wherein: The method also includes receiving a digital model of the audio device and using the digital model of the audio device as input for training the machine learning model.

31. The computer-implemented method of any one of claims 22-30, wherein: At least one preferred listening position in each of the plurality of rooms is used as input for training the machine learning model.

32. The computer-implemented method of any one of claims 22-31, wherein: Training the machine learning model comprises using a plurality of reverberant audio signals according to claim 26, 27 or 28 as input, wherein the training comprises reconstructing a base audio signal as output.

33. The computer-implemented method of claim 32, wherein: The 3D models of the plurality of rooms are used as input and at least one preferred listening point in each of the plurality of rooms is used as input, wherein the method further comprises reconstructing a base audio signal at the at least one preferred listening point as output.

34. The computer-implemented method of any one of claims 22-33, wherein: Training the machine learning model includes generating a compensated impulse response as an output.

35. The computer-implemented method of claim 34, wherein: Generating the compensating impulse response is based on the reverberant audio signal according to claim 26, 27 or 28 and a base audio signal as input for training the machine learning model.

36. The computer-implemented method of claim 34 or 35, wherein: The 3D models of the plurality of rooms are used as input and the at least one preferred listening point in each of the plurality of rooms is used as input, wherein the method further comprises generating a compensated impulse response at the at least one preferred listening point as output.

37. The computer-implemented method of any one of claims 22-36, wherein: The method further comprising receiving the plurality of impulse responses comprises generating the impulse responses for each 3D model of the plurality of rooms according to the method of claims 1-21.

38. The computer-implemented model of any one of claims 22-37, wherein: The method also includes receiving a digital model of the audio device.

39. The computer-implemented method of any one of claims 22-38, wherein: Training the machine learning model includes a neural network.

40. The computer-implemented method of claim 39, wherein: The neural network includes an autoencoder for encoding either one of the inputs to the model or for generating compressed inputs for training a machine learning model for audio compensation.

41. The computer-implemented method of claim 39 or 40, wherein: The neural network also includes training a Generic Adversarial Network (GAN) to generate any of the inputs for training a machine learning model for audio compensation.

42. The computer-implemented method of claim 39, 40, or 41, wherein: The neural network includes a deep neural network, a convolutional neural network, and / or a transformer for training a machine learning model for audio compensation.

43. A method for providing a machine learning model for audio compensation in an audio device, wherein: The audio device comprises at least one microphone, wherein the machine learning model is trained by a method according to any one of claims 22-42, and wherein the method comprises receiving an audio signal at the at least one microphone and generating a compensated audio signal by using the machine learning model for audio compensation.

44. A method for providing machine learning-based audio compensation in an audio device comprising at least one microphone, wherein: Machine learning based audio compensation has been trained using a computer-implemented method for training a machine learning model according to any one of claims 22-42, wherein the method for providing machine learning based audio compensation includes receiving an audio signal at at least one microphone, and generating a compensated audio signal by using the trained machine learning model on the received audio signal.

45. The method according to claim 43 or 44, wherein The compensated audio signal is compensated by convolving the received audio signal with the compensating impulse response of claim 34, 35 or 36.

46. ​​An audio device comprising at least one microphone, wherein: The audio device comprises a processing system for applying machine learning-based audio compensation according to the method of any one of claims 43-45, wherein the audio device is configured to receive an audio signal at the at least one microphone and apply the machine learning-based audio compensation to the audio signal to generate a compensated audio signal.

47. The audio device of claim 46, wherein The audio device includes a communication module configured to transmit a compensated audio signal.

48. An audio device according to claim 46 or 47, wherein The audio device is configured to transmit the compensated audio signal to the cloud, the internet, and / or a network storage center.

49. An audio device according to claim 46, 47 or 48, wherein The audio device is further configured to transmit the compensated audio signal to a remote audio device, wherein the remote audio device includes at least one remote speaker for outputting the compensated audio signal.

50. A system for training a machine learning model for audio compensation, the system comprising: A computer system having a processor coupled to a memory, the processor configured to: receiving a plurality of 3D models of rooms, each of the 3D models comprising at least one sound source and at least one acoustic property; receiving a plurality of impulse responses at a listening position in each of the plurality of rooms; as well as The machine learning model for audio compensation is trained using at least the plurality of impulse responses as input.

51. A computer-implemented method of determining a head-related transfer function, the method comprising: (a) receiving a 3D model of a user's head and a position within the 3D model of at least one sound source representing an eardrum or the like; (b) determining, using a wave-based solver, a plurality of wave-based impulse responses from impulses transmitted at the at least one sound source, wherein the plurality of wave-based impulse responses are determined at a plurality of digital representations of the head receivers; as well as (c) determining a head-related transfer function (HRTF) of the user's head based on the plurality of wave-based impulse responses determined at the plurality of digital representations of the head receivers.

52. The computer-implemented method of claim 51 , wherein: The wave-based solvers use either the Discontinuous Galerkin Finite Element Method (DGFEM) or the Spectral Element Method (SEM).

53. The computer-implemented method of any one of claims 51-52, wherein: The computer-implemented method also includes obtaining a head mesh model representing a geometry of a user's head.

54. The computer-implemented method of claim 53, wherein: The head mesh model is a curved head mesh model.

55. The computer-implemented method of any one of claims 51-54, wherein: The computer-implemented method further includes arranging a digital representation of a head array including a plurality of digital representations of head receivers around the head mesh model such that a distance between any digital representation of the head receivers and the head mesh model is no less than a predetermined distance.

56. The computer-implemented method of any one of claims 51-55, wherein: The computer-implemented method also includes determining a first closest mesh element on the head mesh model that is closest to an eardrum.

57. The computer-implemented method of any one of claims 51-56, wherein: The computer-implemented method also includes arranging a digital representation of a first source correction microphone at a first source distance from the first nearest grid element, wherein the first source distance is less than the predetermined distance.

58. The computer-implemented method of any one of claims 51-57, wherein: The computer-implemented method also includes digitally transmitting a first pulse signal using the first nearest lattice element as a sound source.

59. The computer-implemented method of any one of claims 51-58, wherein: The computer-implemented method also includes determining a first source correction signal using the wave-based solver, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone.

60. The computer-implemented method of any one of claims 51-59, wherein: The computer-implemented method also includes determining a plurality of first source-corrected head impulse responses by source correcting each of the plurality of wave-based impulse responses using the first source correction signal.

61. The computer-implemented method of any one of claims 51-60, wherein: The computer-implemented method also includes generating a head-related transfer function of the user's head for the first eardrum by combining the plurality of first source-corrected head impulse responses.