Generating device- and room-dependent impulse responses

By generating and combining the device's transfer function and the room's impulse response step by step, the problem of low efficiency in existing acoustic simulation tools when dealing with wave-based phenomena is solved, enabling rapid simulation of high-fidelity stereo and flexible evaluation of the device in different rooms.

CN122497954APending Publication Date: 2026-07-31TREBLE TECHNOLOGIES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TREBLE TECHNOLOGIES
Filing Date
2024-05-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing acoustic simulation tools are inefficient and time-consuming when dealing with wave-based phenomena, especially when considering the performance of audio devices in a room. Wave-based simulation methods rely on gridding, which leads to high computational complexity and makes it difficult to parallelize computing resources.

Method used

Device-dependent transfer function (DRTF) and spatial room impulse response (SRIR) are generated stepwise and then combined into device-specific room impulse response (DSRIR). The simulation is performed in a small space using a wave-based solver and combined with geometric acoustic simulation to improve efficiency.

Benefits of technology

It achieves fast simulation of high-fidelity stereo, allowing audio devices to be freely rotated and evaluated in different rooms, reducing simulation time while maintaining high fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497954A_ABST
    Figure CN122497954A_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein the method includes generating at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device received by at least a first microphone, generating a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room received from at least one room sound source in the room and from at least one direction at at least one listening point in the room, and generating the device-specific room impulse response (DSRIR) by combining the device-related transfer function and the spatial room impulse response.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Acoustic simulation refers to computer-based modeling and simulation of sound propagation and interaction in various environments. These simulations are valuable tools in fields such as engineering, architecture, and audio design, allowing researchers and professionals to predict, analyze, and optimize acoustic conditions in real-world scenarios. Acoustic simulations are commonly used to generate impulse responses and transfer functions for simulated objects, whether simulating real-world conditions and identifying the location of acoustic problems or designing new objects (such as buildings and rooms) to avoid acoustic issues in construction.

[0002] The most prevalent solvers used in today's acoustic simulation tools are ray-based or image-based, often referred to as geometry solvers. These are fast and processor-efficient, but because they are ray-based, they do not account for wave-based problems, which makes them very poor at detecting wave-based phenomena such as standing waves, wave cancellation, and similar issues. This is particularly problematic in low or mid-frequency audio.

[0003] However, with the increase in computing power available, for example through cloud computing, the use of so-called wave-based solvers has become feasible. Nevertheless, due to computational constraints, care must be taken when implementing this approach so that it delivers high-fidelity results as quickly as possible.

[0004] Audio device modeling and optimization are key aspects of audio engineering. Current methods require time-consuming simulations or measurements under real-life conditions in specific rooms equipped with several speakers to characterize how an audio device receives sound from multiple sound sources. Furthermore, when considering how the device behaves in a particular room or space, such simulations must take into account both the device geometry associated with its acoustic characteristics and the geometry of the space or room associated with those characteristics, resulting in even more complex and time-consuming simulations.

[0005] Wave-based simulations, such as the finite element method (FEM), finite-difference time-domain method (FDTD), boundary element method (BEM), or finite volume method (FVM), are used to simulate sound waves in high fidelity. One of the challenges of wave-based simulations mentioned above is the computational complexity and the inability to parallelize computational resources, which makes them time-consuming. Some recent improvements have been implemented to reduce simulation time, especially by using the discontinuous Galerkin method (DG), which allows for efficient parallelization of the simulation, thereby greatly reducing simulation time [1]. Wave-based simulations still rely on meshing, which includes the internal volume of the meshed 3D model and the geometry of the elements included in the 3D model, where sound waves need to be simulated by wave-based simulation. The size of the mesh elements or geometric features in the mesh determines the time step of the wave-based simulation. Therefore, if complex geometry is included in the internal volume of the 3D model, the simulation time can become very long, sometimes up to several days, if the simulation can converge in some way.

[0006] Audio devices can sometimes have complex geometries and be relatively small compared to the size of the room in which they are placed, and they can include several microphones. Determining how an analog audio device can capture sound waves from one or more audio sources in a 3D model of the room can be challenging. This becomes particularly complex if the audio device includes several microphones. Summary of the Invention

[0007] Therefore, there is a need to make wave-based simulations and generation of impulse responses and transfer functions more efficient, especially when considering objects such as audio devices in relation to their performance in a space or room (e.g., a conference room). As disclosed herein, this can be achieved by performing separate simulations, generating the impulse response and transfer function of the device in one step and the impulse response and transfer function of the room in another step, so that the results can then be combined to give a complete evaluation of the device and the room together, while allowing for the rotation of the audio device without the cost of additional simulations of the impulse response.

[0008] In one aspect, this disclosure relates to a computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein the device includes at least a first microphone.

[0009] The computer-implemented method further includes the following steps: - Generate at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device received by the at least first microphone. - Generate a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room as received from at least one room sound source in the room and from at least one direction at at least one listening point in the room, and - The device-specific room impulse response (DSRIR) is generated by combining the device-related transfer function and the room impulse response.

[0010] Generating DRTF and SRIR in two separate steps (e.g., through two separate simulations) can have the advantage of avoiding a large simulation, which can be particularly troublesome when using wave-based solvers.

[0011] Wave-based solvers use meshes to solve, and in such cases are limited by the minimum mesh size in the model. Since the apparatus will typically be represented by mesh elements much smaller than a room, running simulations with a meshed model in a room becomes very slow. Furthermore, it is often desirable to run several room simulations, each with different rooms or different listening points, to evaluate how the apparatus performs in such rooms. Therefore, providing a method where DRTF and SRIR are generated separately but can be combined into a DSRIR offers numerous options for generating new rooms, reorienting the apparatus, or using different apparatuses, at a speed greater than simulating everything in the same step while maintaining high fidelity.

[0012] In one embodiment of the above aspects, such advantages become apparent in a computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein the device includes at least a first microphone, the method comprising: - Generate at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device received by the at least first microphone, wherein generating the first device-related transfer function further includes A device mesh model is obtained, representing the geometry of the device and the position of at least the first microphone on the device mesh model. The digital representations of an array of device receivers, including multiple digital representations of the device receivers, are arranged around the device mesh model such that the distance between any digital representation of the device receiver and the device mesh model is not less than a predetermined distance. On the device mesh model, determine the first nearest mesh element that is closest to at least the first microphone. A digital representation of a first source correction microphone positioned at a first source distance from the first nearest grid element, wherein the first source distance is less than the predetermined distance. The first pulse signal is digitally emitted using the first closest grid element as the sound source. A wave-based solver is used to determine a first source correction signal, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone. A wave-based solver is used to determine a plurality of first device impulse responses, wherein each first device impulse response describes the impulse response of the first pulse signal received at a corresponding device receiver. The plurality of first-source-corrected device impulse responses are determined by performing source correction on each of the plurality of first device impulse responses using the first source correction signal. The first device-related transfer function for the first microphone is generated by combining the device impulse responses of the plurality of first source corrections. Determine the energy content at at least one frequency of the relevant transfer function of the first device. - Generate a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room received from at least one room sound source in the room and from at least one direction at at least one listening point in the room, wherein generating the spatial room impulse response further includes Obtain a 3D room model representing the geometry of the room and at least one acoustic property. At least one digital representation of at least one room sound source is arranged in the 3D room model. A digital representation of a room receiver array comprising multiple digital representations of a room receiver, wherein the room receiver array is centered on at least one listening point in the 3D room model, and wherein the number of digital representations of the room receivers is determined based on the energy content of at least one frequency of the first device-related transfer function. The room pulse signal is digitally transmitted from the at least one audio source. At least one wave-based frequency-based solver is used to determine multiple room impulse responses, wherein each room impulse response describes a transmitted room impulse received at a corresponding location in multiple digital representations of the room receiver. A spatial room impulse response is generated based on the multiple room impulse responses. - The device-specific room impulse response (DSRIR) is generated by combining the device-related transfer function and the room impulse response.

[0013] This embodiment is suitable, for example, for high-fidelity stereo because the device receiver array used to generate the first DRTF and the room receiver array SRIR can be sized to use spherical harmonics for encoding and decoding high-fidelity stereo. For example, the number of receivers in the device array is based on (N+1). 2 Determine the highest high-fidelity stereo order N. Similarly, the spatial impulse response of a room can be encoded into high-fidelity stereo using spherical harmonics, where the number of receivers in the room array is also based on (N+1). 2 Determine the highest high-fidelity stereo order N.

[0014] The number of receivers in the device receiver array or room receiver array determines the maximum cutoff order N that can be considered for spherical harmonic decomposition. The sound field can be sampled with at least as many receivers as are used for the extended terms, resulting in a receiver count greater than or equal to (N+1). 2 This condition may be necessary to avoid undersampling, but it does not guarantee accuracy. The optimal number of receivers can depend on the chosen orthogonality rule; exact integration can be performed, for example, using Gaussian orthogonality, where the number of receivers can be greater than or equal to (N+1). 2 Other methods may require fewer samples. Another constraint may involve spatial aliasing. Aliasing can be mitigated by choosing a sufficiently large order N or by using an anti-aliasing spatial filter.

[0015] Using spherical harmonics to encode and decode high-fidelity stereo is generally known, but as mentioned, it can be a computationally intensive process that has recently become more feasible within the context of increased available computing power. However, controlling and using the optimal high-fidelity stereo order N will allow the system to be more efficient and effective.

[0016] Because wave-based simulations used to derive the device-dependent transfer function (DRTF) are performed or conducted in a relatively small space (e.g., within a device receiver array), the high-fidelity stereo order N(device) can typically be higher. However, for room-based simulations used to determine the spatial room impulse response (SRIR), wave-based simulations are typically performed for a larger space. Therefore, determining the optimal high-fidelity stereo order N(room), for example, by using the energy content from the device simulation, would allow for faster yet still high-fidelity results.

[0017] As described above, wave-based simulations can also be partially geometric acoustic simulations. By combining or merging wave-based and geometric acoustic simulations, optimal simulation can be achieved while considering simulation speed and accuracy.

[0018] The merging and blending of wave-based simulations and geometric acoustic simulations can be accomplished by merging and blending the wave-based impulse response obtained from the wave-based simulation and the geometric impulse response obtained from the geometric acoustic simulation. This requires careful consideration to avoid introducing non-physical artifacts into the combined broadband solution that forms the simulated impulse response.

[0019] Simulated wave-based propagation and simulated geometry-based acoustic propagation can be simulated within a first frequency range and a second frequency range, respectively. Simulated wave-based propagation can be performed using wave-based simulation, and simulated geometry-based acoustic propagation can be performed using geometry-based acoustic simulation. The first and second frequency ranges can partially or completely overlap. The transition frequency is defined as the frequency at which the first and second frequency ranges overlap. The transition frequency can also be defined as the center frequency of the overlap between the first and second frequency ranges.

[0020] The transition frequency can be a default transition frequency. The default transition frequency can be proportional to the square or cube root of the volume of the 3D room model. The default transition frequency can be a simplified calculation based on the Schroeder frequency, which is assumed to have a reverberation time of 1 second and multiplied by a factor. The reverberation time can be assumed to be 1 second. The factor can be 0.5 to 2, preferably 0.1 to 5, more preferably 0.1 to 10. The factor can be selected based on the volume of the 3D room model. The reverberation time can be selected as a trade-off between simulation time and simulation accuracy. A low reverberation time may produce inaccurate results, resulting in an inaccurate simulated impulse response, while a high reverberation time may significantly increase the simulation time without increasing the accuracy of the simulated impulse response.

[0021] The first frequency range and the second frequency range can also be referred to as the first audio frequency range and the second audio frequency range, respectively.

[0022] Therefore, the step of merging wave-based impulse responses and geometry-based impulse responses may also include calibrating the acoustic power of at least one room sound source in the 3D model so that they are matched in both the wave-based solver and the geometry-acoustic solver.

[0023] Calibration can be accomplished, for example, by estimating the sound pressure level radiated from at least one room sound source at a distance of 1 meter under free-field conditions. This could be, for instance, in the frequency range near the transition frequency between a first and second sound frequency range. The pressure level can also be estimated at other distances such as 0.5 meters, 2 meters, and / or 3 meters. This distance can be between 0.5 meters and an extension of the 3D room model being simulated.

[0024] The merging step can even include ensuring that the boundary conditions are coherent for both the wave-based solver and the geometric acoustic solver. If the boundary inputs are not aligned, there is a risk of artifacts (such as phase mismatch) and / or undesirable results occurring in the transition frequency region.

[0025] Preferably, the sound pressure level radiated from at least one room sound source at 1 meter can be 94 dB sound pressure level (SPL). The step of calibrating the acoustic power of at least one room sound source in the 3D room model such that the at least one room sound source is matched in both the wave-based solver and the geometric acoustic solver can be performed independently in a source correction step or calibration step included in the wave-based solver and / or the geometric acoustic solver. More preferably, the calibration step can be performed independently in image source simulation and / or ray tracing simulation, wherein the image source simulation and / or ray tracing simulation is included in the geometric acoustic solver.

[0026] Advantageously, the wave-based impulse response can be digitally filtered using a low-pass filter whose cutoff frequency is substantially located at the transition frequency, thereby creating a filtered wave-based impulse response. This can be performed before the step of calibrating the acoustic power. Preferably, the low-pass filter is a fourth-order Butterworth low-pass filter. A Butterworth filter is a type of signal processing filter designed to have the flattest possible frequency response in the passband. The flatness of the frequency response in the passband can be understood in the context of signal processing to minimize distortion and / or preserve the shape of the impulse response. Preferably, for acoustic simulations as discussed herein, a flat passband response of a Butterworth filter may be desirable because it helps avoid coloration of the sound. Any uneven attenuation within the passband can alter the timbre or pitch of the audio signal, potentially leading to a less faithful reproduction of the sound. The order of the filter can be a trade-off between having a sharp cutoff frequency and avoiding ripples in the time domain. The inventors have found that fourth order may be sufficient to achieve the desired filtering.

[0027] Both the wave-based solver and the geometric acoustic solver can output the monoaural and spatial high-fidelity stereo impulse responses for each listening point of a 3D room model, as well as the acoustic parameters derived from the mentioned impulse responses.

[0028] The merging step can digitally filter the geometric acoustic impulse response using a high-pass filter whose cutoff frequency is substantially located at the transition frequency, thereby producing a filtered geometric acoustic impulse response. This can be performed after or before the step of calibrating the acoustic power. Preferably, the step of calibrating the acoustic power can be performed before digitally filtering the geometric acoustic impulse response using a high-pass filter whose cutoff frequency is substantially located at the transition frequency. Preferably, the high-pass filter is a fourth-order Butterworth high-pass filter.

[0029] The merging step can sum the filtered wave-based impulse response and the filtered geometric acoustic impulse response to create a hybrid impulse response, which combines the filtered wave-based impulse response and the filtered geometric acoustic impulse response. The hybrid impulse response can be an impulse response as described herein, where the impulse response can combine wave impulse responses and geometric impulse responses.

[0030] The merging step can be performed independently for all channels included in the spatial high-fidelity stereo impulse response. In other words, the merging step can be performed independently for each of the multiple room impulse responses, where each room impulse response describes the transmitted room impulse received at a corresponding location in multiple digital representations of the room receiver.

[0031] In one embodiment, the device receiver array includes devices that involve internal issues.

[0032] In another embodiment, room receiver array processing involves additional mathematical steps to address external problems.

[0033] When the device-related transfer function and spatial room impulse response are encoded in high-fidelity stereo, they are typically combined directly into a device-specific room impulse response of the lowest high-fidelity stereo order corresponding to either N (device) or N (room). This also allows the device to be freely oriented / rotated relative to the room because the high-fidelity stereo encoding includes spatial positioning information.

[0034] Since the computer-implemented methods discussed herein are executed on a processor, computer, or similar apparatus for performing the computer-implemented methods, it should be understood that, unless otherwise described, these steps are generally performed digitally. Therefore, references to, for example, receivers, microphones, transmitters, speakers, and sound sources should be understood as digital representations simulating or analogizing the functionality of the corresponding physical components.

[0035] Similarly, device models and room models are also digital representations of physical or potentially physical elements, and can be represented as grids or other digital representations.

[0036] Although transmitted, received, analog, or generated signals are also used in digital environments (e.g., within analog environments), these can be processed, for example, by a digital-to-analog converter (DAC) for playback in a physical environment. For instance, a device-specific room impulse response (DSRIR) can be convolved with a silenced sound signal to generate a physical audio experience of how a particular device receives sound in a particular room.

[0037] Detailed description In one embodiment, generating the device-related transfer function includes obtaining a 3D device model representing the geometry of the device and the position of at least the first microphone on the 3D device model. The 3D device model can be, for example, a device mesh model representing the geometry of the device and the position of at least the first microphone on the device mesh model. Meshing can be understood as a common way to generate models in digital environments such as computers. The mesh can preferably be a discretization of the geometry into small, simple shapes. The shapes can be triangles or quadrilaterals in 2D, and / or tetrahedrons or hexahedrons in 3D. Mesh density control can determine an appropriate mesh density because a mesh that is too coarse may lead to inaccurate results, while a mesh that is too fine may increase computational costs and simulation time, or sometimes cause simulations using the mesh to fail to converge.

[0038] A mesh can be understood as a polygonal mesh, which is a collection of vertices, edges, and faces that define the shape of a polyhedral object. Throughout this patent application, the term mesh element can be a face, such as a triangle, quadrilateral, other simple convex polygons, or any other combination thereof.

[0039] In one embodiment, at least one direction is at least two directions, at least three directions, at least four directions, or at least five directions.

[0040] Obtaining a 3D device model can be accomplished in various ways. For example, a physical device can be scanned to obtain a 3D device model. Different 3D scanning methods can be used to scan the physical device, such as handheld 3D scanners, desktop 3D scanners, photogrammetry, and / or LiDAR scanning. A handheld 3D scanner can be a device such as the Artec Eva or a structural sensor, which can capture physical objects in 3D. LiDAR scanning is a laser-based scanning method used to capture large environments or detailed objects. A 3D device model can be digitally obtained by, for example, modeling the 3D device using CAD (Computer-Aided Design) software. A 3D device model can also be obtained by loading a file (such as an STL file, a common format for storing digital models) onto a computer.

[0041] In another embodiment, generating the device-related transfer function includes arranging digital representations of an array of device receivers, comprising multiple digital representations of device receivers, around a 3D device model (such as a device mesh model). The device receiver array is arranged such that any digital representation of a device receiver is at least a predetermined distance from the 3D device model.

[0042] In one embodiment, the array of device receivers, comprising multiple digital representations of the device receivers, is spherical in shape. Alternatively, the shape may be an offset shape, wherein the digital representations of the device receivers are placed / arranged at a predetermined offset distance from the device mesh model. The predetermined offset distance may be the same as the predetermined distance discussed or described herein.

[0043] The predetermined distance may, for example, include the radius of the device receiver array and an additional distance to properly surround the 3D device model. This may be determined, for example, such that the distance from any device receiver on the array to any point on the 3D model of the device is preferably not less than the predetermined distance.

[0044] In one embodiment, the predetermined distance is 0.5-1.5 meters, preferably 0.8-1.2 meters, or most preferably 1 meter. Such a predetermined distance has proven to be a good choice for general room simulation, or preferably for generating and / or simulating the transfer function of the apparatus. Choosing a distance of 1 meter is advantageous because at this predetermined distance, the sound wave can be considered a plane wave. This approximation simplifies aspects such as maintaining consistent amplitude and phase and applying easier boundary conditions. Therefore, the reduced complexity due to the simpler representation and propagation of plane waves makes it possible to simulate plane waves more effectively.

[0045] The number of device receivers in the device receiver array can then be selected by determining the minimum order N, which is (N+1). 2 Yes. This can be achieved by multiplying by a factor, such as 1.5 or 2.0, to obtain higher fidelity, but at the cost of increased simulation time.

[0046] When applying a wave-based solver to generate at least a first device-dependent transfer function for a device including at least one microphone, the number of sound sources typically determines the time and resources required to determine the acoustic simulation. Therefore, if at least one microphone is set up as a microphone, all receivers in the device array must be used as sound sources and emit pulse signals. However, each of these would have to be solved individually, and the time and resources required for solving would increase significantly based on the number of sound sources in the receiver array. Conversely, to determine the device-dependent transfer function as discussed herein, using the reciprocity law, it is advantageous to use the first microphone as a sound source to emit the first pulse signal. This would greatly save acoustic simulation time because only one simulation per microphone would characterize all microphones of the device, rather than one simulation per sound source in the receiver array.

[0047] In one embodiment where the 3D device model is a mesh model, a first nearest mesh element on the 3D device model that is closest to at least the first microphone can be determined. In one embodiment, the first nearest mesh element can be used as a sound source for emitting a first pulse signal. By having and / or identifying the first nearest mesh element on the 3D device model that is closest to at least the first microphone, a sound source for emitting the first pulse signal can be configured.

[0048] In a preferred embodiment, the first closest mesh element can be automatically determined using a computer-implemented method. This computer-implemented method can identify the positions of all mesh elements in the mesh model. When designing a 3D device model, the user can program the position of at least the first microphone. Therefore, the computer-implemented method can calculate the mesh distances between the position of at least the first microphone and multiple mesh elements of the mesh model. The mesh element located at the shortest distance from the position of at least the first microphone can be identified as the first closest mesh element using this computer-implemented method.

[0049] Based on the first pulse response transmitted, a wave-based solver can be used to determine multiple first device pulse responses, each of which can describe the pulse response of the first pulse signal received at a corresponding device receiver. The corresponding device receiver is at least one of a device receiver from an array of device receivers as described herein.

[0050] Therefore, a first device-dependent transfer function for a first microphone can be generated by combining multiple first device impulse responses.

[0051] The first pulse signal should preferably be a perfect pulse with a flat spectrum. However, this may not be possible, and in order to increase the fidelity of the generated device-related transfer function, the first device impulse response should preferably be source-corrected using a reference signal.

[0052] In one embodiment, this can be accomplished by arranging a digital representation of a first source correction microphone located at a first source distance from at least a first microphone or a first source distance from a first closest grid element, wherein the first source distance is less than a predetermined distance.

[0053] Therefore, a wave-based solver can be used to determine the first source correction signal, which describes the first pulse signal received at the first source correction microphone.

[0054] The plurality of first-source-corrected device impulse responses can then be determined by source-correcting each of the plurality of first device impulse responses using the first source correction signal, and the first device-related transfer function of the device for the first microphone can then be determined by combining the plurality of first-source-corrected device impulse responses.

[0055] In one embodiment, generating the device-dependent transfer function includes determining the energy content at at least one frequency of the first device-dependent transfer function. As will be discussed, this can, for example, be used to determine the high-fidelity stereo order N that can be used when generating a spatial room impulse response (SRIR).

[0056] In one embodiment, determining the energy content at at least one frequency of the first device-related transfer function includes determining different high-fidelity stereo orders to identify different levels of energy content.

[0057] In one embodiment, determining the energy content at at least one frequency of the first device-related transfer function includes determining the high-fidelity stereo order N of the energy content at at least one frequency.

[0058] In another embodiment, a high-fidelity stereo order N is determined for multiple frequencies, wherein the energy content of each frequency is determined.

[0059] In yet another embodiment, the high-fidelity stereo order N for the energy content is determined based on the sum of the high-fidelity stereo coefficients for each order N, and then normalized to uniformity for each frequency.

[0060] In one embodiment, the energy content is determined within a frequency range, such as 0 to 20 kHz, 0 to 10 kHz, 10 to 20 kHz, 0 to 9 kHz, 0 to 8 kHz, 0 to 7 kHz, 0 to 6 kHz, 0 to 5 kHz, 0 to 4 kHz, 0 to 3 kHz, 0 to 2 kHz, or 0 to 1 kHz. Preferably, the frequency range can be included within the audible spectrum. The maximum frequency of this range can indicate the desired high-fidelity stereo order N. Typically, a higher maximum frequency requires a larger high-fidelity stereo order N.

[0061] In one embodiment, the device includes a plurality of microphones, such as a second, third, fourth, and fifth microphone. In this case, the computer-implemented method as discussed herein is repeated for each microphone. In other words, each of the plurality of microphones can be considered as a first microphone, such that a plurality of device-dependent transfer functions, such as second, third, fourth, and fifth device-dependent transfer functions of the device, are generated for each of the microphones.

[0062] In one embodiment, generating at least a first device-related transfer function (DRTF) includes: - Obtain a 3D box model including a highly sound-absorbing surface, or a 3D box model with a predetermined size, such that the first pulse signal received by each of the multiple digital representations from the sound source to the device receiver does not include reflections caused by the surface of the 3D box model. - Arrange the device receiver array and device mesh model in a 3D box model.

[0063] By obtaining a 3D box model including highly sound-absorbing surfaces, a first pulse signal emitted from a first nearest grid element is not reflected by the surface of the 3D box model, and therefore is not received by the digital representation of a device receiver array comprising a plurality of digital representations of the device receiver. Preferably, the plurality of digital representations of the device receiver can receive the incoming pulse signal, and preferably do not receive reflections caused by any surface, obstacle, or geometry outside the digital representations of the device receiver array. In another embodiment, the 3D box model has a predetermined size such that the first pulse signal received from the sound source to each of the plurality of digital representations of the device receiver does not include reflections caused by the surface of the 3D box model. The predetermined size, such as the surface being distant from the plurality of digital representations of the device receiver, can be estimated. Making the predetermined size too high may be expensive in terms of computational cost and time, therefore the predetermined size should be estimated and / or calculated, such that the simulation of pulse signal propagation may stop before the pulse signal reaches or substantially reaches the surface of the 3D box model.

[0064] Alternatively, the 3D box model can be a 3D spherical model. Preferably, the 3D spherical model can be composed of spheres.

[0065] In one embodiment, generating a spatial room impulse response includes obtaining a 3D room model representing the geometry of the room and at least one acoustic characteristic or at least one acoustic property. As mentioned above, obtaining a 3D model, such as a 3D room model, can be accomplished in different ways.

[0066] Furthermore, the 3D room model may include at least one acoustic property, which may be, for example, the sound absorption value of different materials included in the 3D room model, and may be different for different elements in the model (such as elements representing windows, carpets, furniture, etc.). At least one acoustic property may be complex surface impedance. The 3D room model may include at least one boundary. At least one boundary may include at least one acoustic property or at least one acoustic feature.

[0067] At least one boundary of a 3D room model can refer to a surface or edge that defines the limits or extent of the 3D room model. At least one boundary can contain a physical barrier that can surround the space or volume of the 3D room model. At least one boundary can be a wall, floor, ceiling, door, window, corner, edge, or any combination thereof.

[0068] In one embodiment, at least one digital representation of at least one room sound source is arranged in a 3D room model. The at least one room sound source can be, for example, an omnidirectional or directional sound source. A directional sound source can be an audio transmitter that emits sound uniformly in a specific direction rather than in all directions. A directional sound source may be characterized by its ability to focus sound waves or acoustic waves along a specific direction or path, which can result in higher sound intensity in that specific direction and lower intensity outside that specific direction. The directional pattern of a directional sound source can be narrow or wide. The directional pattern can include cardioid, supercardioid, hypercardioid, bidirectional, any arbitrary directional pattern, or any combination thereof. Any arbitrary directional pattern can be other directional patterns, which may not be a combination of directional patterns as described herein. The directional sound source can use beamforming, which can manipulate the phase and amplitude of sound or acoustic waves emitted from at least one loudspeaker to create a focused beam. The directional sound source can be modeled using spherical harmonic functions.

[0069] In another embodiment, generating a spatial room impulse response may further include a digital representation of an array of room receivers, comprising multiple digital representations of the room receivers, wherein the array of room receivers is centered on at least one listening point in a 3D room model.

[0070] In one embodiment, the number of digital representations for the room receivers can be determined based on the energy content at at least one frequency of the first device's associated transfer function. Preferably, the energy content can generate information about the high-fidelity stereo order to be used, thereby generating information about the number of receivers to be used in the digital representations for the room receivers.

[0071] In one embodiment, determining the number of digital representations of the room receiver based on the energy content of at least one frequency of the device-dependent transfer function (DRTF) further includes determining the number of digital representations of the room receiver based on the high-fidelity stereo order N, wherein the number of digital representations of the room receiver can be (N+1). 2 1.5 (N+1) 2 or 2 (N+1) 2 More digital representations with room receivers can provide better accuracy at the cost of generating more data.

[0072] In another embodiment, generating a spatial room impulse response may include digitally transmitting a room impulse signal from at least one audio source. Multiple room impulse responses may be determined using at least one wave-based frequency-based solver, wherein each room impulse response describes the transmitted room impulse signal received at a corresponding location in a plurality of digital representations of the room receiver. The spatial room impulse response may be based on multiple room impulse responses.

[0073] In another embodiment, generating the space room impulse response may include using a geometric acoustic solver with at least one geometric acoustic frequency to determine a second plurality of room impulse responses.

[0074] Multiple impulse responses generated using a wave-based solver and a second set of multiple room impulse responses generated using a geometric acoustic solver can be combined in one embodiment to generate multiple combined room impulse responses.

[0075] For example, in another embodiment, multiple room impulse responses can be generated in the low frequencies of the acoustic spectrum using a wave-based solver, and a second plurality of impulse responses can be generated in the high frequencies of the acoustic spectrum using a geometric acoustic solver.

[0076] The sound spectrum can be, for example, between 0 and 20 kHz, such as between 0 and 15 kHz, such as between 0 and 12 kHz, such as between 0 and 10 kHz, such as between 0 and 8 kHz, such as between 0 and 6 kHz, such as between 20 Hz and 20 kHz, such as between 20 Hz and 15 kHz, such as between 20 Hz and 12 kHz, such as between 20 Hz and 10 kHz, such as between 20 Hz and 8 kHz, such as between 20 Hz and 6 kHz. Preferably, the sound spectrum as defined herein can preferably be the sound spectrum heard by humans. Some frequency ranges of the sound spectrum may be more useful for such acoustic applications (such as human speech), where most human speech frequencies can generally be included between 100 and 17 kHz, which may include the fundamental frequency and harmonics of human speech. Male speech may cover a frequency range of 100 Hz to 8 kHz, while female speech may cover a frequency range of 350 Hz to 17 kHz.

[0077] In one embodiment, the low frequencies of the sound spectrum include those between 0 and 20 kHz, such as between 0 and 15 kHz, such as between 0 and 12 kHz, such as between 0 and 10 kHz, such as between 0 and 8 kHz, such as between 0 and 6 kHz, such as between 20 Hz and 20 kHz, such as between 20 Hz and 15 kHz, such as between 20 Hz and 12 kHz, such as between 20 Hz and 10 kHz, such as between 20 Hz and 8 kHz, such as between 20 Hz and 6 kHz, such as between 20 Hz and 5 kHz, such as between 20 Hz and 4 kHz, such as between 20 Hz and 3 kHz, such as between 20 Hz and 2 kHz, such as between 20 Hz and 1.5 kHz, such as between 20 Hz and 1 kHz.

[0078] In one embodiment, the high frequencies of the sound spectrum include those between 1 kHz and 20 kHz, such as between 1.5 kHz and 20 kHz, such as between 2 kHz and 20 kHz, such as between 3 kHz and 20 kHz, such as between 4 kHz and 20 kHz, such as between 5 kHz and 20 kHz, such as between 6 kHz and 20 kHz, such as between 8 kHz and 20 kHz, such as between 10 kHz and 20 kHz, such as between 12 kHz and 20 kHz, such as between 1 kHz and 15 kHz, such as between 1.5 kHz and 15 kHz, such as between 2 kHz and 15 kHz, such as between 3 kHz and 15 kHz, etc. Such as between 4kHz and 15kHz, such as between 5kHz and 15kHz, such as between 6kHz and 15kHz, such as between 8kHz and 15kHz, such as between 10kHz and 15kHz, such as between 12kHz and 15kHz, such as between 1kHz and 12kHz, such as between 1.5kHz and 12kHz, such as between 2kHz and 12kHz, such as between 3kHz and 12kHz, such as between 4kHz and 12kHz, such as between 5kHz and 12kHz, such as between 6kHz and 12kHz, such as between 8kHz and 12kHz, such as between 10kHz and 12kHz.

[0079] As discussed in this paper, using high-fidelity stereo to encode and decode acoustic signals can offer numerous advantages, particularly when processing high-fidelity acoustic data and signals, providing a flexible means of communication for many different applications and uses. For example, since the device-dependent transfer function and spatial impulse response of the first (and possibly subsequent) device are generated separately, but through relationships, for example, that can be used to determine the energy content of the high-fidelity stereo order N, the device can be rotated relative to the room. A new device-dependent transfer function can be generated for a new device, and combined with an already generated spatial room impulse response having a corresponding high-fidelity stereo order N (or higher), and vice versa, a new spatial room impulse response can be generated for a new room.

[0080] In one embodiment, high-fidelity stereo can therefore be used to encode and decode the generated device-specific room impulse response. However, generating a device-specific room impulse response (DSRIR) by combining the device-dependent transfer function and the spatial room impulse response can be avoided, and at least the first device-dependent transfer function (DRTF) and spatial impulse response (SRIR) can be processed separately or further processed in a matrix. For example, DSRIR can be used to analyze the sound field in a room or around the device, or to generate spatial sound field visualizations.

[0081] In one embodiment of the above aspects, such advantages become apparent in a computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein the device includes at least a first microphone, the method comprising: - Generate at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device received by the at least first microphone, wherein generating the first device-related transfer function further includes A device mesh model is obtained, representing the geometry of the device and the position of at least the first microphone on the device mesh model. The digital representations of an array of device receivers, including multiple digital representations of the device receivers, are arranged around the device mesh model such that the distance between any digital representation of the device receiver and the device mesh model is not less than a predetermined distance. On the device mesh model, determine the first nearest mesh element that is closest to at least the first microphone. A digital representation of a first source correction microphone positioned at a first source distance from the first nearest grid element, wherein the first source distance is less than the predetermined distance. The first pulse signal is digitally emitted using the first closest grid element as the sound source. A wave-based solver is used to determine a first source correction signal, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone. A wave-based solver is used to determine a plurality of first device impulse responses, wherein each first device impulse response describes the impulse response of the first pulse signal received at a corresponding device receiver. The plurality of first-source-corrected device impulse responses are determined by performing source correction on each of the plurality of first device impulse responses using the first source correction signal. The first device-related transfer function for the first microphone is generated by combining the device impulse responses of the plurality of first source corrections. Determine the energy content at at least one frequency of the relevant transfer function of the first device. - Generate a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room received from at least one room sound source in the room and from at least one direction at at least one listening point in the room, wherein generating the spatial room impulse response further includes Obtain a 3D room model representing the geometry of the room and at least one acoustic property. At least one digital representation of at least one room sound source is arranged in the 3D room model. A digital representation of a room receiver array comprising multiple digital representations of a room receiver, wherein the room receiver array is centered on at least one listening point in the 3D room model, and wherein the number of digital representations of the room receivers is determined based on the energy content of at least one frequency of the first device-related transfer function. The room pulse signal is digitally transmitted from the at least one audio source. At least one wave-based frequency-based solver is used to determine multiple room impulse responses, wherein each room impulse response describes a transmitted room impulse received at a corresponding location in multiple digital representations of the room receiver. A spatial room impulse response is generated based on the multiple room impulse responses. - The device-specific room impulse response (DSRIR) is generated by combining the device-related transfer function and the room impulse response.

[0082] Device-specific room impulse responses can be convolved with anechoic sounds to generate device-specific room sounds. This provides audio rendering of how sound can be captured by devices in a room. Using high-order high-fidelity stereo allows for appropriate high-fidelity stereo rendering, generating high-fidelity, device-specific room sounds.

[0083] It is possible to evaluate the device-specific room impulse response to consider many different scenarios with device location and orientation in multiple 3D room models. This allows users to evaluate the performance of any device in any room and at any orientation based on a room simulation and a device simulation, which can be defined separately as SRIR generation and DRTF generation.

[0084] On the other hand, a computer-implemented method for recovering the device-dependent transfer function (DRTF) of a master device is disclosed. The computer-implemented method for recovering the DRTF of a master device, wherein the DRTF of the master device includes at least one loss magnitude level, may include the following steps: - Obtain multiple amplitude levels of the DRTF of the master device as a function of multiple elevation angles and multiple azimuth angles, wherein the DRTF of the master device includes at least one lost amplitude level; - Obtain the matrix including the spherical harmonic basis functions; - Applying the truncated singular value decomposition method to the matrix comprising spherical harmonic basis functions, wherein the truncated singular value decomposition method includes the following steps: Singular values ​​are generated based on the matrix including spherical harmonic basis functions according to singular value decomposition (SVD). Select the top singular value and the corresponding set of singular vectors, wherein the top singular value and the corresponding set of singular vectors are selected based on a predetermined threshold; Based on singular value decomposition (SVD), a secondary matrix including spherical harmonic basis functions is generated based on the top selection of the singular values ​​and the corresponding set of singular vectors. - Multiply the matrix corresponding to multiple amplitude levels of the DRTF of the main device with the inverse matrix of the secondary matrix to generate multiple high-fidelity stereo coefficients; - Generate multiple restored amplitude levels of the DRTF of the master device based on the multiple high-fidelity stereo coefficients, wherein the multiple restored amplitude levels of the DRTF of the master device include an image of the amplitude levels of the DRTF of the master device and a reconstruction of the at least one lost amplitude level.

[0085] In one embodiment, the amplitude level of the DRTF of the master device is expressed as a function of elevation and azimuth.

[0086] In another embodiment, the amplitude level of the DRTF of the main device is used for the DRTF frequency. The DRTF frequency can include between 20 Hz and 20 kHz. Preferably, the DRTF frequency can include between 100 Hz and 10 kHz.

[0087] In a preferred embodiment, the amplitude level is a complex value, or the amplitude level is represented in the form of a complex number.

[0088] The DRTF of the master device disclosed herein can be generated according to the computer implementation of the method for generating device-related transfer functions as described herein. Attached Figure Description

[0089] In the following description, embodiments and examples will be presented in more detail with reference to the accompanying drawings: Figure 1 An embodiment of a method for generating a device-specific room impulse response as disclosed herein is illustrated schematically.

[0090] Figure 2A-2B The amplitude level of the master device's DRTF at 1000 Hz as a function of elevation and azimuth with missing data is shown, as well as the amplitude level of the same DRTF reconstructed from the high-fidelity stereo representation of the slave device at 1000 Hz as a function of elevation and azimuth, wherein the missing data is recovered using the amplitude level of the DRTF reconstruction or recovery method as described herein. Detailed Implementation

[0091] The computer-implemented method 100 for generating device-specific room impulse responses, as discussed in this paper, is... Figure 1 It is shown schematically in the diagram.

[0092] This method can be considered to be formed by two sub-methods, one showing method 101 for generating device-related transfer function (DRTF) and the other showing method 102 for generating space room impulse response (SRIR).

[0093] To generate the device-related transfer function, a 3D model of device 110 is provided as input to the method. The 3D model of device 110 includes three microphones: a first microphone 111, a second microphone 112, and a third microphone 113.

[0094] When entering this method, if the 3D model of device 110 has not yet been meshed as a model, it is meshed and placed in a simulation tool that applies a wave-based solver. As should be understood herein, the wave-based solver applies a wave-based method that can apply numerical techniques to directly solve the governing partial differential equations describing wave motion in a virtual domain, such as those representing air volume. This could be, for example, the wave equation in the time domain or the Helmholtz equation in the frequency domain. Thus, the concept of a different wave-based method is to divide the virtual domain of interest into small subdomains (discretization) and solve algebraic equations in each subdomain. Therefore, the wave-based method used in the wave-based solver disclosed herein can be understood as a method of solving partial differential equations using discretization techniques. The Treble simulation tool can be used, for example, to perform such wave-based simulations using a wave-based solver.

[0095] Device receiver array 114 is arranged around the 3D model of device 110 and includes multiple device receivers 114' (not all are indicated by reference numerals in the figures for simplicity). In this case, the device receiver array 114 forms a spherical pattern around the 3D model of device 110. A spherical pattern is generally preferred because it facilitates encoding signals into high-fidelity stereo using spherical harmonics. Preferably, the spherical pattern can be a Lebedev grid.

[0096] The radius of the receiver array was determined such that the distance from any receiver on the array to any point on the 3D model of the device is no less than a predetermined distance set to 1 meter in this example. This has proven to be a good choice for general room simulation to ensure far-field acoustic radiation conditions. The formula N>2 can be used. pi f R(device) / c determines the high-fidelity stereo order N, where 'R(device)' is the maximum distance from the center of the receiver array to any point on the 3D model of the device, 'f' is the frequency considered, and 'c' is the speed of sound (typically 344 m / s). The desired number of receivers in the array can then be selected using the minimum order N, which is (N+1). 2Yes. This can be achieved by multiplying by a factor, such as 1.5 or 2.0, to obtain higher fidelity, but at the cost of increased data generation and simulation time.

[0097] However, other patterns can be used, where the additional step of transposing the shape onto a sphere can be used for spherical harmonics. Such a pattern can be offset from the surface of the 3D model of the device by a predetermined distance, such as the 1 meter discussed.

[0098] Microphones 111, 112, and 113 on the 3D model of device 110 are then configured to serve as sound sources. Although not shown, this can be accomplished, for example, by determining the nearest mesh element on the mesh model for each microphone and using the nearest mesh element to virtually emit pulse signals, as will be described. A remeshing step can be performed to remesh the nearest mesh element to a specific size or shape so as to be as close as possible to the size of the microphone. Preferably, an initial meshing step is performed, for example, the location of the microphone is specified as input to the meshing tool being used, thereby generating appropriate mesh elements located at the microphone locations.

[0099] Switching the function of the microphones on the device to act as sound sources, such as speakers or other sound emitters, significantly increases the speed of subsequent wave-based solvers because multiple sound sources significantly impact processing speed. Therefore, when the microphones on the device are set to function as microphones, each receiver 114' in the receiver array 114 must be used as a sound source, which greatly increases processing speed because the number of receivers in the array typically far exceeds the number of microphones on the device. Thus, the ability to switch the function for wave-based simulation significantly reduces time and is a major advantage of using computer-implemented simulation tools such as Treble software.

[0100] The method then performs wave-based simulations one at a time for each microphone 111, 112, 113, where the microphone, or the nearest grid element in this case, emits a pulse, and the signal received at each receiver 114' in the array is recorded.

[0101] Ideally, the pulse would have a flat spectrum; however, this is generally not possible. Therefore, source correction of the signal received at each receiver is performed [2]. Reference signals for source correction are recorded using source correction receivers 121, 122, 123 placed very close to the microphone placement used as a loudspeaker (e.g., 1 mm in front of the microphone or grid used as a transmitter). Thus, in the current embodiment, the three source correction receivers 121, 122, and 123 are placed 1 mm in front of the microphones 111, 112, and 113 or the corresponding closest grid.

[0102] The source correction signal of the pulse from one of the microphones on the device, received at each array receiver, forms the transfer function describing that particular microphone. Thus, in the present case, since there are three microphones on the device, three different transfer functions are generated, which are also described herein as first, second, and third device-dependent transfer functions 145, 146, and 147, and together they form a general device-dependent transfer function, which can, for example, be stored as a three-dimensional matrix.

[0103] In addition to the device-related transfer functions of the first, second, and third devices, the method for generating the device-related transfer functions also generates an energy map after encoding into high-fidelity stereo. An energy map 140 is generated, where the energy at different frequencies (Hz) is used to determine the high-fidelity stereo order (n), and can be used in room simulation (when generating spatial room impulse responses, as will be discussed) to provide efficient simulation and allow for free rotation of the device in high-fidelity stereo encoding and decoding, as will be discussed below. For example, in the current case, the frequencies of the high-fidelity stereo curve 141 are determined, indicating an energy content of 95% at the corresponding frequencies along the x-axis, and are used to determine the high-fidelity stereo order (n) on the y-axis.

[0104] Method 102 for generating spatial room impulse responses uses a 3D room model 150, which represents the geometry of the room as input. Sound sources 152 and listening points 151 are arranged in the room model. The 3D room model may also include the geometry of furniture, such as tables and chairs, door openings, and / or monitors. It may also include the acoustic properties of different geometries and materials, such as windows, carpets, different wall materials, etc.

[0105] Subsequently, the space room impulse response 170 is determined based on the wave-based space room impulse response 171 for the low-to-mid frequency range and the geometric acoustic space room impulse response 172 for the mid-to-high frequency range of the audible spectrum.

[0106] Spatial impulse responses 171 and 172 embed spatiotemporal information about the direction of arrival of the incoming sound waves at the receiver location. Typically, a spatial impulse response comprises multiple single-channel impulse responses, each recording sound from a specific direction or angle at the same listening point.

[0107] The wave-based spatial impulse response 171 can be constructed in a simulation by emitting a pulse signal from a sound source 152 and recording multiple room impulse responses at multiple room receivers 160' surrounding the listening point 151 in a room receiver array 160. The room receiver array comprises room receivers 160' arranged in a spherical array shape around the listening position 151 (not all are indicated by reference numerals in the figures for simplicity). The receivers can be omnidirectional or have a cardioid directional pattern to optimize the operating frequency range of the array.

[0108] The number and size of one or more receivers used in the room receiver array are initially determined by the high-fidelity stereo order N derived from energy diagram 140, where the high-fidelity stereo curve illustrates the order N that determines the desired frequency range of the wave-based spatial impulse response 171. Given the order N, as discussed above, the number of receivers 160' can be determined by (N+1). 2 It can be determined that it can be multiplied by a factor, such as 1.5 or 2.0, to obtain higher fidelity.

[0109] Furthermore, the radius R (of the array) can be expressed using the formula N>2 discussed above. pi f R(array) / c is used to determine this. For a given high-fidelity stereo order N and radius R(array), this means the maximum frequency is constrained to f due to spatial aliasing. <N c / (2) pi R (array)). Therefore, R (array) must be selected based on the maximum frequency of interest. Once the impulse response has been recorded for all receivers in the array, the spatial impulse response can be encoded into high-fidelity stereo.

[0110] The geometric acoustic spatial impulse response 172 can be determined by analyzing the incoming directions of all image sources and rays at listening point 151 using common image source and ray tracing techniques. The geometric acoustic spatial impulse response can then be directly generated and encoded into high-fidelity stereo.

[0111] The wave-based space room impulse response 171 and the geometric acoustic space room impulse response 172 can then be combined or mixed into a combined space room impulse response 170. In some cases, the wave-based space room impulse response or the geometric acoustic space room impulse response can be used independently.

[0112] Therefore, the device-specific room impulse response 180 is provided by the first, second, and third device-related transfer functions 145, 146, and 147 herein, as well as the combined space room impulse response 170, and the first, second, and third device-related transfer functions 145, 146, and 147 together form a general device-related transfer function, which may, for example, be stored as a matrix.

[0113] Figure 2A-2B The amplitude level of the master device's DRTF at 1000 Hz as a function of elevation and azimuth with missing data is shown, as well as the amplitude level of the same DRTF reconstructed from the high-fidelity stereo representation of the slave device at 1000 Hz as a function of elevation and azimuth, wherein the missing data is recovered using the amplitude level of the DRTF reconstruction or recovery method as described herein.

[0114] Figure 2A The amplitude level of the DRTF of a master device at 1000Hz, as a function of elevation and azimuth, is shown, indicating a loss of amplitude level or missing data. For all azimuth angles, the loss of amplitude level is essentially located below -45 degrees at elevation angles. This can correspond to DRTF measurement or DRTF simulation cases where the device is arranged on a table or support, preferably horizontally, which would impair or challenge the measurement or simulation of DRTF for elevation angles, which could correspond to a position below the table or support on which the master device is arranged. Other cases can be similar, such as when the master device is close to a wall, where other angles, such as azimuth, will be lost or potentially damaged for positions behind a wall. If spherical harmonic / high-fidelity stereo decomposition is performed, then... Figure 2A The amplitude level of the DRTF of the master device shown may present a challenge. A missing amplitude level will render the spherical harmonic / high-fidelity stereo decomposition suboptimal. This challenge can be addressed, for example, using DRTF reconstruction or recovery methods, where the missing amplitude level is filled with a random amplitude level, preferably a low amplitude level, more preferably an amplitude level lower than that measured or simulated in the DRTF of the master device. This method can be cumbersome, situation-dependent, and may not prove sufficiently accurate.

[0115] In one embodiment, during or before the high-fidelity stereo decomposition process, a truncated singular value decomposition method or approach is used to reconstruct or recover the lost amplitude level of the master device's DRTF.

[0116] Truncated singular value decomposition (SVD) methods are based on singular value decomposition (SVD). SVD can be used for dimensionality reduction, data compression, noise reduction, or solving linear systems. At the magnitude level of DRTF, such as... Figure 2AThe graph shown can be assimilated into a matrix of amplitude levels associated with each direction, where each amplitude level is an element of the matrix, and where the elements of the matrix can be complex-valued. The complex-valued elements can include phase and / or amplitude information. High-fidelity stereo decomposition of the matrix (up to order N) can involve finding (N+1) harmonic contributions associated with each spherical harmonic. 2 Coefficients. The coefficients can be directly obtained as a multiplication between the matrix of energy levels and the inverse of the matrix containing the values ​​of the spherical harmonic basis functions, evaluated for each direction of the receivers in the device array. Prior to the multiplication stage, for regularization purposes, the inverse of the matrix containing the values ​​of the spherical harmonic basis functions can be processed by SVD. By selecting an appropriate number of singular values ​​based on their relative amplitudes, proper tuning of the basis function matrix is ​​ensured, allowing for accurate estimation of the high-fidelity stereo coefficients and thus maintaining a good approximation of the amplitude levels of the DRTF after reconstruction or restoration.

[0117] As described in the preceding paragraph, this method can be described as Truncated Singular Value Decomposition (TSVD), where truncated Singular Value Decomposition can be a method of reconstructing or recovering a matrix, preferably using a finite number of singular values ​​in the context of processing lost data, such as the magnitude level of the DRTF of the master device. As described herein, the truncated Singular Value Decomposition method involves using a top selection of singular values ​​and their corresponding singular vectors, rather than using all singular values ​​and vectors. This truncation helps capture the most important features of the matrix, thereby capturing the most important features of the magnitude level of the DRTF with the magnitude level of the loss, such as... Figure 2A As shown in the diagram. By applying TSVD, lost amplitude levels can be reconstructed, simulated, or recovered without affecting non-lost amplitude levels because the TSVD method can fill the lost amplitude levels with reasonable amplitude levels derived from measured or simulated amplitude levels and reported in the original DRTF, such as... Figure 2A As shown in the image.

[0118] like Figure 2B As shown, for a frequency of 1 kHz, the TSVD method is used to achieve the amplitude level of the main device's DRTF as a function of elevation and azimuth. Figure 2B The amplitude level of the DRTF of the main device shown can be used efficiently to perform spherical harmonic / high-fidelity stereo decomposition, with smooth and efficient results.

[0119] Reference list: [1] F. Pind, “Wave-based Virtual Acoustics”, 2020, Technical University of Denmark – https: / / orbit.dtu.dk / en / publications / wave-based-virtual-acoustics .

[0120] [2] S. Sakamoto et al., “Calculation of impulse responses and acoustic parameters in a hall by the finite-difference time-domain method”, Acoust.Sci.&Tech. 29, 4 (2008).

[0121] Example List The following embodiments are disclosed in this document. 1. A computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein the device includes at least a first microphone, the method comprising: - Generate at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device received by the at least first microphone. - Generate a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room received from at least one room sound source in the room and from at least one direction at at least one listening point in the room.

[0122] 2. The computer-implemented method according to Project 1, wherein the at least one direction is at least two directions, at least three directions, at least four directions, or at least five directions.

[0123] 3. The computer-implemented method according to Project 1, wherein the computer-implemented method further includes: - The device-specific room impulse response (DSRIR) is generated by combining the device-related transfer function and the room impulse response.

[0124] 4. The computer-implemented method according to Project 1, wherein generating the device-related transfer functions includes: - Obtain a device mesh model, which represents the geometry of the device and the position of at least the first microphone on the device mesh model. - Arrange the digital representations of a device receiver array, including multiple digital representations of the receiver, around the device grid model such that the distance between any digital representation of the receiver and the device grid model is not less than a predetermined distance. - Determine the first nearest grid element on the device grid model that is closest to at least the first microphone. - A digital representation of a first source correction microphone positioned at a first source distance from the first nearest grid element, wherein the first source distance is less than the predetermined distance. - Use the first closest grid element as the sound source to digitally emit the first pulse signal. - A wave-based solver is used to determine a first source correction signal, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone. - A wave-based solver is used to determine a plurality of first device impulse responses, wherein each first device impulse response describes the impulse response of the first pulse signal received at a corresponding receiver. -The plurality of first-source-corrected device impulse responses are determined by performing source correction on each of the plurality of first device impulse responses using the first source correction signal. - The first device-related transfer function for the device used for the first microphone is generated by combining the device impulse responses of the plurality of first source corrections.

[0125] 5. The computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer function includes determining the energy content of at least one frequency of the first device-related transfer function.

[0126] 6. The computer-implemented method according to any one of the foregoing items, wherein generating the spatial room impulse response comprises, - Obtain a 3D room model representing the geometry of the room and at least one acoustic property. - At least one digital representation of at least one room sound source is arranged in the 3D room model. - A digital representation of a room receiver array comprising multiple digital representations of the receivers, wherein the room receiver array is centered on at least one listening point in the 3D room model, and wherein the number of digital representations of the receivers is determined based on the energy content of at least one frequency of the first device's associated transfer function. - Digitally transmit room pulse signals from the at least one audio source. - At least one frequency-based wave-based solver is used to determine multiple room impulse responses, wherein each room impulse response describes the transmitted room impulse received at a corresponding location in multiple digital representations of the receiver. - Generate spatial room impulse responses based on the multiple room impulse responses.

[0127] 7. The computer-implemented method according to any one of the preceding items, wherein the predetermined distance is between 0.5 and 1.5 meters, preferably between 0.8 and 1.2 meters, or most preferably 1 meter.

[0128] 8. A computer-implemented method according to any one of the preceding items, wherein the apparatus includes a plurality of microphones, such as a second, third, fourth, and fifth microphone, each of the plurality of microphones being regarded as the first microphone, such that a plurality of apparatus-related transfer functions, such as the second, third, fourth, and fifth apparatus-related transfer functions of the apparatus, are generated for each microphone.

[0129] 9. The computer-implemented method according to any one of the preceding items, wherein determining the energy content of at least one frequency of the associated transfer function of the first device includes determining different high-fidelity stereo orders to identify different levels of energy content.

[0130] 10. A computer-implemented method according to any one of the preceding items, wherein the energy content is determined for a frequency range, such as 0 to 20 kHz, such as 0 to 10 kHz, such as 10 to 20 kHz, such as 0 to 9 kHz, such as 0 to 8 kHz, such as 0 to 7 kHz, such as 0 to 6 kHz, such as 0 to 5 kHz, such as 0 to 4 kHz, such as 0 to 3 kHz, such as 0 to 2 kHz, such as 0 to 1 kHz.

[0131] 11. The computer-implemented method according to any one of the preceding items, wherein generating at least a first device-related transfer function (DRTF) comprises: - Obtain a 3D box model including a highly sound-absorbing surface or a 3D box model with a predetermined size, such that the first pulse signal received from each of the plurality of digital representations from the sound source to the receiver does not include reflections caused by the surface of the 3D box model. Arrange the device receiver array and the device mesh model in the 3D box model.

[0132] 12. The computer-implemented method according to Project 11, wherein the 3D box model is a 3D spherical model.

[0133] 13. The computer-implemented method according to any one of the preceding items, wherein generating at least a first device-related transfer function (DRTF) comprises arranging an array of device receivers, including a plurality of digital representations of the receivers, in a spherical or offset shape, wherein the digital representations of the receivers are placed / arranged at a predetermined offset distance from the device mesh model.

[0134] 14. The computer-implemented method according to any one of the preceding items, wherein the method further comprises determining a second plurality of room impulse responses using at least one geometric acoustic solver at at least one geometric acoustic frequency.

[0135] 15. A computer-implemented method according to any one of the preceding items, wherein the plurality of impulse responses generated using the wave-based solver and the second plurality of room impulse responses generated using the geometric acoustic solver are combined to generate a plurality of combined room impulse responses.

[0136] 16. The computer-implemented method according to any one of the preceding items, wherein the plurality of room impulse responses generated using the wave-based solver are generated in the low frequency of the acoustic spectrum, and the second plurality of impulse responses generated using the geometric acoustic solver are generated in the high frequency of the acoustic spectrum.

[0137] 17. The computer-implemented method according to Item 16, wherein the sound spectrum includes the range between 0 and 20 kHz, such as between 0 and 15 kHz, such as between 0 and 12 kHz, such as between 0 and 10 kHz, such as between 0 and 8 kHz, such as between 0 and 6 kHz, such as between 20 Hz and 20 kHz, such as between 20 Hz and 15 kHz, such as between 20 Hz and 12 kHz, such as between 20 Hz and 10 kHz, such as between 20 Hz and 8 kHz, such as between 20 Hz and 6 kHz.

[0138] 18. The computer-implemented method according to Item 16, wherein the low frequencies of the sound spectrum include those between 0 and 20 kHz, such as between 0 and 15 kHz, such as between 0 and 12 kHz, such as between 0 and 10 kHz, such as between 0 and 8 kHz, such as between 0 and 6 kHz, such as between 20 Hz and 20 kHz, such as between 20 Hz and 15 kHz, such as between 20 Hz and 12 kHz, such as between 20 Hz and 10 kHz, such as between 20 Hz and 8 kHz, such as between 20 Hz and 6 kHz, such as between 20 Hz and 5 kHz, such as between 20 Hz and 4 kHz, such as between 20 Hz and 3 kHz, such as between 20 Hz and 2 kHz, such as between 20 Hz and 1.5 kHz, such as between 20 Hz and 1 kHz.

[0139] 19. The computer-implemented method according to Item 16, wherein the high frequencies of the sound spectrum include those between 1 kHz and 20 kHz, such as between 1.5 kHz and 20 kHz, such as between 2 kHz and 20 kHz, such as between 3 kHz and 20 kHz, such as between 4 kHz and 20 kHz, such as between 5 kHz and 20 kHz, such as between 6 kHz and 20 kHz, such as between 8 kHz and 20 kHz, such as between 10 kHz and 20 kHz, such as between 12 kHz and 20 kHz, such as between 1 kHz and 15 kHz, such as between 1.5 kHz and 15 kHz, such as between 2 kHz and 15 kHz, such as between 3 kHz and 15 kHz. Between kHz, such as between 4kHz and 15kHz, such as between 5kHz and 15kHz, such as between 6kHz and 15kHz, such as between 8kHz and 15kHz, such as between 10kHz and 15kHz, such as between 12kHz and 15kHz, such as between 1kHz and 12kHz, such as between 1.5kHz and 12kHz, such as between 2kHz and 12kHz, such as between 3kHz and 12kHz, such as between 4kHz and 12kHz, such as between 5kHz and 12kHz, such as between 6kHz and 12kHz, such as between 8kHz and 12kHz, such as between 10kHz and 12kHz.

[0140] 20. The computer-implemented method according to any one of the preceding items, wherein a high-fidelity stereo is used to encode and decode the generated device-specific room impulse response.

[0141] 21. A computer-implemented method according to any one of the preceding items, wherein determining the energy content of at least one frequency of the associated transfer function of the first device includes determining the high-fidelity stereo order N of the energy content of the at least one frequency.

[0142] 22. The computer-implemented method according to Project 21, wherein a high-fidelity stereo order N is determined for multiple frequencies, wherein the energy content of each frequency is determined.

[0143] 23. The computer-implemented method according to item 21 or 22 includes determining the high-fidelity stereo order N of the energy content based on the sum of the high-fidelity stereo coefficients for each order N, and then normalizing it to one for each frequency.

[0144] 24. The computer-implemented method according to any one of items 21-23, wherein determining the number of digital representations of the receiver based on the energy content of at least one frequency of the device-related transfer function further comprises determining the number of digital representations of the receiver based on the high-fidelity stereo order N, wherein the number of digital representations of the receiver is (N+1). 2 1.5 (N+1) 2 or 2 (N+1) 2 .

[0145] 25. The computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer function comprises: - Obtain a device mesh model, which represents the geometry of the device and the position of at least the first microphone on the device mesh model. - Determine the first nearest grid element on the device grid model that is closest to at least the first microphone.

[0146] 26. The computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer function comprises: - Arrange the digital representations of a device receiver array, including multiple digital representations of the receiver, around the device grid model such that the distance between any digital representation of the receiver and the device grid model is not less than a predetermined distance.

[0147] 27. The computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer function comprises: - Use the first closest grid element as the sound source to digitally emit the first pulse signal.

[0148] 28. A computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer functions comprises: - A digital representation of a first source correction microphone positioned at a first source distance from the first nearest grid element, wherein the first source distance is less than the predetermined distance. - A wave-based solver is used to determine a first source correction signal, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone. -The plurality of first-source-corrected device impulse responses are determined by source-correcting each of the plurality of first device impulse responses using the first source correction signal. - The first device-related transfer function for the first microphone is generated by combining the device impulse responses of the plurality of first source corrections.

[0149] 29. The computer-implemented method according to any one of the preceding items, wherein generating the device-related transfer function comprises: - A wave-based solver is used to determine a plurality of first device impulse responses, wherein each first device impulse response describes the impulse response of the first pulse signal received at a corresponding receiver.

[0150] 30. The computer-implemented method according to any one of the preceding items, wherein generating the spatial room impulse response comprises, - Obtain a 3D room model representing the geometry of the room and at least one acoustic property.

[0151] 31. The computer-implemented method according to any one of the preceding items, wherein generating the space room impulse response comprises, - At least one digital representation of at least one room sound source is arranged in the 3D room model.

[0152] 32. The computer-implemented method according to any one of the preceding items, wherein generating the spatial room impulse response comprises, - A digital representation of a room receiver array comprising multiple digital representations of the receiver, wherein the room receiver array is centered on at least one listening point in the 3D room model, and wherein the number of digital representations of the receiver is determined based on the energy content of at least one frequency of the first device-related transfer function.

[0153] 33. The computer-implemented method according to any one of the preceding items, wherein generating the spatial room impulse response comprises, - Digitally transmit room pulse signals from the at least one audio source.

[0154] 34. The computer-implemented method according to any one of the foregoing items, wherein generating the spatial room impulse response comprises, - Use at least one wave-based frequency-based solver to determine multiple room impulse responses, wherein each room impulse response describes the transmitted room impulse received at a corresponding location in a plurality of digital representations of the receiver.

[0155] 35. The computer-implemented method according to any one of the preceding items, wherein generating the spatial room impulse response comprises, - Generate spatial room impulse responses based on the multiple room impulse responses.

[0156] 36. A computer-implemented method for recovering the device-dependent transfer function (DRTF) of a master device, wherein the DRTF of the master device includes at least one loss magnitude level, the method comprising the steps of: - Obtain multiple amplitude levels of the DRTF of the master device as a function of multiple elevation angles and multiple azimuth angles, wherein the DRTF of the master device includes at least one lost amplitude level; - Obtain the matrix including the spherical harmonic basis functions; - Applying the truncated singular value decomposition method to the matrix comprising spherical harmonic basis functions, wherein the truncated singular value decomposition method includes the following steps: - Based on the singular value decomposition (SVD), singular values ​​are generated on the matrix including the spherical harmonic basis functions; - Select the top singular value and the corresponding set of singular vectors, wherein the top singular value and the corresponding set of singular vectors are selected based on a predetermined threshold; - Based on the singular value decomposition (SVD), a secondary matrix including spherical harmonic basis functions is generated based on the top selection of the singular values ​​and the corresponding set of singular vectors; - Multiply the matrix corresponding to multiple amplitude levels of the DRTF of the main device with the inverse matrix of the secondary matrix to generate multiple high-fidelity stereo coefficients; - Generate multiple restored amplitude levels of the DRTF of the master device based on the multiple high-fidelity stereo coefficients, wherein the multiple restored amplitude levels of the DRTF of the master device include an image of the amplitude levels of the DRTF of the master device and a reconstruction of the at least one lost amplitude level.

[0157] 37. The computer-implemented method according to item 36, wherein the amplitude level of the DRTF of the main device is expressed as a function of elevation and azimuth.

[0158] 38. The computer-implemented method according to items 36-37, wherein the amplitude level of the DRTF of the main device is used for the DRTF frequency.

[0159] 39. The computer-implemented method according to item 38, wherein the DRTF frequency is between 20 Hz and 20 kHz.

[0160] 40. The computer-implemented method according to items 36-39, wherein the amplitude level is a complex value, or wherein the amplitude level is represented in the form of a complex number.

[0161] 41. The computer-implemented method according to items 36-40, wherein the DRTF of the main device is generated according to any one of items 1-35.

Claims

1. A computer-implemented method for generating a device-specific room impulse response (DSRIR) describing the acoustic characteristics of the device and the room received by the device, wherein, The device includes at least a first microphone, and the method includes: - Generate at least a first device-related transfer function (DRTF), wherein the at least first device-related transfer function describes the acoustic characteristics of the device as received by the at least first microphone, wherein generating the first device-related transfer function further includes: A device mesh model is obtained, the device mesh model representing the geometry of the device and the position of at least the first microphone on the device mesh model. The digital representations of an array of device receivers, including multiple digital representations of the device receivers, are arranged around the device mesh model such that the distance between any digital representation of the device receiver and the device mesh model is not less than a predetermined distance. On the device mesh model, determine the first nearest mesh element that is closest to at least the first microphone. A digital representation of a first source correction microphone positioned at a first source distance from the first nearest grid element, wherein the first source distance is less than a predetermined distance. The first pulse signal is digitally emitted using the first closest grid element as the sound source. A wave-based solver is used to determine a first source correction signal, wherein the first source correction signal describes the first pulse signal received at the first source correction microphone. A wave-based solver is used to determine a plurality of first device impulse responses, wherein each first device impulse response describes the impulse response of the first pulse signal received at a corresponding device receiver. The plurality of first-source-corrected device impulse responses are determined by performing source correction on each of the plurality of first device impulse responses using the first source correction signal. The first device-related transfer function for the first microphone is generated by combining the device impulse responses of the plurality of first source corrections. Determine the energy content at at least one frequency of the relevant transfer function of the first device. - Generating a spatial room impulse response (SRIR), wherein the spatial room impulse response describes the acoustic characteristics of the room received from at least one room sound source in the room and from at least one direction at at least one listening point in the room, wherein generating the spatial room impulse response further includes: Obtain a 3D room model representing the geometry of the room and at least one acoustic property. At least one digital representation of at least one room sound source is arranged in the 3D room model. A digital representation of a room receiver array comprising multiple digital representations of a room receiver, wherein the room receiver array is centered on at least one listening point in the 3D room model, and wherein the number of digital representations of the room receivers is determined based on the energy content of at least one frequency of the first device-related transfer function. The room pulse signal is digitally transmitted from the at least one audio source. At least one wave-based frequency-based solver is used to determine multiple room impulse responses, wherein each room impulse response describes a transmitted room impulse received at a corresponding location in a plurality of digital representations of the room receiver. A spatial room impulse response is generated based on the multiple room impulse responses. - The device-specific room impulse response (DSRIR) is generated by combining the device-related transfer function and the space room impulse response.

2. The computer-implemented method according to claim 1, wherein, The predetermined distance is between 0.5 and 1.5 meters, preferably between 0.8 and 1.2 meters, or most preferably 1 meter.

3. The computer-implemented method according to claim 1 or 2, wherein, The device includes a plurality of microphones, such as a second, third, fourth, and fifth microphone, wherein each of the plurality of microphones is regarded as the first microphone, such that a plurality of device-related transfer functions, such as the second, third, fourth, and fifth device-related transfer functions of the device, are generated for each microphone.

4. The computer-implemented method according to claim 1, 2, or 3, wherein, Determining the energy content at at least one frequency of the relevant transfer function of the first device includes determining different high-fidelity stereo orders to identify different levels of energy content.

5. The computer-implemented method according to claim 3 or 4, wherein, The energy content is determined for the frequency range, such as 0 to 20 kHz, such as 0 to 10 kHz, such as 10 to 20 kHz, such as 0 to 9 kHz, such as 0 to 8 kHz, such as 0 to 7 kHz, such as 0 to 6 kHz, such as 0 to 5 kHz, such as 0 to 4 kHz, such as 0 to 3 kHz, such as 0 to 2 kHz, such as 0 to 1 kHz.

6. The computer-implemented method according to any one of the preceding claims, wherein, Generating at least the first device-related transfer function (DRTF) includes: - Obtain a 3D box model including a highly sound-absorbing surface or a 3D box model with a predetermined size, such that the first pulse signal is received only once by the device's receiver array. - Arrange the device receiver array and the device mesh model in the 3D box model.

7. The computer-implemented method according to any one of the preceding claims, wherein, Generating at least a first device-related transfer function (DRTF) includes arranging an array of device receivers, comprising multiple digital representations of the device receivers, in a spherical or offset shape, wherein the digital representations of the device receivers are placed / arranged at a predetermined offset distance from the device mesh model.

8. The computer-implemented method according to any one of the preceding claims, wherein, The method also includes using a geometric acoustic solver with at least one geometric acoustic frequency to determine a second plurality of room impulse responses.

9. The computer-implemented method according to claim 8, wherein, The plurality of impulse responses generated using the wave-based solver and the second plurality of room impulse responses generated using the geometric acoustic solver are combined to generate a plurality of combined room impulse responses.

10. The computer-implemented method according to any one of claims 8 or 9, wherein, The plurality of room impulse responses generated using the wave-based solver are generated in the low frequencies of the acoustic spectrum, and the second plurality of impulse responses generated using the geometric acoustic solver are generated in the high frequencies of the acoustic spectrum.

11. The computer-implemented method according to any one of the preceding claims, wherein, The generated room-specific impulse response is encoded and decoded using high-fidelity stereo.

12. The computer-implemented method according to any one of the preceding claims, wherein, Determining the energy content of at least one frequency of the relevant transfer function of the first device includes determining the high-fidelity stereo order N of the energy content of the at least one frequency.

13. The computer-implemented method according to claim 12, wherein, The high-fidelity stereo order N is determined for multiple frequencies, where the energy content of each frequency is determined.

14. The computer-implemented method of claim 12 or 13, comprising determining the high-fidelity stereo order N of the energy content based on the sum of the high-fidelity stereo coefficients for each order N, and then normalizing it to one for each frequency.

15. The computer-implemented method according to any one of claims 12-14, wherein, Determining the number of digital representations of the room receiver based on the energy content of at least one frequency of the device's associated transfer function further includes determining the number of digital representations of the room receiver based on the high-fidelity stereo order N, wherein the number of digital representations of the room receiver is (N+1). 2 1.5 (N+1) 2 or 2 (N+1) 2 .