Room geometry inference method based on direct sound source and first-order reflection sound source positioning
Through the CA-SpatialNet network based on microphone array and cross attention layer, high-precision direct sound source and first-order reflected sound source positioning under unknown sound source location and room impulse response prior conditions are achieved, which solves the problem of large DOA and TDOA estimation errors in the prior art, and improves the accuracy of room geometric inference.
Patent Information
- Application Number
- CN202510556943.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art under unknown sound source location and room impulse response prior conditions, the DOA estimation accuracy of direct sound source and first-order reflected sound source is insufficient. TDOA estimation relies on explicit signal extraction to lead to large errors and low room geometry inference accuracy.
The room geometric inference method based on the positioning of direct sound sources and first-order reflective sound sources is adopted, and the reverberation signal is collected using a microphone array, and the time-frequency covariance matrix is calculated through the SpatialNet network for DOA estimation. Combined with the CA-SpatialNet network of the cross attention layer, the TDOA features are implicitly extracted, the sound source distance and position are estimated, and finally the position and normal vector of the reflection surface are estimated through the neural network.
It realizes higher accuracy direct sound source and first-order reflected sound source positioning, improves the accuracy of room geometric inference, avoids errors caused by explicit signal separation, and improves the accuracy of DOA and TDOA estimation.
Smart Images

Figure CN120539673A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning and room parameter estimation, and specifically relates to a room geometry parameter inference method and system based on implicit extraction of the direction of arrival (DOA) and time difference of arrival (TDOA) of a sound source and its primary reflection signal. Background Art
[0002] Traditional room geometry inference methods mainly rely on two types of methods. The first type of methods includes the following two:
[0003] 1. Methods based on room impulse response (RIR): This method requires the known sound source location or relies on the average RIR of multiple measurements, which limits its application in practical scenarios. 2. Methods based on microphone array signals, such as end-to-end neural network models based on high-order ambisonics (HOA), are only applicable to regular rooms (such as shoebox rooms). Furthermore, since the direct sound and reflected sound characteristics are not effectively separated, the computational burden is high and the accuracy is low.
[0004] The second category of methods uses deconvolutional neural networks (DCNNs) and time-domain convolutional networks (TCNs) to locate direct sound sources and first-order reflections, achieving room geometry inference. However, this second category of methods has the following drawbacks: 1. DOA estimation accuracy is insufficient. Traditional time-domain methods fail to fully exploit the frequency-domain correlation of signals, resulting in large DOA estimation errors and high missed detection rates. 2. TDOA estimation relies on explicit signal extraction, requiring beamforming to separate direct and reflected sound signals. This error is significant in reverberant environments.
[0005] Due to the above reasons or the shortcomings of existing methods, a method is needed that can achieve higher-precision positioning of direct sound sources and first-order reflected sound sources, thereby achieving more accurate room geometry inference. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to achieve high-precision DOA estimation for direct sound and first-order reflections, given unknown sound source locations and a priori RIRs. Furthermore, how to implicitly extract TDOA features to estimate the sound source distance, avoiding the errors introduced by explicit signal separation. To address this problem, the present invention provides a room geometry inference method based on direct and first-order reflection source localization.
[0007] The technical solution adopted in the present invention is:
[0008] A room geometry inference method based on direct sound source and first-order reflected sound source localization, the steps of which include:
[0009] 1) Use the microphone array to collect the reverberation signal of the target room and convert it into a FOA signal, calculate the time-frequency covariance matrix of the FOA signal and input it into the DOA estimation module; the DOA estimation module estimates the direction of arrival (DOA) of the direct sound source s based on the time-frequency covariance matrix. s and the direction of arrival DOA of n first-order reflected sound sources s′ s′ ; The FOA signal is a mixed first-order Ambisonics signal;
[0010] 2) Use the sound source distance estimation module to estimate the direction of arrival (DOA) of the direct sound source s s , the arrival direction DOA of the first-order reflected sound source s′ s′ and the height d of the microphone array h Estimate the distance d to the direct sound source s s ; Then according to the distance d to the direct sound source s s and the direction of arrival DOA of the direct sound source s s Determine the position P of the direct sound source s s ;
[0011] 3) Using the reflected sound source distance estimation module based on the FOA signal and the direction of arrival DOA of the direct sound source s s , the arrival direction DOA of the i-th first-order reflected sound source s′ s′ and the distance d to the direct sound source s s Estimate the i-th first-order reflected sound source s′
[0012] The distance d s′ ; Then according to the distance d of the i-th first-order reflected sound source s′ s′ and the arrival direction DOA of the i-th first-order reflected sound source s′ s′ Determine the position P of the i-th first-order reflected sound source s′ s′ ; i = 1 ~ n;
[0013] 4) According to the position P of the direct sound source s s and the position P of each first-order reflected sound source s′ s′ , and obtain the geometric space structure of the target room.
[0014] Furthermore, the self-attention layer in the narrowband block of the SpatialNet network is replaced with a cross-attention layer to obtain the sound source distance estimation module; the sound source distance estimation module is based on the arrival direction DOA of the direct sound source s. s , the arrival direction DOA of the first-order reflected sound source s′ from the bottom plate of the target room s′ and the height d of the microphone array h , estimate the distance d to the direct sound source s s .
[0015] Furthermore, the self-attention layer in the narrowband block of the SpatialNet network is replaced with a cross-attention layer to obtain the reflected sound source distance estimation module.
[0016] Furthermore, first according to the position P of the direct sound source s s and the position P of each first-order reflected sound source s′ s′ , calculate the position P of each reflective surface in the target room w and normal vector n w ; Then according to the position P of each reflecting surface w and normal vector n w Determine the geometric spatial structure of the target room.
[0017] Further,
[0018] The room geometry inference method based on direct sound source and first-order reflection sound source localization (LOCRGI) of the present invention includes the following modules:
[0019] 1. DOA estimation module
[0020] The network structure is SpatialNet, the structure is as follows Figure 2 As shown in the figure, SpatialNet is a well-known network that jointly models spatiotemporal correlations using cross-band blocks and narrow-band blocks. The network extracts inter-band correlations related to the sound source's DOA through cross-band blocks (consisting of a frequency-domain convolutional layer, a frequency-domain fully connected layer, and another frequency-domain convolutional layer), and extracts temporal information through narrow-band blocks (consisting of a time-domain self-attention layer and a time-domain convolutional layer). Based on this network, we effectively extract spatial information and sound source content information, thereby more accurately estimating the DOA of direct sound sources and first-order reflection sources.
[0021] The network input is a first-order ambisonics (FOA) signal. We first convert the signal to the time-frequency domain, calculate the covariance matrix features of the multi-channel signal at each time-frequency point, and input it into the SpatialNet network. The features after passing through L layers of cross-band blocks and L layers of narrowband blocks are averaged along the time and frequency dimensions. Then, through a fully connected layer, the output is the DOA estimation vector (unit 3D coordinates) for the direct sound and six first-reflection sound sources. The vector modulus is used to determine the presence of the sound source (threshold > 0.5).
[0022] 2. Direct sound source distance estimation module and first-order reflection sound source distance estimation module
[0023] This part includes an improved SpatialNet network based on the cross-attention mechanism (that is, the self-attention in the narrowband block of the original SpatialNet is replaced by the cross-attention layer. The improved SpatialNet network is called CA-SpatialNet). It implicitly extracts features related to the TDOA of the direct sound source and the first-order reflected sound source, avoiding explicit signal extraction. The CA-SpatialNet network structure is as follows Figure 3 The network input is the FOA signal and the DOA estimation result of the direct sound. s , the DOA estimation result of a first-order reflected sound DOA s′ , and a distance-related input d c We use CA-SpatialNet to first estimate the DOA based on the DOA of the direct sound and the first-order reflected sound from the floor. s 、DOA f , and the height d of the microphone array h , estimate the distance d of the sound source s , and then estimate the DOA based on the DOA of the direct sound and other first-order reflected sounds s 、DOA s′ , and the distance d of the direct sound s , estimate the distance d of the first-order reflected sound s′ .
[0024] After calculating the DOA and distance estimation of the direct sound source s and any first-order reflected sound source s′, the position coordinates P of the direct sound source can be directly calculated. s And the position coordinate P of the first-order reflected sound source s′ s’ .
[0025] CA-SpatialNet first passes the input FOA signal through a one-dimensional convolutional layer to obtain audio signal features. The DOA estimates of the direct sound source and first-order reflection sound source are then passed through two different fully connected layers with SiLU activation functions to obtain DOA features for the direct sound source and first-order reflection sound source. The audio signal features are multiplied by the two DOA features to obtain audio features in both directions. The audio signal then passes through two different cross-band blocks to extract inter-band features.
[0026] We replace the self-attention in the narrowband block of the original SpatialNet with a cross-attention layer, thereby transforming the narrowband block into a cross-attention block, which receives the audio features in two directions after passing through the cross-band block as input, implicitly extracting features related to the TDOA of the direct sound source and the first-order reflected sound source.
[0027] Furthermore, the microphone array height is fed into a fully connected layer to obtain the array height feature, which is then concatenated with the TDOA-related features and passed through a fully connected layer to obtain the sound source distance. Alternatively, the sound source distance is fed into a fully connected layer to obtain the distance feature, which is then concatenated with the TDOA-related features and passed through a fully connected layer to obtain the first-order reflection sound source distance.
[0028] 3. Room Boundary Inference Module
[0029] Since there are errors in the DOA of the direct sound source, the DOA of the first-order reflected sound source, and the distance estimation, it is unreasonable to directly use them to calculate the position of the reflecting surface. Therefore, we use a neural network to estimate the normal vector and position of the reflecting surface. Given the position P of the direct sound source and the first-order reflected sound source s ,P s′ , the position of the reflecting surface P w and normal vector n w Defined as.
[0030]
[0031] The network structure is as follows Figure 4 As shown in the figure, the network input is the position of the direct sound source and the first-order reflected sound source (three-dimensional coordinates calculated by DOA and distance). After the two positions are respectively passed through the fully connected layer to obtain features, they pass through a multi-layer fully connected neural network to output the position and normal vector of each wall (i.e., the reflecting surface).
[0032] Furthermore, we use the mean squared error (MSE) loss and the Adam optimizer to train the three networks.
[0033] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the above method.
[0034] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0035] The beneficial effects of the present invention are:
[0036] 1. Based on the SpatialNet network, spatial information is used to achieve effective DOA estimation.
[0037] 2. Based on the CA-SpatialNet network, features related to TDOA are implicitly extracted to estimate the distances of the direct sound source and the first-order reflection sound source, avoiding the explicit sound source signal estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1This is the structural diagram of this method.
[0039] Figure 2 This is the structure diagram of the DOA estimation network for direct sound sources and first-order reflected sound sources based on SpatialNet.
[0040] Figure 3 This is the network structure diagram for estimating the distance between direct sound sources and first-order reflected sound sources based on CA-SpatialNet.
[0041] Figure 4 This is the network structure diagram for estimating the position and normal vector of the reflecting surface.
[0042] Figure 5 is a schematic diagram of the size parameters of different shaped rooms in the experimental setup of the present invention;
[0043] (a) Trapezoidal room, (b) convex room, (c) concave room, (d) L-shaped room. DETAILED DESCRIPTION
[0044] The following describes a room geometry inference method based on direct sound source and first-order reflection sound source localization provided by the present invention in conjunction with the accompanying drawings:
[0045] Step 1: Signal preprocessing. Use an Eigenmike array to collect the reverberation signal of the target room and convert it into a FOA signal. The FOA signal is subjected to a short-time Fourier transform (STFT) to calculate the time-frequency covariance matrix in the time-frequency domain, which serves as the input to the DOA estimation module. The covariance matrix is calculated as:
[0046] R(t,f)B(t,f)B(t,f) H , (2)
[0047] in(·) H is the conjugate transpose, and B(t,f) is the time-frequency spectrum of the signal.
[0048] Step 2: DOA estimation: Input the covariance matrix into SpatialNet, which outputs the DOA vectors of the direct sound and the first reflected sound.
[0049] Step 3: Sound source distance estimation: The FOA signal, direct sound DOA estimate, first-order floor reflection DOA estimate, and microphone array height are input into CA-SpatialNet to obtain the distance estimate of the sound source and determine the location of the direct sound source.
[0050] Step 4: First-order reflection distance estimation: The FOA signal, direct sound DOA estimate, DOA estimates of other first-order reflections, and the distance estimate of the sound source are input into CA-SpatialNet to obtain the distance estimates of other first-order reflection sources, thereby determining the location of the first-order reflection source.
[0051] Step 5: Input the position of the direct sound source and the position of the first-order reflected sound source into the reflecting surface estimation network to obtain the position and normal vector estimation of the reflecting surface.
[0052] Method evaluation experiment
[0053] The evaluation of the method of the present invention is carried out through simulation experiments. The data set used in the experiment is the FSD18K data set, which contains 41 different types of sounds. The experiment is based on the mirror model and uses the pyroomacoustics library to simulate the transfer function. A total of five different types of rooms are simulated: shoebox room (Rectangular), trapezoidal room (Trapezoidal), convex room (Convex), concave room (Concave) and L-shaped room (L-shape). The parameters of the trapezoidal, convex, concave and L-shaped rooms are as follows: Figure 5 As shown. The room size is random, the size range is:
[0054] Rectangle: length 2.0m≤L<7.5m, width 2.0m≤W<7.5m;
[0055] Trapezoid: 2.0m≤y1<7.5m, 3.0m≤x2<8.0m, 0.5m≤x1 <x2;
[0056] Convex: 1.0m≤x1 <x2<x3<8.0m,|x3-x1|≥0.2m,1.0m≤y1<y2<8.0m;
[0057] Concave: The parameters are the same as the convex shape, but the geometric structure is concave;
[0058] L-type: 1.0m≤x1 <x2<8.0m,1.0m≤y1<y2<8.0m。
[0059] The sound source remains stationary and randomly distributed in the room, with a distance greater than 0.1m from the wall and greater than 0.3m from the microphone.
[0060] The baseline method adopted in the present invention is:
[0061] 1. DCNN+TD-CNN: DOA estimation and beamforming explicit signal separation method based on time domain deconvolution network (DCNN);
[0062] 2. CRNN end-to-end model: Direct room geometry inference method based on convolutional recurrent neural network (CRNN) (Reference.
[0063] The evaluation indicators used in this invention are:
[0064] 1. DOA estimation: Recall, Precision, and Angular Error;
[0065] 2. Sound source distance estimation: absolute distance error (Distance error), relative error (Rela-Error);
[0066] 3. Room geometry inference: wall position error, normal vector angle error.
[0067] Table 1. Comparison of sound source detection performance between the proposed method and the baseline method
[0068]
[0069] Table 1 shows the recall and precision of sound source detection for different methods. Experimental results demonstrate that, compared to the baseline method, the proposed method achieves accurate sound source detection, with both recall and precision approaching 100%. In contrast, the DCNN-based method exhibits poor recall, indicating that the baseline method misses a significant number of sound sources, which results in the baseline method missing many room reflective surfaces during boundary estimation.
[0070] Table 2. Comparison of DOA estimation performance between the proposed method and the baseline method
[0071]
[0072] Table 2 shows the DOA estimation errors of different methods for detected sound sources. It can be seen that the proposed method achieves a smaller angle estimation error than the baseline method. This is because SpatialNet effectively extracts spatial information by capturing frequency-domain correlations, while the temporal self-attention network better extracts contextual information from the source signal, thereby utilizing time periods that are beneficial for localization (e.g., periods of high source signal energy).
[0073] Table 3. Comparison of direct sound source distance estimation performance between the proposed method and the baseline method
[0074]
[0075] Table 4. Comparison of the first-order reflection sound source distance estimation performance of the proposed method and the baseline method
[0076]
[0077] Table 3 shows the distance error and relative distance error of sound source distance estimation using different methods. Table 4 shows the distance estimation error and relative error of the two methods for the first-order reflected sound source. It can be seen that based on the Ca-SpatialNet proposed in the present invention, the method achieves more accurate sound source distance estimation. Although the DCNN+TD+CNN method also realizes sound source distance estimation by combining the microphone height and TDOA estimation between the direct sound source and the first-order reflected sound source, its TDOA estimation method is based on signal extraction, and its beamformer is difficult to achieve a narrow enough beam. The difficulty of signal extraction greatly limits the performance of this method. In contrast, the method proposed in the present invention implicitly extracts features related to the signal TDOA through the proposed Ca-SpatialNet, without the need to extract the sound source signals of the direct sound source and the first-order reflected sound.
[0078] Table 5. Comparison of the reflective surface position estimation performance of the proposed method and the baseline method
[0079]
[0080] Table 6. Comparison of the reflection surface normal vector estimation performance of the proposed method and the baseline method
[0081]
[0082] Tables 5 and 6 show the room boundary positioning distance error and the angular error of the detected boundary normal vector estimation for different methods in various room scenarios. It can be seen that compared with the end-to-end room boundary prediction method, the solution based on direct sound source and first-order reflection sound source localization in the present invention achieves smaller errors in both the position and normal direction estimation of the room boundary, which verifies the theoretical basis for decomposing room boundary estimation into localization subtasks. Although the DCNN+TD+CNN method is also based on the localization of direct sound sources and first-order reflection sound sources, its performance significantly decreases in the more complex room shapes set by the present invention. This is due to its network structure design and the signal extraction-based TDOA estimation method, which even causes its performance to be lower than the end-to-end method. In comparison, the method of the present invention uses a more efficient DOA estimation network and implicit TDOA-related feature estimation based on Ca-SpatialNet (without explicitly extracting direct sound sources and first-order reflection signals), ultimately achieving more accurate first-order reflection distance estimation, thereby achieving better performance.
[0083] Although the specific embodiments and drawings of the present invention are disclosed for illustrative purposes, and are intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the preferred embodiments and drawings disclosed.
Claims
1. A room geometry inference method based on direct sound source and first-order reflection sound source localization, comprising the following steps: 1) Use the microphone array to collect the reverberation signal of the target room and convert it into a FOA signal, calculate the time-frequency covariance matrix of the FOA signal and input it into the DOA estimation module; the DOA estimation module estimates the direction of arrival (DOA) of the direct sound source s based on the time-frequency covariance matrix. s and the direction of arrival DOA of n first-order reflected sound sources s′ s′ ; The FOA signal is a mixed first-order Ambisonics signal; 2) Use the sound source distance estimation module to estimate the direction of arrival (DOA) of the direct sound source s s , the arrival direction DOA of the first-order reflected sound source s′ s′ and the height d of the microphone array h Estimate the distance d to the direct sound source s s ; Then according to the distance d to the direct sound source s s and the direction of arrival DOA of the direct sound source s s Determine the position P of the direct sound source s s ; 3) Using the reflected sound source distance estimation module based on the FOA signal and the direction of arrival DOA of the direct sound source s s , the arrival direction DOA of the i-th first-order reflected sound source s′ s′ and the distance d to the direct sound source s s Estimate the distance d of the i-th first-order reflected sound source s′ s′ ; Then according to the distance d of the i-th first-order reflected sound source s′ s′ and the arrival direction DOA of the i-th first-order reflected sound source s′ s′ Determine the position P of the i-th first-order reflected sound source s′ s′ ; i = 1 ~ n; 4) According to the position P of the direct sound source s s and the position P of each first-order reflected sound source s′ s′ , and obtain the geometric space structure of the target room.
2. The method according to claim 1, characterized in that The sound source distance estimation module is obtained by changing the self-attention layer in the narrowband block of the SpatialNet network to a cross-attention layer; the sound source distance estimation module is based on the arrival direction DOA of the direct sound source s. s , the arrival direction DOA of the first-order reflected sound source s′ from the bottom plate of the target room s′ and the height d of the microphone array h , estimate the distance d to the direct sound source s s .
3. The method according to claim 1, characterized in that The reflected sound source distance estimation module is obtained by replacing the self-attention layer in the narrowband block of the SpatialNet network with a cross-attention layer.
4. The method according to claim 1, 2 or 3, characterized in that: First, according to the position P of the direct sound source s s and the position P of each first-order reflected sound source s′ s′ , calculate the position P of each reflective surface in the target room w and normal vector n w ; Then according to the position P of each reflecting surface w and normal vector n w Determine the geometric spatial structure of the target room.
5. The method according to claim 1, 2 or 3, characterized in that:
6. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Positioning method and device for physical loudspeaker and virtual loudspeaker, equipment and storage medium
CN122063541A
A positioning method, device and equipment of a physical speaker and a virtual speaker and a storage medium
CN122063541B