Mobile sound source real-time tracking and positioning method based on space sound field reconstruction
By combining composite acoustic sensing networks, compressed sensing, and particle filtering algorithms with dynamic beamforming technology, the problems of positioning accuracy and real-time performance of mobile sound sources in open indoor environments are solved, achieving efficient sound source tracking and positioning and sound image reconstruction, and improving speech intelligibility and sound field immersion.
Patent Information
- Application Number
- CN202511177495.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-09
AI Technical Summary
Existing sound source localization methods struggle to achieve continuous localization and sound image reconstruction of moving sound sources in open indoor environments, resulting in low localization accuracy and poor real-time performance, which affects speech intelligibility and sound field immersion.
A composite acoustic sensing network combined with a time difference algorithm is used for initial screening and localization. Compressed sensing theory and particle filtering algorithm are used for trajectory smoothing. A dynamic beamforming network is designed and spatial sound field is reconstructed using a ray tracing algorithm to achieve real-time tracking and localization of the sound source.
It achieves efficient mobile sound source tracking and positioning in open indoor environments, with positioning delay ≤15ms, spatial sound image positioning error <3°, and beam pointing accuracy of ±5°, improving speech intelligibility and sound field immersion, and enhancing the system's adaptability to complex environments.
Smart Images

Figure CN121091202A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of acoustic positioning, more particularly, the present application relates to a mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction. BACKGROUND
[0002] Spatial sound field reconstruction refers to the process of reconstructing or simulating the sound field in a specific space through a series of technical means. In conference simultaneous interpretation, immersive performance and other scenarios, there is a high demand for real-time tracking and positioning of mobile sound sources. Traditional sound source tracking systems mostly rely on fixed arrays, and in open indoor environments, it is difficult to achieve continuous positioning and sound image reconstruction of mobile sound sources, resulting in spatial inconsistency between the perceived sound source position of the audience and the actual sound source, affecting the intelligibility of speech and the sense of immersion of the sound field.
[0003] The existing technology has the following deficiencies:
[0004] The existing sound source positioning method has deficiencies in positioning accuracy, real-time performance and adaptability to the environment. Some methods have high positioning delay and cannot meet the real-time tracking requirements; some methods perform poorly in spatial sound image positioning error and beam pointing accuracy, making it difficult to provide a good immersive experience. Therefore, a new mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction is needed to solve the above problems.
[0005] In view of the above problems, the present application provides a solution. SUMMARY
[0006] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction, which provides a mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction to solve the problems raised in the background art.
[0007] To achieve the above purpose, the present application provides the following technical scheme:
[0008] A mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction, comprising the following steps:
[0009] Deploy a composite acoustic perception network in the perception layer, and use a time difference algorithm to preliminarily position the sound source coordinates and collect environmental acoustic fingerprint data;
[0010] Construct a spatial spectrum estimation model based on compressed sensing theory in the analysis layer, and use a particle filtering algorithm to smooth the preliminary screening coordinates to generate a continuous sound source motion vector;
[0011] In the rendering layer, a dynamic beamforming network is designed, the phase and amplitude weights of the microphone channels are adjusted in real time according to the motion vector of the sound source, and a directional sound beam is formed through the virtual array technology.
[0012] The ray tracing algorithm is used for online calibration of the indoor reflecting surface to generate the spatial sound field impulse response compensation parameters, and accurate reconstruction of the spatial sound field is realized.
[0013] In a preferred embodiment, the composite acoustic perception network is composed of a microphone array and ultrasonic sensors; the microphone array is arranged in a uniform circular distribution, and the ultrasonic sensors are uniformly distributed around the microphone array.
[0014] In a preferred embodiment, the initial screening positioning process of the sound source coordinates is as follows through the time difference algorithm:
[0015] The microphone array receives audible sound wave signals, and records the time when each microphone receives the signal The ultrasonic sensor receives ultrasonic signals, and records the time when each sensor receives the signal and obtains the time t0 when the sound source starts to sound;
[0016] According to the time when each microphone receives the signal Combined with the time t0 when the sound source starts to sound and the propagation speed of audible sound in air, the distance from the sound source to the i-th microphone is calculated;
[0017] According to the time when each sensor receives the signal Combined with the time t0 when the sound source starts to sound, the interval of the signals received by the sensor is obtained, and the distance from the sound source to the j-th ultrasonic sensor is obtained by multiplying the propagation speed of ultrasonic waves;
[0018] The distance from the sound source to the i-th microphone and the distance from the sound source to the j-th ultrasonic sensor are combined, the time t0 when the sound source starts to sound is eliminated, and the relationship between the time difference and the distance of the signals received by the microphone array and the ultrasonic sensor is obtained;
[0019] Let the sound source coordinates be S=(x, y, z), combine the coordinates of the microphone and the ultrasonic sensor, and calculate the distance from the sound source to the microphone and the ultrasonic sensor according to the distance formula;
[0020] The relationship between the time difference and the distance, and the distance from the sound source to the microphone and the ultrasonic sensor calculated according to the distance formula are combined into an equation group, the equation group is solved, and the least square method is used to solve the initial screening coordinates of the sound source
[0021] In a preferred embodiment, the relationship between time difference and distance is combined into an equation group according to the distance formula of the sound source to the microphone and the ultrasonic sensor distance calculated according to the distance formula as follows:
[0022]
[0023] wherein (x, y, z) is the sound source coordinate, is the coordinate of the jth ultrasonic sensor, is the coordinate of the ith microphone, is the distance from the sound source to the ith microphone, is the distance from the sound source to the jth ultrasonic sensor, c a is the propagation speed of sound in air, c u is the propagation speed of ultrasonic wave in air, is the time when the ith microphone receives the signal, is the time when the jth ultrasonic sensor receives the ultrasonic signal.
[0024] In a preferred embodiment, the environmental acoustic fingerprint data includes environmental noise power spectral density, and the acquisition process is as follows:
[0025] The environmental noise signal is collected and integrated with the sampling duration to obtain the environmental noise power spectral density, and the calculation formula is as follows:
[0026]
[0027] wherein P(f) is the environmental noise power spectral density; T is the sampling duration; N(t, f) is the signal value of the environmental noise at time t and frequency f.
[0028] In a preferred embodiment, the analysis layer constructs a spatial spectrum estimation model based on compressed sensing theory, including regarding the spatial distribution of the sound source as a sparse signal, estimating the spatial distribution of the sound source by processing and reconstructing the collected signal, and the process is as follows:
[0029] The monitoring space is divided into a plurality of grid points, and the grid point coordinates are set as G k = (k x Δx, k y Δy, k z Δz), wherein Δx, Δy, Δz are grid spacings, k x , k y , k z is the grid index;
[0030] A sparse sensing matrix Φ ∈ R 32×K is constructed, wherein K is the total number of grid points, and the matrix element is the distance from the ith microphone to the kth grid point, and λ is the wavelength of the sound wave;
[0031] The sparse vector α is solved by minimizing the l1 norm, i.e., min ||α||1, subject to the constraint ||Φα-s||2≤∈, where s is the microphone array received signal vector, and ∈ is the noise threshold;
[0032] The grid point corresponding to the maximum value in α is taken as the candidate position of the sound source at this moment, and the spatial distribution of the sound source is obtained.
[0033] In a preferred embodiment, the combined particle filtering algorithm performs trajectory smoothing processing on the preliminary screening coordinates to generate a continuous sound source motion vector process as follows:
[0034] First, the particles are initialized, and the initial particle set is Each particle is subject to a Gaussian distribution with a mean value of and a covariance matrix Σ, i.e.,
[0035] According to the sound source motion model, the next state ξ n (t) of the particle is predicted, and the formula is ξ n (t) = ξ n (t-1) + v(t-1)Δt + M n (t), where v(t-1) is the sound source velocity at time t-1, Δt is the time interval, M n (t) is the process noise, ξ n (t-1) is the particle state at time t-1;
[0036] The particle weight w n (t) is updated according to the signal likelihood, where w n (t) ∝ p(z(t)|ξ n (t)), z(t) is the observation value at time t, and p(z(t)|ξ n (t)) is the likelihood function;
[0037] The particles are normalized, and the formula is The particle set is regenerated
[0038] The sound source motion vector at time t is calculated to generate a continuous sound source motion vector, and the formula is:
[0039]
[0040] In the formula, is the sound source motion vector, and ξ′ n(t) is the particle set regenerated at time t, ξ' n (t-1) is the particle set regenerated at time t-1.
[0041] In a preferred embodiment, the process of adjusting the phase and amplitude weight of the microphone channel in real time according to the sound source motion vector is as follows:
[0042] According to the sound source motion vector obtained by the analysis layer Determine the azimuth angle θ(t) of the sound source at time t;
[0043] According to the azimuth angle θ(t) of the sound source at time t, combined with the sound wave wavelength λ and the angle position θ i of the i-th microphone in the circular array, the phase weight φ i (t) of the i-th microphone channel is obtained, the formula is φ i (t) = 2πrcos(θ(t)-θ i ) / λ, and the amplitude weight a i (t) = sinc(πNcos(θ(t)-θ i ) / 2), where N is the number of microphones, and r is the radius of the microphone deployment circle.
[0044] In a preferred embodiment, the process of online calibration of indoor reflective surfaces using ray tracing algorithm is as follows:
[0045] According to the ray tracing algorithm, simulate the propagation path of sound waves from the sound source to the microphone, when setting the initial parameters, input the coordinates of the sound source and the microphone, and preset the initial equation of the indoor reflective surface;
[0046] And emit rays covering a certain range at 1° intervals with the sound source as the vertex, respectively simulate 0 reflection path (direct sound), 1 reflection path, 2 and 3 reflection paths;
[0047] For indoor reflective surfaces, calculate the reflection points of sound waves on each reflective surface according to the simulated propagation path, and optimize the equation parameters of the reflective surface by the least square method to realize online calibration of the reflective surface.
[0048] In a preferred embodiment, the process of generating spatial sound field impulse response compensation parameters to realize accurate reconstruction of the spatial sound field is as follows:
[0049] According to the calibrated reflective surface type, query the sound reflection coefficient of the corresponding material, and calculate the attenuation coefficient according to the sound reflection coefficient and the simulation times of the reflection path;
[0050] Generate an impulse response for each microphone, construct a discrete impulse response sequence and convert it to a frequency response to obtain the spatial sound field impulse response compensation parameter where α kis the attenuation coefficient of the kth reflection, t k is the time of the kth reflected sound wave to reach the microphone, and δ(·) is the Dirac function;
[0051] According to the spatial sound field impulse response compensation parameter, the real-time signal is filtered, and the spatial sound field is synthesized in combination with the virtual array technology to realize the spatial consistency reconstruction of the sound source position and the sound image.
[0052] The technical effects and advantages of the mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction are as follows:
[0053] 1. The present application constructs a layered functional architecture through the technical fusion of multi-modal perception and dynamic sound field rendering, realizes the whole-process closed-loop control from sound source signal acquisition, coordinate preliminary screening to motion trajectory optimization and sound field accurate reconstruction. The composite acoustic perception network of the perception layer combines the time difference algorithm to provide reliable initial coordinates for positioning; the compressed sensing spatial spectrum estimation and particle filtering algorithm of the analysis layer significantly improve the continuity and accuracy of the mobile sound source trajectory tracking; the dynamic beamforming and ray tracing technology of the rendering layer realizes the spatial consistency of the sound image through the sound field impulse response compensation, solves the problem of deviation between the perceived position and the actual position of the sound source in the traditional method, and greatly improves the speech intelligibility and sound field immersion.
[0054] 2. The present application realizes the efficient tracking and positioning of the mobile sound source in the open indoor environment through the sound field processing architecture. The sound source positioning delay is ≤15ms, the spatial sound image positioning error is <3°, and the beam pointing accuracy reaches ±5°, which can meet the high requirements of real-time and accuracy in conference simultaneous transmission, immersive performance and other scenes. At the same time, the dynamic updating mechanism of the environmental acoustic fingerprint and the online calibration function of the reflection surface enhance the adaptability of the system to complex environments, avoid the positioning drift caused by fixed scene parameters, and provide a general and efficient solution for mobile sound source tracking in different indoor scenes. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 It is a structure schematic diagram of the mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction. DETAILED DESCRIPTION
[0056] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0057] Embodiment 1, Figure 1The application discloses a mobile sound source real-time tracking and positioning method based on spatial sound field reconstruction.
[0058] A composite acoustic perception network is deployed in the perception layer, time difference algorithm is used to preliminarily screen and position the sound source coordinates and collect environmental acoustic fingerprint data;
[0059] The perception layer deploys a composite acoustic perception network composed of an array of 32 microphones and 8 ultrasonic sensors; the 32 microphones are deployed in a uniform circular distribution, with a diameter of 50 cm, i.e. a radius r = 25 cm, and the coordinates of the i-th microphone are M i =(rcosθ i ,rsinθ i ,0), wherein i = 1, 2,..., 32;
[0060] The 8 ultrasonic sensors are uniformly distributed around the microphone array, and the coordinates of the j-th ultrasonic sensor are U j =((r+d)cosφ j ,(r+d)sinφ j ,0), wherein d is the distance from the edge of the microphone array to the ultrasonic sensor,
[0061] The time difference between the reception of the mobile sound source signal by the microphone array and the ultrasonic sensor is calculated by the time difference algorithm, the propagation speed of audible sound in air is about 343 m / s, the propagation speed of ultrasonic waves is about 331.5 m / s, the distance and the direction of the sound source to the perception network are calculated according to the time difference and the propagation speed, the preliminary screening and positioning of the sound source coordinates are realized, and the process is as follows:
[0062] The microphone array receives audible sound wave signals, and records the time when each microphone receives the signal The ultrasonic sensor receives ultrasonic wave signals, and records the time when each sensor receives the signal and the time t0 when the sound source starts to sound is obtained;
[0063] The time when each microphone receives the signal is combined with the time t0 when the sound source starts to sound, and the propagation speed of audible sound in air is used to calculate the distance from the sound source to the i-th microphone;
[0064] The time when each sensor receives the signal is combined with the time t0 when the sound source starts to sound, and the interval between the reception of the signals by the sensors is obtained, and the distance from the sound source to the j-th ultrasonic sensor is obtained by multiplying the interval by the propagation speed of ultrasonic waves;
[0065] The distance from the sound source to the i-th microphone and the distance from the sound source to the j-th ultrasonic sensor are associated, and the time t0 when the sound source starts to sound is eliminated to obtain the relationship between the time difference and the distance of the microphone array and the ultrasonic sensor receiving the moving sound source signal;
[0066] Suppose the sound source coordinates are S=(x, y, z), and the coordinates of the microphone and ultrasonic sensor are combined to calculate the distance from the sound source to the microphone and ultrasonic sensor according to the distance formula;
[0067] The relationship between the time difference and the distance of the microphone array and the ultrasonic sensor receiving the moving sound source signal, and the distance from the sound source to the microphone and ultrasonic sensor calculated according to the distance formula are combined into an equation group, and the initial screening coordinates of the sound source are obtained by solving the equation group and using the least square method The equation group is as follows:
[0068]
[0069] In the formula, (x, y, z) is the sound source coordinates, is the coordinates of the j-th ultrasonic sensor, is the coordinates of the i-th microphone, is the distance from the sound source to the i-th microphone, is the distance from the sound source to the j-th ultrasonic sensor, c a is the propagation speed of sound in air, c u is the propagation speed of ultrasonic waves in air, is the time when the i-th microphone receives the signal, is the time when the j-th ultrasonic sensor receives the ultrasonic signal.
[0070] The environmental acoustic fingerprint data includes environmental noise power spectral density, and the acquisition process is as follows:
[0071] First, the environmental noise signal is collected, and the environmental noise power spectral density is calculated by integrating the sampling duration, and the calculation formula is as follows:
[0072]
[0073] Where P(f) is the environmental noise power spectral density; T is the sampling duration, which is 0.1s here; N(t, f) is the signal value of the environmental noise at time t and frequency f.
[0074] It should be noted that the perception layer collects the sound source signal and the environmental acoustic fingerprint by using the composite acoustic perception network, and obtains the initial screening coordinates of the sound source by using the time difference algorithm, thereby providing initial position information for sound field reconstruction.
[0075] A spatial spectrum estimation model based on compressed sensing theory is constructed in the analysis layer, and the particle filtering algorithm is combined to smooth the trajectory of the preliminary screening coordinates to generate continuous sound source motion vectors.
[0076] A spatial spectrum estimation model based on compressed sensing theory is constructed in the analysis layer, and the particle filtering algorithm is combined to smooth the trajectory of the preliminary screening coordinates to generate continuous sound source motion vectors.
[0077] The monitoring space is divided into a plurality of grid points, and the grid point coordinates are G k =(k x Δx,k y Δy,k z Δz), wherein Δx, Δy, Δz are grid spacings, which are all taken as 0.1 m, k x ,k y ,k z is the grid index;
[0078] A sparse sensing matrix Φ ∈ R 32×K is constructed, wherein K is the total number of grid points, and the matrix element is the distance from the i-th microphone to the k-th grid point, and λ is the wavelength of audible sound;
[0079] The sparse vector α is solved by minimizing the l1 norm, that is, min||α||1, and the constraint condition is ||Φα-s||2≤∈, wherein s is the microphone array received signal vector, and ∈ is the noise threshold;
[0080] The grid point corresponding to the maximum value in α is taken as the candidate position of the sound source at that moment.
[0081] 1000 particles are set to sample the preliminary screening coordinates, the particle weights are updated according to the signal likelihood, the particles with high weights are resampled, the trajectory smoothing processing is realized, and continuous sound source motion vectors are generated, and the process is as follows:
[0082] First, the particles are initialized, 1000 particles are set, and the initial particle set is Each particle obeys a Gaussian distribution with the preliminary screening coordinates as the mean value, and the covariance matrix Σ as the covariance matrix, that is, The covariance matrix Σ is constructed by calculating the variance and covariance of the coordinate deviation through fixing the real coordinates of the sound source and collecting a plurality of preliminary screening coordinate samples;
[0083] According to the sound source motion model, the next state ξ n (t) of the particle is predicted as n (t-1)+v(t-1)Δt+M n(t), where v(t-1) is the sound source velocity at time t-1; At is the time interval; M n (t) is the process noise, which is subject to Gaussian distribution with mean 0 and covariance Q; ξ n (t-1) is the particle state at time t-1;
[0084] The particle weight w is updated according to the likelihood of the signal n (t), where w n (t)∝p(z(t)|ξ n (t)), z(t) is the observation value at time t (i.e. the candidate position obtained by spatial spectrum estimation), p(z(t)|ξ n (t)) is the likelihood function, and a Gaussian likelihood function p(z(t)|ξ n (t)) = N(z(t); ξ n (t), R) is used here, where R is the observation noise covariance matrix;
[0085] The particles are normalized, and the formula is Particles with higher weights are retained, and a new particle set is generated
[0086] The sound source motion vector at time t is calculated, and a continuous sound source motion vector is generated, and the formula is:
[0087]
[0088] In the formula, is the sound source motion vector, ξ′ n (t) is the newly generated particle set at time t, ξ′ n (t-1) is the newly generated particle set at time t-1.
[0089] It should be noted that the analysis layer constructs a spatial spectrum estimation model based on the compressed sensing theory, optimizes the preliminary screening coordinates in combination with the particle filtering algorithm, generates a continuous sound source motion vector, and further determines the motion trajectory of the sound source in space, thereby providing a dynamic position basis for sound field reconstruction.
[0090] The rendering layer designs a dynamic beamforming network, adjusts the phase and amplitude weights of the microphone channels in real time according to the sound source motion vector, and forms a directional sound beam through virtual array technology;
[0091] The rendering layer designs a dynamic beamforming network, adjusts the phase and amplitude weights of the 32 microphone channels in real time according to the sound source motion vector obtained by the analysis layer, expands the effective aperture of the microphone array to 2m through virtual array technology, forms a directional sound beam, and the process is as follows:
[0092] According to the sound source motion vector obtained by the analysis layer Determine the azimuth angle θ(t) of the sound source at time t; according to the azimuth angle θ(t) of the sound source at time t, combine the sound wave wavelength and the angle position of the i th microphone in the circular array, and obtain the phase weight φ of the i th microphone channel i (t), the formula is φ i (t) = 2πrcos(θ(t)-θ i ) / λ, the amplitude weight is a i (t) =sinc(πNcos(θ(t)-θ i ) / 2), where N is
[0093] The number of microphones, here N = 32;
[0094] The effective aperture of the microphone array is expanded by the virtual array technology, forming a directional sound beam with an effective aperture of 2m.
[0095] And use ray tracing algorithm for online calibration of indoor reflecting surface, generate spatial sound field impulse response compensation parameters, realize accurate reconstruction of spatial sound field.
[0096] Using the ray tracing algorithm, set the reflection number of sound waves to 3 times, simulate the propagation path of sound waves in the room, and perform online calibration of the reflecting surfaces such as walls, ceilings, and floors in the room to generate spatial sound field impulse response compensation parameters, the process is as follows:
[0097] According to the ray tracing algorithm, simulate the propagation path of sound waves from the sound source to the microphone, when setting the initial parameters, input the coordinates of the sound source and the microphone, preset the indoor reflecting surface as 6 (front wall, back wall, left wall, right wall, ceiling, floor), the initial equation is:
[0098] Floor: z = 0, ceiling: z = H (H is the room height, the initial value is set to 3m);
[0099] Front wall: x = 0, back wall: x = L x (L x is the room length, the initial value is set to 10m);
[0100] Left wall: y = 0, right wall: y = L y (L y is the room width, the initial value is set to 8m);
[0101] And with the sound source as the vertex by 1° interval emission covering a certain range of rays, respectively, 0 times of reflection path (direct sound) simulation, 1 time of reflection path simulation, 2 times and 3 times of reflection path simulation; Among them, 0 times of reflection path, namely direct sound simulation, to calculate the ray direction vector, path length and arrival time, 1 time of reflection path simulation needs to calculate the intersection of the ray and the reflection surface, verify the reflection law and filter the effective path, 2 times and 3 times of reflection path simulation is to iteratively calculate the reflection point, to 3 times of reflection cutoff, and remove the path;
[0102] For the reflection surface of the indoor wall, ceiling, floor and the like, according to the simulation propagation path, the reflection point of the sound wave on each reflection surface is calculated, and the reflection surface equation parameters are optimized by the least square method to realize the online calibration of the reflection surface, the process is as follows: first, collect the reflection point coordinates, obtain the reflection sound arrival time through the microphone array, and back calculate the reflection point coordinates according to the simulation path and the measured time; Set the reflection surface equation as Ax+By+Cz+D=0 (the unit normal vector (A, B, C) satisfies A 2 +B 2 +C 2 =1), D is a constant term, for the m reflection points on the surface The error function is The partial derivative of J(A, B, C, D) is taken and set to 0 to obtain a linear equation group, and the optimized reflection surface parameters are obtained by solving the linear equation group by the least square method, and iteratively updated until the parameter change is less than 10 -4 ;
[0103] According to the calibrated reflection surface type (such as concrete wall, glass, wooden ceiling), the sound reflection coefficient of the corresponding material is queried, and the attenuation coefficient is calculated according to the sound reflection coefficient (0 times of reflection a0=1, 1 times of reflection a1=R1, 2 times of reflection a2=R1R2, 3 times of reflection a3=R1R2R3); Generate an impulse response for each microphone, construct a discrete impulse response sequence and convert it to a frequency response to obtain a spatial sound field impulse response compensation parameter Where a k is the attenuation coefficient of the kth reflection, t k is the time of the kth reflected sound wave reaching the microphone, and δ(·) is the Dirac function.
[0104] According to the spatial sound field impulse response compensation parameter, the audio signal is compensated and the sound field is reconstructed, the real-time signal is filtered first, then combined with the virtual array technology to synthesize the spatial sound field, and the spatial consistency reconstruction of the sound source position and the sound image is realized.
[0105] It should be noted that the rendering layer adjusts the beamforming parameters according to the sound source motion vector to form a directional sound beam, and at the same time calibrates the reflection surface through a ray tracing algorithm and generates an impulse response compensation parameter to realize accurate reconstruction of the spatial sound field. Through the cooperative work of each layer, the sound source position information is constantly updated, thereby realizing real-time tracking and positioning of the moving sound source, so that the sound source position perceived by the listener is consistent with the actual sound source in space. Each functional module builds a closed-loop control system through a data bus, wherein the sound source positioning delay is less than or equal to 15 ms, the spatial sound image positioning error is less than 3°, and the beam pointing accuracy is ± 5°.
[0106] The above formulas are all dimensionless numerical calculations, and the formulas are obtained by software simulation of a large amount of data to obtain a formula closest to the actual situation. The preset parameters in the formula are set by a person skilled in the art according to the actual situation.
[0107] The above embodiments can be realized wholly or partially by software, hardware, firmware or any combination thereof. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product.
[0108] Those skilled in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. A person skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0109] In addition, the functional modules in each embodiment of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0110] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0111] Finally, the above is only a preferred embodiment of the present application and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction, characterized in that, Includes the following steps: A composite acoustic sensing network is deployed in the sensing layer to perform initial screening and localization of sound source coordinates and collect environmental acoustic fingerprint data using a time difference algorithm. In the analysis layer, a spatial spectrum estimation model based on compressed sensing theory is constructed, and the particle filtering algorithm is used to smooth the trajectory of the initial screening coordinates to generate continuous sound source motion vectors. A dynamic beamforming network is designed in the rendering layer to adjust the phase and amplitude weights of the microphone channel in real time according to the sound source motion vector, and to form a directional sound beam through virtual array technology. The ray tracing algorithm is used to calibrate the indoor reflective surface online, generate spatial sound field impulse response compensation parameters, and realize accurate reconstruction of the spatial sound field.
2. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 1, characterized in that, The composite acoustic sensing network consists of a microphone array and an ultrasonic sensor; the microphone array is deployed in a uniform circular distribution, and the ultrasonic sensor is uniformly distributed around the microphone array.
3. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 2, characterized in that, The initial screening and localization process of sound source coordinates using the time difference algorithm is as follows: The microphone array receives audible sound wave signals and records the time it takes for each microphone to receive the signal. The ultrasonic sensor receives ultrasonic signals and records the time when each sensor receives the signal. And obtain the time t0 when the sound source starts emitting sound; Based on the time each microphone receives the signal The distance from the sound source to the i-th microphone is calculated by combining the time t0 when the sound source starts emitting sound with the speed of audible sound waves in the air. The time when each sensor receives a signal The interval between the sensor receiving signals is obtained by taking the time t0 when the sound source starts emitting sound, and multiplying it by the speed of ultrasonic wave propagation to obtain the distance from the sound source to the j-th ultrasonic sensor. By combining the distance from the sound source to the i-th microphone and the distance from the sound source to the j-th ultrasonic sensor, and eliminating the time t0 when the sound source starts to emit sound, we can obtain the relationship between the time difference between the microphone array and the ultrasonic sensor receiving the moving sound source signal and the distance. By combining the coordinates of the microphone and the ultrasonic sensor, the distance from the sound source to the microphone and the ultrasonic sensor is calculated according to the distance formula. The relationship between time difference and distance, the distance from the sound source to the microphone and the ultrasonic sensor calculated according to the distance formula are combined into a system of equations. By solving the system of equations and using the least squares method, the initial screening coordinates of the sound source are obtained.
4. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 3, characterized in that, The relationship between time difference and distance, calculated using the distance formula, is combined into a system of equations as follows: In the formula, (x,y,z) are the coordinates of the sound source. These are the coordinates of the j-th ultrasonic sensor. Let be the coordinates of the i-th microphone. Let be the distance from the sound source to the i-th microphone. Let c be the distance from the sound source to the j-th ultrasonic sensor. a Let c be the speed at which sound waves travel in air. u This represents the speed at which ultrasound travels through the air. Let i be the time when the i-th microphone receives the signal. Let be the time when the j-th ultrasonic sensor receives the ultrasonic signal.
5. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 4, characterized in that, The environmental acoustic fingerprint data includes the environmental noise power spectral density, and the acquisition process is as follows: The environmental noise signal is collected and integrated with the sampling time to obtain the environmental noise power spectral density.
6. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 5, characterized in that, The analysis layer constructs a spatial spectrum estimation model based on compressed sensing theory, which includes treating the spatial distribution of sound sources as sparse signals, and estimating the spatial distribution of sound sources by processing and reconstructing the acquired signals. The process is as follows: The monitoring space is divided into several grid points, and a sparse sensing matrix is constructed. The sparse vector α is solved by minimizing the l1 norm. The grid point corresponding to the maximum value of α is taken as the candidate position of the sound source at that moment, and thus the spatial distribution of the sound source is obtained.
7. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 6, characterized in that, The process of combining the particle filtering algorithm to smooth the trajectory of the initial screening coordinates and generate continuous sound source motion vectors is as follows: First, the particles are initialized, and each particle follows a Gaussian distribution with the initial screening coordinates as the mean and the covariance matrix as Σ. Based on the sound source motion model, predict the next state of the particle and update the particle weights according to the likelihood of the signal. The particles are normalized, and the particle set is regenerated. Calculate the sound source motion vector at time t to generate continuous sound source motion vectors.
8. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 7, characterized in that, The process of adjusting the phase and amplitude weights of the microphone channel in real time based on the sound source motion vector is as follows: Based on the sound source motion vector obtained from the analysis layer, determine the azimuth angle of the sound source at time t; Based on the azimuth angle of the sound source at time t, combined with the wavelength of the sound wave and the angular position of the i-th microphone in the circular array, the phase weight and amplitude weight of the i-th microphone channel are obtained.
9. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 8, characterized in that, The process of online calibration of indoor reflective surfaces using the ray tracing algorithm is as follows: The propagation path of sound waves from the sound source to the microphone is simulated using the ray tracing algorithm. When setting the initial parameters, the coordinates of the sound source and the microphone need to be entered, and the initial equation of the indoor reflective surface needs to be preset. Rays covering a certain range were emitted at 1° intervals with the sound source as the vertex, and reflection path simulations were performed for 0, 1, 2 and 3 times respectively. For indoor reflective surfaces, the reflection points of sound waves on each reflective surface are calculated based on the simulated propagation path. The parameters of the reflective surface equation are optimized using the least squares method to achieve online calibration of the reflective surface.
10. The method for real-time tracking and localization of a moving sound source based on spatial sound field reconstruction according to claim 9, characterized in that, The process of generating spatial sound field impulse response compensation parameters to achieve accurate reconstruction of the spatial sound field is as follows: Query the acoustic reflection coefficient of the corresponding material based on the calibrated reflective surface type, and calculate the attenuation coefficient based on the acoustic reflection coefficient and the number of simulations of the reflection path; An impulse response is generated for each microphone, a discrete impulse response sequence is constructed and converted into a frequency response, and the spatial sound field impulse response compensation parameters are obtained. By filtering the real-time signal based on the spatial sound field impulse response compensation parameters and then combining it with virtual array technology to synthesize the spatial sound field, the spatial consistency reconstruction of the sound source location and sound image is achieved.