Sound environment imaging method based on sound array
By acquiring sound wave signals through a distributed audio array, constructing a spatial topology model and performing phase compensation, reconstructing the three-dimensional density field of sound wave energy, and generating a multimodal imaging spectrum, the problem of information loss and single-dimensional imaging in existing sound wave imaging technologies is solved, and comprehensive and accurate imaging of the sound wave environment is achieved.
Patent Information
- Application Number
- CN202511573474.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-31
AI Technical Summary
Existing acoustic imaging technologies have shortcomings in signal acquisition, spatial modeling, energy distribution characterization, and 3D imaging. They are unable to fully capture the spatial distribution characteristics of acoustic signals, reflect the propagation laws of acoustic waves, adapt to the dynamic changes of acoustic signals, and accurately mark discontinuous boundaries. As a result, the imaging results have a single dimension of information and cannot meet the environmental perception needs in complex scenarios.
Multi-channel acoustic signals are acquired by a distributed audio array, a spatial topology model containing the geometric constraints between the acoustic reflector and the obstacle is constructed, the acoustic energy attenuation gradient is calculated and dynamic partitions are divided, phase compensation is performed on the signal, the three-dimensional density field of acoustic energy is reconstructed, and the time-frequency domain characteristic parameters of the multi-channel signals are fused to generate a multimodal imaging spectrum.
It achieves comprehensive and accurate imaging of the acoustic environment, can track changes in acoustic energy in real time, accurately identify the location of key structures, and generate multimodal imaging maps that include time-frequency domain variation patterns and spatial energy distribution, meeting the environmental perception needs in complex scenarios.
Smart Images

Figure CN121028097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of acoustic imaging technology, specifically to an acoustic environment imaging method based on an acoustic array. Background Technology
[0002] Acoustic imaging technology, as a non-contact environmental sensing method, has been widely applied in various fields such as indoor space monitoring, industrial environmental detection, and security early warning due to its advantages such as strong environmental adaptability and controllable cost. Its core logic lies in inferring the environmental structure and energy distribution state within the target space by collecting and analyzing acoustic signals, providing intuitive evidence for subsequent environmental assessment and decision-making.
[0003] However, existing acoustic imaging technologies still face many unresolved problems. In the signal acquisition stage, traditional techniques often employ single-point or small-scale centralized sensor arrays, making it difficult to comprehensively capture the spatial distribution characteristics of acoustic signals within the target space. This results in missing time-frequency domain information, failing to fully reflect the propagation patterns of sound waves. In the spatial modeling stage, existing methods often neglect the geometric constraints between acoustic reflectors and obstacles, leading to deviations between the constructed propagation path model and the actual environment, thus affecting the accuracy of subsequent energy calculations.
[0004] In terms of energy distribution characterization, traditional techniques often employ static partitioning to define the sound wave energy range, which fails to adapt to the dynamic changes in sound wave signals over time, resulting in lag and limitations in energy distribution characterization. Simultaneously, sound waves are susceptible to phase distortion due to the propagation path, and existing techniques lack targeted compensation mechanisms, significantly reducing the accuracy of subsequent signal analysis and reconstruction. In 3D imaging, current methods often only provide rough estimates of the sound wave energy density field, failing to accurately mark discontinuous boundaries or clearly distinguish key structures such as obstacles and reflecting surfaces. Furthermore, existing imaging technologies often rely solely on single-dimensional signal features to generate maps, lacking the fusion of multimodal information, resulting in imaging results with limited information dimensions, unable to fully meet the environmental perception needs of complex scenes. Summary of the Invention
[0005] The purpose of this invention is to provide a sound wave environment imaging method based on an acoustic array to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides a method for acoustic environment imaging based on an acoustic array, the method comprising:
[0007] Multi-channel acoustic signals within the target space are acquired using a distributed audio array, and time-frequency domain characteristic parameters are extracted from the acoustic signals.
[0008] A spatial topology model of the sound wave propagation path is constructed based on time-frequency domain characteristic parameters. The spatial topology model includes the geometric constraint relationship between the sound wave reflecting surface and the obstacle.
[0009] The energy attenuation gradient of sound waves in the target space is calculated based on the spatial topology model, and the dynamic partition of sound wave energy distribution is divided.
[0010] Phase compensation is performed on the acoustic signal within the dynamic partition to generate a compensated acoustic signal sequence.
[0011] The three-dimensional density field of sound wave energy in the target space is reconstructed based on the compensated sound wave signal sequence, and the discontinuous boundary of the three-dimensional density field is marked.
[0012] By fusing the time-frequency domain characteristic parameters of multi-channel acoustic signals with the three-dimensional density field, a multimodal imaging spectrum of the acoustic environment is generated.
[0013] Preferably, the time-frequency domain feature parameters include:
[0014] The spectral entropy value of the acoustic signal within a preset time window;
[0015] The time delay difference between the arrival of the sound wave signal at each node of the distributed audio array;
[0016] The harmonic distortion coefficient of a sound wave signal in the frequency domain.
[0017] Preferably, the specific steps for constructing the spatial topology model of the sound wave propagation path are as follows:
[0018] Based on the time delay difference in the time-frequency domain characteristic parameters, calculate the path length difference of the sound wave signal to each sound node;
[0019] The spatial azimuth angle of the sound wave reflecting surface is inferred from the path length difference, and the geometric constraint relationship between the reflecting surface and the obstacle is established.
[0020] The geometric constraints are mapped to a topological connectivity graph of the sound wave propagation path, which includes the material attenuation coefficient of the reflecting surface.
[0021] Preferably, the step of dynamically partitioning the sound wave energy distribution includes:
[0022] The target space is divided into high decay region, medium decay region and low decay region according to the energy decay gradient.
[0023] The scattered components of the acoustic signal are extracted in the high attenuation region, and the spatial correlation of the scattered components is calculated.
[0024] The threshold values for the partitioning boundary between the medium attenuation zone and the low attenuation zone are dynamically adjusted based on spatial correlation.
[0025] Preferably, the step of performing phase compensation on the acoustic signal within the dynamic partition is as follows:
[0026] Identify phase abrupt changes in the compensated acoustic signal sequence;
[0027] The phase shift of the acoustic signal sequence is corrected based on the matching results between the phase abrupt change point and the discontinuous boundary of the three-dimensional density field.
[0028] The corrected phase offset is superimposed on the original acoustic signal sequence to generate a phase-aligned acoustic signal sequence.
[0029] Preferably, the steps for reconstructing the three-dimensional density field of acoustic wave energy include:
[0030] The compensated acoustic signal sequence is divided into multiple sub-sequences according to a time window;
[0031] Calculate the energy projection vector of each subsequence in three-dimensional space;
[0032] The initial grid of the density field is generated based on the spatial superposition of the energy projection vectors;
[0033] Topology optimization of the initial mesh based on discontinuous boundaries yields an optimized three-dimensional density field.
[0034] Preferably, the specific steps for calculating the energy projection vector are as follows:
[0035] Extract the dominant frequency component of the acoustic signal from the subsequence;
[0036] Based on the geometric constraint relationship between the dominant frequency component and the spatial topology model, the propagation direction vector of the sound wave energy is determined;
[0037] The propagation direction vector is multiplied by the energy amplitude of the acoustic signal to obtain the energy projection vector.
[0038] Preferably, the step of marking the discontinuous boundary of the three-dimensional density field includes:
[0039] Detect grid nodes in a three-dimensional density field whose energy gradient changes exceed a preset threshold;
[0040] Perform consistency verification on the direction of energy gradient change between adjacent grid nodes;
[0041] Mark the start and end positions of the non-continuous boundary based on the verification results.
[0042] Preferably, the steps for generating multimodal imaging atlases are as follows:
[0043] Map the time-frequency domain feature parameters to the corresponding grid nodes of the three-dimensional density field;
[0044] Assign rendering passes for acoustic environment imaging based on the parameter type of the mesh nodes;
[0045] Multimodal imaging atlases are generated by fusing imaging data from different rendering channels.
[0046] Preferably, the specific steps for allocating rendering channels are as follows:
[0047] The first rendering channel is assigned to the spectral entropy value to characterize the frequency domain distribution of the acoustic signal;
[0048] A second rendering channel is assigned to the harmonic distortion coefficients to characterize the degree of waveform distortion of the acoustic signal;
[0049] The energy value of the three-dimensional density field is assigned to the third rendering channel to characterize the spatial density of sound wave energy.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] By acquiring multi-channel acoustic signals through a distributed audio array, compared with traditional single-point or centralized acquisition methods, acoustic information can be acquired simultaneously from multiple spatial locations. The extracted time-frequency domain feature parameters are more comprehensive and representative, and can completely capture the propagation characteristics and variation patterns of acoustic signals in the target space, providing rich and reliable basic data for subsequent modeling and imaging.
[0052] The spatial topology model constructed based on time-frequency domain characteristic parameters incorporates the geometric constraints between the sound wave reflecting surface and obstacles, which can accurately reproduce the actual propagation path of the sound wave in the target space. This avoids the propagation simulation deviation caused by neglecting geometric constraints in traditional models, making the modeling of the sound wave propagation process more in line with the actual environmental conditions, and providing accurate model support for subsequent energy calculation and zoning.
[0053] Based on the spatial topology model, the acoustic energy attenuation gradient is calculated and dynamic partitions are divided. This method breaks through the limitations of traditional static partitioning, can track the changing trend of acoustic energy in real time, accurately capture the dynamic distribution characteristics of acoustic energy in different regions, and make the energy partitioning results more reflective of the real-time state of the acoustic environment, effectively improving the adaptability to dynamic acoustic scenarios.
[0054] Phase compensation of acoustic signals within dynamic partitions can effectively eliminate phase distortion caused by path differences and obstacle obstruction during acoustic propagation, making the compensated acoustic signal sequence closer to the original propagation state. This provides a high-quality signal source for subsequent three-dimensional density field reconstruction, ensuring the accuracy of imaging from the source.
[0055] By reconstructing the three-dimensional density field of sound wave energy and marking discontinuous boundaries through the compensated signal sequence, the three-dimensional distribution of sound wave energy in the target space can be clearly presented, and the location and outline of key structures such as obstacles and reflective surfaces can be accurately identified. This solves the problem that traditional reconstruction methods are difficult to accurately distinguish key spatial structures, and makes the spatial characteristics of the sound wave environment more intuitive and clear.
[0056] By integrating the time-frequency domain characteristic parameters of multi-channel acoustic signals with a three-dimensional density field to generate a multimodal imaging atlas, a combination of time, frequency, and spatial dimensional information is achieved, breaking the limitations of traditional single-dimensional imaging. The generated imaging atlas contains both the time-frequency domain variation patterns of the acoustic signals and the three-dimensional characteristics of spatial energy distribution, enabling a comprehensive depiction of the overall state of the acoustic environment from multiple perspectives and providing richer information support for environmental perception needs in different scenarios. Attached Figure Description
[0057] Figure 1 This is a schematic diagram illustrating the working principle of the acoustic environment imaging method based on an acoustic array as described in this invention.
[0058] Figure 2 A flowchart for constructing a spatial topology model of the sound wave propagation path;
[0059] Figure 3 A flowchart for reconstructing the three-dimensional density field of sound wave energy. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] Please see Figure 1This invention provides a method for acoustic environment imaging based on an acoustic array. The method includes: constructing a distributed acoustic array by deploying multiple acoustic nodes at different locations in a target space; synchronously acquiring acoustic wave signals propagating in the space by the distributed acoustic array; and sending the acquired multi-channel acoustic wave signals to a digital signal processing unit after analog-to-digital conversion. The digital signal processing unit preprocesses the acoustic wave signal of each channel, including noise reduction and signal normalization, and then extracts time-frequency domain feature parameters characterizing the acoustic wave characteristics from the preprocessed acoustic wave signals. These time-frequency domain feature parameters are input into a spatial modeling module as the basic data for constructing an environmental model. Based on the physical laws of sound wave propagation, the spatial modeling module uses the path information contained in the time-frequency domain feature parameters to deduce the geometric constraint relationships between reflecting surfaces and obstacles encountered by the sound wave in space, thereby constructing a spatial topology model describing the sound wave propagation path. The spatial topology model not only includes geometric information but also incorporates the attenuation characteristics of the reflecting surface materials on the sound waves. Based on the established spatial topology model, the energy calculation module simulates the propagation process of sound waves in the target space, calculates the attenuation gradient of sound wave energy as the propagation distance and number of reflections increase, and divides the target space into dynamic partitions with different energy characteristics according to the spatial distribution differences of the energy attenuation gradient. The phase processing module compensates for the phase inconsistency problem of sound wave signals in the dynamic partitions. The compensation operation takes into account the phase delay caused by the sound wave traveling through different paths, generating a sound wave signal sequence with better phase consistency. The imaging reconstruction module uses the compensated sound wave signal sequence to reconstruct the density distribution of sound wave energy in three-dimensional space through spatial interpolation and meshing algorithms, forming a three-dimensional density field, and identifies discontinuous boundaries with drastic energy changes in the density field. Finally, the multimodal fusion module fuses the original time-frequency domain feature parameters with the reconstructed three-dimensional density field at the data layer, assigns different visualization rendering channels to different feature parameters, and generates a multimodal imaging atlas of the sound wave environment that comprehensively reflects the frequency domain characteristics, waveform quality, and spatial energy distribution of the sound waves.
[0062] Example 1: See Figure 2The distributed audio array consists of multiple high-precision microphone nodes, deployed according to a predetermined geometric pattern at key locations on the boundary or within the target space. Each microphone node is equipped with independent signal conditioning circuitry and an analog-to-digital converter to ensure synchronous acquisition and digitization accuracy of multi-channel acoustic signals. The acquired multi-channel acoustic signals are transmitted to the central processing unit (CPU) in pulse-code modulation (PCM) format. The CPU runs digital signal processing algorithms to preprocess the raw data. Preprocessing includes bandpass filtering using a finite-length impulse response (FIR) filter to remove low-frequency environmental noise and high-frequency thermal noise from the circuitry, and amplitude normalization to eliminate signal amplitude inconsistencies caused by differences in microphone node sensitivity. The preprocessed acoustic signals then enter the feature extraction stage. The feature extraction algorithm is based on a short-time Fourier transform (SFT) framework. The system divides the continuous acoustic signal stream into a series of overlapping time windows, the length of which is dynamically adjusted according to the dominant frequency of the acoustic wave to ensure sufficient waveform period is included. For each signal segment within a time window, the Fast Fourier Transform (FFT) algorithm transforms it from the time domain to the frequency domain to obtain the power spectral density distribution. The calculation of the spectral entropy value is based on the Shannon entropy formula to quantize the power spectral density distribution. The quantization result reflects the complexity and uncertainty of the frequency components of the sound signal within a specific time window. The measurement of the time delay difference relies on the generalized cross-correlation function. The algorithm performs cross-correlation calculations on the same signal segment received by any two microphone nodes in the distributed audio array. By detecting the peak position of the cross-correlation function, the relative time difference of the sound wave arriving at the two nodes is determined. The product of this time difference and the sound wave propagation speed is the path length difference. The analysis of the harmonic distortion coefficient adopts the total harmonic distortion (THD) calculation method. The algorithm first determines the fundamental frequency and amplitude of the sound signal through spectral peak detection, and then extracts the amplitudes of the second harmonic, third harmonic, and so on up to the preset order of harmonic components. The THD is defined as the ratio of the root mean square value of each harmonic amplitude to the fundamental amplitude.
[0063] The spatial topology modeling unit receives the time-frequency domain feature parameter dataset from the feature extraction unit. The modeling process uses the path length difference obtained from the time delay difference conversion as the core input data. The path length difference constitutes an overdetermined system of equations, where the unknowns are the spatial coordinates of the sound source or reflection point relative to the microphone node array. Solving this system of equations employs an iterative least squares algorithm. Initially, the algorithm assumes an initial sound source position and calculates the sum of squared residuals between the theoretical path difference and the actual measured path difference. Gradient descent is then used to continuously adjust the sound source position estimate until the residuals are minimized. The process of inferring the spatial azimuth of the sound wave reflecting surface combines geometric acoustics principles and the mirror method. The algorithm equates the actual sound wave reflection path to a straight-line propagation path from the virtual mirror point of the sound source to the microphone node. By analyzing the coordinates of the virtual mirror points of the reflection path received by multiple microphone nodes, a plane fitting algorithm is used to calculate the normal vector of the reflecting surface, thereby determining its spatial azimuth. Establishing the geometric constraint relationship between the reflecting surface and obstacles requires the introduction of a spatial geometric inference engine. The engine treats the reflecting surface as an infinitely extending plane and models obstacles as convex polyhedra in three-dimensional space. The derivation of geometric constraints is accomplished through intersection tests of the reflector plane equation and the obstacle boundary box. The test results can determine whether there are parallel, perpendicular, or oblique spatial relationships between the reflector and the obstacle. For example, when the dot product of the reflector normal vector and the normal vector of a certain surface of the obstacle is close to zero, it is considered a parallel relationship; when the absolute value of the dot product is close to one, it is considered a perpendicular relationship. The assignment of material attenuation coefficients is achieved by querying a pre-built acoustic material database. The database stores the sound absorption coefficients and reflection coefficients of common building materials such as concrete, glass, and wood at different frequencies. The system retrieves the corresponding attenuation parameters from the database based on the material identifier of the reflector.
[0064] The topology connectivity graph is constructed using graph theory data structures. The vertex set of the graph includes the sound source, each microphone node, and the calculated reflection points. The edge set represents all possible propagation paths of sound waves between vertices. Each edge is accompanied by a set of attribute values, including the physical length of the path, the number of reflections experienced by the path, and the product of the material attenuation coefficients of each reflecting surface along the path. The connectivity of the graph is verified using a ray tracing algorithm. The algorithm emits a large number of virtual sound rays from the sound source and tracks the propagation trajectory of the sound rays between reflecting surfaces. Only those paths that ultimately reach at least one microphone node are retained in the topology connectivity graph. The final topology connectivity graph is stored in memory in the form of an adjacency matrix. The non-zero elements of the adjacency matrix represent the comprehensive attenuation weight of the path, which is used for subsequent sound wave energy propagation simulation and attenuation gradient calculation. The entire implementation process relies on a high-precision time synchronization mechanism. The microphone nodes of the distributed audio array achieve microsecond-level synchronization of the sampling clock through a precise clock protocol. Time synchronization errors directly lead to distortion in the time delay difference measurement, thus affecting the accuracy of the spatial topology model. The signal processing link latency is rigorously calibrated, ensuring a constant total delay from analog signal input to digital feature parameter output. This constant delay facilitates the alignment and correlation of time series data. The feature extraction unit and the spatial topology modeling unit communicate via a high-speed data bus, which employs a timestamp matching mechanism to ensure accurate association between each feature parameter set and its corresponding acoustic signal segment.
[0065] The spatial topology model update mechanism supports dynamic environmental adaptation. The system periodically re-executes the feature extraction and modeling process, detecting the movement or addition of obstacles in the environment by comparing changes in the topology connection graph over consecutive periods. During dynamic updates, the algorithm maintains a historical model cache, compares the current model with the historical model, marks changed vertices and edges of the graph, and evaluates the significance of the changes. Only changes with significance exceeding a threshold trigger a global update of the entire acoustic environment imaging model. This incremental update strategy balances computational efficiency with the response speed to environmental changes. The model accuracy verification is integrated within the topology modeling unit. The verification method adopts a cross-validation strategy. The system randomly hides the measurement data of some microphone nodes, reconstructs the topology model using the data of the remaining nodes, and then compares the path difference of the hidden nodes predicted by the reconstructed model with the actual measurement values. Error statistics such as root mean square error and mean absolute error are calculated and recorded in real time. When the error exceeds the preset tolerance, the system triggers the model reconstruction process instead of directly applying the current model. Quality monitoring of feature parameters is carried out throughout the entire process of this embodiment. The system sets reasonable value ranges and noise thresholds for spectral entropy, time delay difference, and harmonic distortion coefficients, respectively. Any abnormal parameter value that exceeds the range will be marked as invalid data and trigger data resampling or algorithm parameter adjustment. The monitoring logic establishes a control chart based on statistical process control methods to track the quality fluctuations of characteristic parameters in real time and correct systematic deviations in a timely manner.
[0066] Example 2: The energy calculation module receives a topology diagram of the sound wave propagation path from the spatial topology modeling unit. This diagram contains geometric information and material attenuation coefficients for the propagation paths from all sound source points to microphone nodes. The energy attenuation gradient is calculated based on a propagation loss model of sound waves in free space and during reflection. The model considers two basic mechanisms: geometric diffusion attenuation and air absorption attenuation. Geometric diffusion attenuation follows the amplitude attenuation law of spherical or cylindrical waves, while the air absorption attenuation coefficient is calculated using international standard formulas based on ambient temperature, humidity, and sound wave frequency. For each sound wave propagation path, the algorithm integrates the path length, number of reflections, and material attenuation coefficients of each reflecting surface to calculate the total energy attenuation value of the sound wave propagating from the sound source point along the path to any specified point in the target space. The target space is discretized into a three-dimensional grid. The energy calculation module calculates the energy attenuation value of the sound wave reaching each grid node through all effective paths in parallel, and takes the energy value of the path with the smallest attenuation as the final energy estimate for that grid node.
[0067] Based on the energy attenuation distribution of grid nodes, the energy calculation module uses a gradient descent algorithm to identify the spatial rate of change of energy attenuation values in the three-dimensional grid space. The algorithm calculates the energy attenuation difference between each grid node and its six neighboring nodes, and obtains the amplitude and direction of the energy attenuation gradient through three-dimensional difference operations. The dynamic partitioning of the sound wave energy distribution is based on two threshold boundaries set according to the energy attenuation gradient amplitude: a gradient amplitude greater than the high threshold is defined as a high attenuation zone, a gradient amplitude between the high and low thresholds is defined as a medium attenuation zone, and a gradient amplitude less than the low threshold is defined as a low attenuation zone. The initial values of the threshold boundaries are set based on historical experience data and dynamically adjusted during system operation. The high attenuation zone typically corresponds to areas where sound waves need to cross physical barriers or undergo multiple reflections, the medium attenuation zone represents areas with single reflections or medium-distance propagation, and the low attenuation zone is mainly the direct sound field area near the sound source.
[0068] In the high attenuation region, the signal processing algorithm needs to extract the scattered components of the acoustic signal. The extraction of scattered components employs independent component analysis (ICA) blind source separation technology. The algorithm treats the mixed signal received by multiple microphone nodes of the distributed acoustic array as a linear mixture of several independent source signals. ICA separates the source signals by maximizing the non-Gaussianity of the signal components. Components without obvious directionality and with weak correlation among the separated signal components are identified as scattered components. Calculating the spatial correlation of the scattered components requires analyzing the statistical relationship between the scattered signals received by different microphone nodes. Spatial correlation calculation is based on the cross-correlation matrix of the scattered signal segments. Eigenvalue decomposition is performed on the cross-correlation matrix, and the ratio of the largest eigenvalue to the sum of eigenvalues is used as a quantitative indicator of spatial correlation. A higher spatial correlation ratio indicates that the scattered energy mainly comes from a few dominant reflection paths, while a lower spatial correlation ratio reflects a more diffuse distribution of scattered energy. Based on the spatial correlation calculation results, the system dynamically adjusts the boundary thresholds between the medium and low attenuation regions. When the spatial correlation of scattered components is high in the high attenuation region, the algorithm appropriately increases the boundary thresholds between the medium and low attenuation regions, causing the medium attenuation region to expand towards the low attenuation region. The boundary threshold adjustment employs a proportional-integral-derivative (PI-DE) control strategy. The proportional term responds to the instantaneous changes in spatial correlation, the integral term accumulates historical trends, and the derivative term predicts future directions of change. The controller's output acts on the offset of the partition boundary threshold, achieving smooth adaptive changes in the partition boundaries. The dynamic partitioning results are stored in the form of a three-dimensional mask matrix. The value of each element in the matrix identifies the partition category to which the corresponding grid node belongs. This mask matrix will be used to guide subsequent differential signal processing.
[0069] The phase processing module receives the dynamic partitioning results and the corresponding acoustic signal data. Phase compensation is performed separately for the signal within each partition. Phase abrupt change detection employs a phase difference method. The algorithm performs a first-order difference operation on the instantaneous phase of the acoustic signal, and points with difference values exceeding a preset threshold are marked as phase abrupt changes. Phase abrupt change detection is performed jointly in the time and frequency domains. The signal first undergoes a short-time Fourier transform to obtain the time spectrum, and phase difference operations are performed independently in each frequency band. Events where phase abrupt changes are detected simultaneously in multiple frequency bands are assigned higher confidence. Detected phase abrupt changes need to be matched with the discontinuous boundary of the three-dimensional density field. The matching process calculates the spatial distance between the acoustic propagation path corresponding to the timestamp of the phase abrupt change and the discontinuous boundary of the three-dimensional density field. Matches with a distance less than the tolerance value are considered valid matches. The phase offset of the acoustic signal sequence is corrected based on the matching results, calculated based on the geometric path difference of acoustic propagation. For each matched phase abrupt change, the algorithm recalculates the theoretical propagation time of the acoustic wave from the sound source to the microphone node based on the spatial location of the discontinuous boundary. The difference between the theoretical propagation time and the actual detected abrupt change time is converted into a phase offset. The phase offset is calculated considering the center frequency of the sound wave. The time difference is multiplied by the angular frequency to obtain a phase correction value in radians. This correction value is applied to a segment of the original sound wave signal sequence within a time window centered on the phase abrupt change point. Phase compensation is performed in the frequency domain. The algorithm transforms the signal segment to the frequency domain and applies a linear phase rotation to each frequency component of the spectrum. The rotation angle is determined by the calculated phase offset. After frequency domain processing, the signal is restored to the time domain using an inverse Fourier transform.
[0070] Generating a phase-aligned acoustic signal sequence requires integrating the compensation results from all partitions and time windows. The integration process employs an overlap-addition method, applying a window function to each compensated signal segment and weighted averaging the overlapping signals to eliminate boundary effects caused by segmentation. The final phase-aligned acoustic signal sequence maintains the same sampling rate and length as the original sequence, but with significantly improved phase continuity. The phase-aligned sequence is stored in a circular buffer for use by the subsequent 3D density field reconstruction module. The entire phase compensation process is implemented on a field-programmable gate array (FPGA) hardware platform, utilizing a parallel processing architecture to meet real-time requirements. Parameters of the phase compensation algorithm, such as abrupt change detection threshold and matching tolerance, can be adjusted online via a configuration interface to adapt to different acoustic environments and signal characteristics. The system continuously monitors the phase compensation effect, using methods including calculating the coherence index and phase variance of the signal segments before and after compensation. The coherence index assesses the consistency between signals from different microphone nodes, while the phase variance reflects the phase stability within a single signal segment. The monitoring results are fed back to the parameter adjustment logic of the phase processing module, forming a closed-loop control. When the monitored indicators are lower than expected, the system automatically fine-tunes the phase compensation parameters or triggers a re-compensation process. This self-optimization mechanism ensures that the phase compensation operation maintains good performance when environmental conditions change, providing a fundamental guarantee for the accuracy of acoustic environmental imaging.
[0071] Example 3: See Figure 3 The imaging reconstruction module receives a phase-aligned acoustic signal sequence from the phase processing module. This sequence is stored in a buffer as a discrete-time signal with a fixed sampling rate. The segmentation of the acoustic signal sequence into multiple sub-sequences employs an overlapping sliding window mechanism. The window length is dynamically adjusted based on the dominant frequency period of the acoustic signal to ensure each window contains an integer multiple of the period. A Hanning window is used to reduce spectral leakage, and the overlapping area is set to 50% of the window length to ensure temporal continuity. Each sub-sequence carries a timestamp and is associated with the location of a potential sound source in three-dimensional space. This association is calculated based on the time delay difference between the arrival times of the sound waves at each node of the distributed acoustic array. Calculating the energy projection vector of each sub-sequence in three-dimensional space requires extracting the dominant frequency component of the acoustic signal. The identification of the dominant frequency component uses autocorrelation function analysis. The algorithm calculates the autocorrelation function of the sub-sequence signal and detects its maximum peak position. The reciprocal of the dominant frequency component corresponds to the time delay of the maximum peak of the autocorrelation function. The amplitude of the dominant frequency component is obtained by calculating the short-time Fourier transform modulus of the sub-sequence signal at the corresponding frequency. The frequency resolution is determined by the window length. Determining the propagation direction vector of sound wave energy relies on the geometric constraints provided by the spatial topological model. These constraints include the direction of the normal vector of the reflecting surface and the definition of the incident plane of the sound wave. The calculation of the propagation direction vector combines the specular reflection law of sound rays; the dot product of the incident direction and the normal vector of the reflecting surface determines the vector expression of the reflection direction.
[0072] The energy projection vector is generated through vector construction, combining a scalar representing the magnitude of the sound wave intensity with a unit vector representing the propagation direction. The mathematical expression for the sound intensity projection vector is:
[0073]
[0074] in: The sound intensity projection vector represents the directionality of the sound wave energy flux density in space, and its dimension is watts per square meter [W / m²]. It represents the sound intensity amplitude, is a scalar, is calculated based on the square of the root mean square value of the sound pressure signal, and has the dimension of watts per square meter [W / m²]. The unit vector representing the direction of propagation is a dimensionless quantity.
[0075] The spatial superposition of sound intensity projection vectors is performed in a three-dimensional mesh space, the resolution of which is pre-set according to the target space size and accuracy requirements. Each mesh cell maintains an accumulator, which sums all sound intensity projection vectors projected onto that cell. Projection determination is based on a ray casting algorithm, which emits rays from the potential location of the sound source along the propagation direction vector; mesh cells through which the rays pass are marked as affected cells. The initial mesh is generated using inverse distance weighted interpolation, with the accumulated energy value of each mesh cell used as the node value for surface fitting to generate a continuous three-dimensional energy distribution representation. Topology optimization of the initial mesh based on discontinuous boundaries requires the introduction of adaptive mesh refinement technology; this technology automatically increases the mesh density in areas where discontinuous boundaries are detected, while maintaining a coarser mesh in areas with gradual energy changes. The optimization process employs a mesh refinement strategy based on error estimation. The error estimator calculates the second derivative of the energy value within each mesh cell; regions with large second derivatives indicate drastic changes in energy distribution. The topology optimization algorithm inserts new grid nodes at discontinuous boundaries. The energy value of the new node is calculated from the neighboring nodes through cubic spline interpolation, thus maintaining the continuity and smoothness of the energy distribution.
[0076] The optimized 3D density field is stored in voxel grid form, with each voxel containing energy intensity value and spatial coordinate information. Visualization of the 3D density field employs direct volume rendering technology, which uses a ray casting algorithm to calculate the color and transparency of each pixel, generating a perspective image of the 3D energy distribution. The 3D density field data is simultaneously output to a multimodal fusion module for data fusion with frequency domain feature parameters. The entire reconstruction process includes multiple quality control steps, achieved by calculating reconstruction error metrics. These metrics include energy conservation checks and boundary consistency checks. The energy conservation check compares the total energy of the input signal with the energy difference of the volume integral of the 3D density field, while the boundary consistency check verifies the alignment of discontinuous boundaries with the physical boundaries in the acoustic image. When the error metrics exceed the allowable range, the system automatically triggers a re-execution of the reconstruction process and adjusts relevant algorithm parameters such as grid resolution or interpolation methods.
[0077] The computational efficiency of the reconstruction algorithm is optimized through parallel computing technology. The 3D mesh space is divided into multiple sub-regions and assigned to different computing units for parallel processing. Graphics processor hardware acceleration is used for the calculation of sound intensity projection vectors and mesh interpolation operations, enabling real-time processing of large-scale 3D datasets. The system supports dynamic resolution adjustment, automatically balancing reconstruction accuracy and processing speed based on available computing resources. When computing resources are limited, a multi-resolution hierarchical processing strategy is adopted. The 3D density field is stored using a hierarchical data structure, with the bottom layer storing full-resolution data and the upper layer storing downsampled multi-resolution representations. The hierarchical structure supports fast data retrieval and visualization, allowing users to automatically switch to the appropriate resolution data representation based on the viewing scale. A data compression algorithm is applied to the persistent storage of the 3D density field. The compression algorithm is based on predictive coding and entropy coding of energy values, reducing storage space requirements while maintaining accuracy. The reliability of the reconstruction results is verified using cross-validation. This method randomly divides the nodes of the distributed acoustic array into training and test sets. After reconstructing the 3D density field using the training set node data, the energy values of the test set node positions are predicted and compared with the actual measured values. The verification results generate a consistency report, which identifies regions with high confidence and regions with potential uncertainties in the 3D density field. The confidence information is stored as metadata along with the 3D density field for reference by subsequent processing modules.
[0078] Example 4: The boundary detection module receives optimized 3D density field data from the imaging reconstruction module. The 3D density field data is organized in a regular grid format, with each grid node storing a normalized acoustic energy density value ranging from zero to one. The process of marking discontinuous boundaries begins with calculating the energy gradient field of the 3D density field. The energy gradient field is calculated using a central difference algorithm to calculate the partial derivatives of each grid node in the three coordinate axes. For internal grid nodes, the algorithm uses a six-neighbor template to calculate the gradient components in the X, Y, and Z directions. For boundary grid nodes, a forward or backward difference format is used to avoid out-of-bounds access. Detecting grid nodes where the energy gradient change exceeds a preset threshold requires setting dual judgment conditions. The first condition judges whether the gradient magnitude exceeds a global threshold, and the gradient magnitude is calculated as the Euclidean norm of the gradient components in the three directions. The second condition judges the consistency of the gradient directions, filtering out isolated noise points by analyzing the angle between the gradient vectors of adjacent grid nodes. The global threshold is determined based on the statistical distribution of the overall gradient magnitude of the 3D density field. The system calculates the mean and standard deviation of the gradient magnitude and sets the threshold to the mean plus three times the standard deviation to capture significant gradient changes. A local adaptive threshold supplements weak boundaries that the global threshold may miss. The local threshold is dynamically adjusted based on the gradient magnitude range within a small neighborhood around each grid node.
[0079] The consistency verification of the energy gradient change direction of adjacent grid nodes is performed using vector field analysis. This method calculates the average dot product of the gradient vector of each grid node and the gradient vectors of its six immediate neighbors. A dot product close to one indicates a high degree of consistency in gradient directions; a dot product close to negative one indicates opposite gradient directions; and a dot product close to zero indicates orthogonal gradient directions. An angle tolerance parameter is set for the consistency verification; the verification is considered successful only when the angle between the gradient vectors of more than half of the neighboring grid nodes and the center node is less than 15 degrees. Grid nodes that fail the consistency verification are considered noise points and excluded, even if their gradient magnitude exceeds a threshold. Marking the start and end positions of discontinuous boundaries based on the verification results requires a boundary tracing algorithm. This algorithm starts from any seed point that passes the threshold detection and consistency verification and grows the region along a plane perpendicular to the gradient direction. The boundary tracing algorithm adopts a 3D version of the Cannibal edge detection approach. First, non-maximum suppression is performed to retain grid nodes with locally maximum gradient magnitudes. Then, strong edge nodes are combined with connected weak edge nodes to form a complete boundary profile through double-threshold connections. Post-processing of boundary segments includes removing short, isolated segments, connecting broken boundary gaps, and smoothing sawtooth fluctuations in the boundary curve.
[0080] The marking results are stored as a set of spatial coordinates of boundary grid nodes, with each boundary segment accompanied by topological connectivity information describing the adjacency relationships between segments. The system calculates a geometric feature descriptor for each boundary segment, including parameters such as the average curvature, total length, bounding box size, and normal vector direction. These geometric features are used for boundary type classification and rendering style allocation in subsequent multimodal imaging atlas generation. Quality control of the boundary marking process is achieved through a multi-scale verification mechanism, which repeatedly executes the boundary detection algorithm on grid representations at different resolutions. The consistency of detected boundary positions at different scales is compared, and only those boundaries that consistently appear at multiple scales are confirmed as valid discontinuous boundaries. Multi-scale verification effectively suppresses false boundaries introduced by grid discretization, improving the robustness and reliability of boundary marking. The marking results of discontinuous boundaries in the 3D density field are stored in a dedicated boundary database, which uses a spatial index structure to support fast range queries and nearest neighbor searches. Boundary data and the original 3D density field data are linked through spatial coordinates, allowing for rapid retrieval of the corresponding energy density distribution characteristics based on the boundary location. The database also records the detection confidence score of the boundary, which is calculated based on the gradient magnitude, consistency check score, and multi-scale verification consistency. The computational performance of the boundary marking algorithm is optimized through spatial partitioning, dividing the 3D mesh space into appropriately sized sub-blocks. Each sub-block performs boundary detection independently before merging the global results. Overlapping regions are set at the boundaries of sub-blocks to prevent boundary segments from being cut off, and the merging process achieves seamless connection by matching boundary nodes within the overlapping regions. The parallel computing architecture fully utilizes multi-core processors and graphics processing units to accelerate vector operations and neighborhood lookup operations in boundary detection. The system provides a visual configuration interface for boundary marking parameters, allowing users to interactively adjust parameters such as gradient threshold, consistency tolerance, and minimum boundary length, and view the marking effect in real time. The parameter adjustment history and corresponding boundary detection results are automatically recorded to form a parameter-effect mapping relationship. This mapping relationship is analyzed through machine learning algorithms to form an automatic parameter recommendation function, helping users quickly obtain satisfactory boundary marking results.
[0081] Referring to Table 1, the labeling results of discontinuous boundaries are not only used for visualization rendering but also provide structural information for acoustic environment analysis. Hard boundaries correspond to the surface of physical obstacles, while soft boundaries correspond to regions with varying acoustic impedance. Different types of boundaries are distinguished and displayed with different colors and textures in the multimodal imaging spectrum. Boundary curvature information reveals the surface's unevenness, and the boundary orientation distribution reflects the structural orientation characteristics of space.
[0082] Table 1: Geometric Feature Descriptors for Discontinuous Boundaries
[0083] Feature Name Calculation method Physical meaning Data types Boundary segment length Cumulative Euclidean distances between adjacent nodes on the boundary path Reflecting the extent of the boundary space extension Floating-point number (meters) Mean curvature Arithmetic mean of the curvature at each point on the boundary path Characterizing the degree of boundary curvature Floating-point number (1 / meter) Normal vector consistency The standard deviation of the angle between the normal vector at each point on the segment and the mean normal vector Describe the smoothness of the boundary surface Floating-point number (degrees) Gradient magnitude mean Arithmetic mean of gradient magnitudes at boundary nodes Indicator of energy change intensity at the boundary floating-point numbers Boundary dimension ratio Ratio of the longest side to the shortest side of the bounding box Reflecting the anisotropy of boundary shape floating-point numbers
[0084] The boundary marking module is integrated with other modules of the system through a standardized data interface, which defines the format, coordinate system conventions, and update mechanism of the boundary data. When the 3D density field is updated, the boundary marking module automatically triggers recalculation. The incremental update algorithm only performs local boundary remarking on regions where the density field changes to improve efficiency. The history of the marking results is recorded for analyzing the dynamic evolution of the acoustic environment.
[0085] Example 5: The fusion rendering module receives time-frequency domain feature parameters from the feature extraction unit and three-dimensional density field data from the imaging reconstruction module. The time-frequency domain feature parameters include attributes such as the spectral entropy value and harmonic distortion coefficient associated with each grid node. The three-dimensional density field data includes the spatial coordinates and energy density values of the grid nodes. The generation of the multimodal imaging atlas begins with the data alignment step. Data alignment maps the time-frequency domain feature parameters to the corresponding grid nodes of the three-dimensional density field based on a unified space-time coordinate system. Each grid node is associated with a set of feature values through a spatial coordinate index. The feature values include energy density values, spectral entropy values, and harmonic distortion coefficients, forming a set of data points with multidimensional attributes. Mapping the time-frequency domain feature parameters to the corresponding grid nodes of the three-dimensional density field requires solving the data scale difference problem. The time-frequency domain feature parameters are derived from the acoustic signal processing results, and the three-dimensional density field is derived from the spatial reconstruction algorithm. The data mapping employs a spatial interpolation method to resample the time-frequency domain feature parameters to the grid node positions of the three-dimensional density field. The interpolation method considers the physical characteristics of sound wave propagation. For each grid node, the system searches for its neighboring sound wave sampling points and calculates the node's feature parameter values using a weighted average based on the propagation path length. The mapping process ensures that each grid node obtains a complete feature vector, which serves as the data foundation for multimodal rendering.
[0086] Assigning rendering channels for acoustic environment imaging based on the parameter types of mesh nodes requires designing a color mapping function. This function converts different types of feature parameters into visually distinguishable color components. The first rendering channel is assigned to the spectral entropy values, which uses a continuous color gradient scheme to represent the numerical range of the spectral entropy values. Low spectral entropy values are mapped to a blue hue, medium spectral entropy values to a green hue, and high spectral entropy values to a red hue. This color gradient is achieved through linear interpolation in the HSV color space. The second rendering channel is assigned to the harmonic distortion coefficients, which uses a brightness gradient scheme to represent the degree of waveform distortion. Low harmonic distortion coefficients are displayed in a dark hue, and high harmonic distortion coefficients are displayed in a bright hue. The brightness value and harmonic distortion coefficient have a linear relationship. The third rendering channel is assigned to the energy values of the three-dimensional density field, which uses a transparency gradient scheme to represent the spatial density of acoustic energy. Low-energy areas are displayed with high transparency, and high-energy areas are displayed with low transparency. The transparency change follows an exponential decay curve to enhance visual contrast.
[0087] The generation of multimodal imaging atlases by fusing imaging data from different rendering channels employs a pixel-level overlay algorithm. This algorithm weights and mixes the color contributions of each grid node across the three rendering channels. Color mixing is performed in the RGBA color space, and the final color value of each grid node is determined by the hue of the first rendering channel, the brightness of the second rendering channel, and the transparency of the third rendering channel. The mixing weights are dynamically adjusted based on the observer's analytical needs: increasing the weight of the first rendering channel when emphasizing frequency domain characteristic analysis, increasing the weight of the second rendering channel when emphasizing waveform quality assessment, and increasing the weight of the third rendering channel when emphasizing energy distribution observation.
[0088] The visualization of multimodal imaging atlases is based on volumetric rendering technology, which generates a 2D projected image of a 3D scene using a ray casting algorithm. Each ray emanating from the viewpoint passes through the 3D data field, sampling multiple points along the ray path. The color and transparency of each sampled point are determined by trilinear interpolation of the multimodal color value of its corresponding grid node. All sampled points on the ray are color-blended in either a front-to-back or back-to-front order to ultimately form the color value of the pixel. The rendering process is performed in real time, supporting viewpoint rotation, scaling, and translation, providing multi-angle observation capabilities of the acoustic environment. The system provides an adjustable interface for rendering parameters, allowing users to independently adjust parameters such as contrast, brightness, and color balance of the three rendering channels using slider controls. The effect of parameter adjustments is displayed in real time in the display window, allowing users to find the optimal visualization configuration through interactive exploration. Commonly used parameter combinations can be saved as preset schemes for quick loading and use in specific application scenarios. Preset schemes include dedicated configurations such as spectrum analysis mode, distortion detection mode, and energy distribution mode.
[0089] The multimodal imaging atlas supports layered display, allowing users to individually enable or disable the display of a specific rendering channel. Layered display facilitates the analysis of spatial correlations between different feature parameters; for example, the spectral entropy distribution can be displayed separately, then overlaid with the energy distribution contour to observe the correspondence between frequency domain characteristics and energy distribution. The display control interface provides channel transparency adjustment, enabling a smooth transition between different channel display intensities. The atlas generation process includes an automatic optimization mechanism that analyzes the data distribution characteristics under the current view and automatically adjusts rendering parameters to enhance visualization. When a large dynamic range of data is detected, the system automatically enables logarithmic color mapping to enhance the display contrast of low-value areas; when the data distribution is relatively concentrated, linear color mapping is used to maintain a direct correspondence between numerical values and colors. Automatic optimization reduces the workload of manual adjustments by users and improves atlas generation efficiency.
[0090] The multimodal imaging atlas supports multiple standards for output formats, including 2D image sequences, 3D volumetric data files, and interactive web visualization formats. 2D image sequences generate animation frames rotating around the model, used for creating demonstration videos; 3D volumetric data files store complete multimodal data for further processing by professional analysis software; the interactive web format, implemented using WebGL technology, allows for 3D viewing and basic analysis operations via a browser. Meta-information of the atlas data is fully recorded, including data acquisition time, spatial extent, coordinate system definition, rendering parameter configuration, and eigenvalue statistics. This meta-information is embedded in the output file in JSON format, ensuring data traceability and repeatability. Users can quickly understand the atlas's generation conditions and data characteristics through the meta-information. Observers can simultaneously obtain multi-dimensional information from a single view, such as the spatial distribution of acoustic energy, regional differences in spectral characteristics, and spatial patterns of waveform distortion. Integrated visualization enhances the understanding of complex acoustic environments and reveals the physical mechanisms of the interaction between sound waves and spatial structures. Multimodal imaging atlases have application value in fields such as architectural acoustic design, noise source localization, and abnormal acoustic event detection. The atlases support quantitative analysis functions, allowing users to measure the characteristic parameter values of specific areas in a 3D scene and generate statistical reports to support professional decision-making processes.
[0091] Performance optimization of the rendering system ensures smooth interaction in large-scale data scenarios. Level of detail technology automatically adjusts the mesh resolution based on view distance, and frustum clipping avoids unnecessary rendering of invisible areas. A multi-threaded architecture distributes data loading, feature calculation, and image rendering tasks across different threads for parallel execution, maintaining a responsive user interface. Graphics processor hardware acceleration enables efficient execution of ray casting algorithms, supporting real-time rendering of datasets with millions of mesh nodes. A bidirectional link is established between the multimodal imaging atlas and the raw acoustic data. Users can click on regions of interest in the atlas, and the system automatically retrieves and plays the corresponding raw acoustic signal. This bidirectional link directly links visual analysis with auditory perception, helping users verify the auditory correspondence of the visualization results. The click-to-query function displays detailed feature parameter values for the current location, overlaid in a table format within the 3D scene.
[0092] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0093] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for acoustic environment imaging based on an acoustic array, characterized in that, Includes the following steps: Multi-channel acoustic signals within the target space are acquired using a distributed audio array, and time-frequency domain characteristic parameters are extracted from the acoustic signals. A spatial topology model of the sound wave propagation path is constructed based on time-frequency domain characteristic parameters. The spatial topology model includes the geometric constraint relationship between the sound wave reflecting surface and the obstacle. The energy attenuation gradient of sound waves in the target space is calculated based on the spatial topology model, and the dynamic partition of sound wave energy distribution is divided. Phase compensation is performed on the acoustic signal within the dynamic partition to generate a compensated acoustic signal sequence. The three-dimensional density field of sound wave energy in the target space is reconstructed based on the compensated sound wave signal sequence, and the discontinuous boundary of the three-dimensional density field is marked. By fusing the time-frequency domain characteristic parameters of multi-channel acoustic signals with the three-dimensional density field, a multimodal imaging spectrum of the acoustic environment is generated.
2. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The time-frequency domain feature parameters include: The spectral entropy value of the acoustic signal within a preset time window; The time delay difference between the arrival of the sound wave signal at each node of the distributed audio array; The harmonic distortion coefficient of a sound wave signal in the frequency domain.
3. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The specific steps for constructing a spatial topology model of the sound wave propagation path are as follows: Based on the time delay difference in the time-frequency domain characteristic parameters, calculate the path length difference of the sound wave signal to each sound node; The spatial azimuth angle of the sound wave reflecting surface is inferred from the path length difference, and the geometric constraint relationship between the reflecting surface and the obstacle is established. The geometric constraints are mapped to a topological connectivity graph of the sound wave propagation path, which includes the material attenuation coefficient of the reflecting surface.
4. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The steps for dynamically partitioning the sound wave energy distribution include: The target space is divided into high decay region, medium decay region and low decay region according to the energy decay gradient. The scattered components of the acoustic signal are extracted in the high attenuation region, and the spatial correlation of the scattered components is calculated. The threshold values for the partitioning boundary between the medium attenuation zone and the low attenuation zone are dynamically adjusted based on spatial correlation.
5. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The steps for phase compensation of acoustic signals within a dynamic partition are as follows: Identify phase abrupt changes in the compensated acoustic signal sequence; The phase shift of the acoustic signal sequence is corrected based on the matching results between the phase abrupt change point and the discontinuous boundary of the three-dimensional density field. The corrected phase offset is superimposed on the original acoustic signal sequence to generate a phase-aligned acoustic signal sequence.
6. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The steps for reconstructing the three-dimensional density field of acoustic wave energy include: The compensated acoustic signal sequence is divided into multiple sub-sequences according to a time window; Calculate the energy projection vector of each subsequence in three-dimensional space; The initial grid of the density field is generated based on the spatial superposition of the energy projection vectors; Topology optimization of the initial mesh based on discontinuous boundaries yields an optimized three-dimensional density field.
7. The acoustic environment imaging method based on an acoustic array according to claim 6, characterized in that, The specific steps for calculating the energy projection vector are as follows: Extract the dominant frequency component of the acoustic signal from the subsequence; Based on the geometric constraint relationship between the dominant frequency component and the spatial topology model, the propagation direction vector of the sound wave energy is determined; The propagation direction vector is multiplied by the energy amplitude of the acoustic signal to obtain the energy projection vector.
8. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The steps for marking the discontinuous boundary of a three-dimensional density field include: Detect grid nodes in a three-dimensional density field whose energy gradient changes exceed a preset threshold; Perform consistency verification on the direction of energy gradient change between adjacent grid nodes; Mark the start and end positions of the non-continuous boundary based on the verification results.
9. The acoustic environment imaging method based on an acoustic array according to claim 1, characterized in that, The steps for generating a multimodal imaging atlas are as follows: Map the time-frequency domain feature parameters to the corresponding grid nodes of the three-dimensional density field; Assign rendering passes for acoustic environment imaging based on the parameter type of the mesh nodes; Multimodal imaging atlases are generated by fusing imaging data from different rendering channels.
10. The acoustic environment imaging method based on an acoustic array according to claim 9, characterized in that, The specific steps for allocating rendering channels are as follows: The first rendering channel is assigned to the spectral entropy value to characterize the frequency domain distribution of the acoustic signal; A second rendering channel is assigned to the harmonic distortion coefficients to characterize the degree of waveform distortion of the acoustic signal; The energy value of the three-dimensional density field is assigned to the third rendering channel to characterize the spatial density of sound wave energy.
Citation Information
Patent Citations
Ultrasonic sensor adjusting method, distance measuring method, medium and electronic equipment
CN112285679A
High-voltage electrical equipment inspection method and system based on microphone array measurement
CN120252946A
Building sound wave reflection path analysis and noise source positioning method and system
CN120256868A
Three-dimensional sound event accurate identification and positioning method based on AI and multi-microphone array
CN120428168A
Remote intelligent lock control method based on cloud platform
CN120472567A
Cited By
Data processing-based mochi bread fermentation quality prediction system
CN121256284A
A data processing-based fermented quality prediction system for steamed bread
CN121256284B