A method and system for complex sound field reconstruction based on a physically constrained diffusion model

By using a physical constraint diffusion model-based approach, a conditional denoising network and angular spectrum method are employed to reconstruct the complex sound field, thus solving the problems of high computational complexity and multiple solutions in complex amplitude hologram reconstruction and achieving efficient and stable complex sound field reconstruction.

CN122395537APending Publication Date: 2026-07-14BEIJING INSTITUTE OF GRAPHIC COMMUNICATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INSTITUTE OF GRAPHIC COMMUNICATION
Filing Date
2026-04-14
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing acoustic holography techniques suffer from high computational complexity and limited convergence speed in complex amplitude hologram reconstruction, and are also subject to multiple solutions and ill-posedness, making it difficult to achieve high-precision reconstruction of complex sound fields.

Method used

A method based on a physical constraint diffusion model is adopted. The noise component is predicted by a conditional denoising network, and the acoustic propagation model is combined with the angular spectrum method to gradually reconstruct the source surface complex sound field. The diffusion model is used to learn the data distribution characteristics and introduce physical constraints to achieve a stable inverse mapping process.

Benefits of technology

By relying solely on amplitude measurements, complex sound fields can be reconstructed stably and efficiently, maintaining the physical integrity of the sound field and the diversity of solutions, thus improving the accuracy and efficiency of reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122395537A_ABST
    Figure CN122395537A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of sound field reconstruction, and particularly discloses a complex sound field reconstruction method and system based on a physically constrained diffusion model, which comprises the following steps: adopting a complex sound pressure form to represent a source plane complex sound field as an original clean sample of a diffusion model; gradually injecting Gaussian noise into the source plane complex sound field to construct a forward diffusion process, and training a conditional denoising network to predict corresponding noise components under the condition of a given diffusion time step and a target plane amplitude; using the predicted noise components of the conditional denoising network to back-propagate a clean sample estimation corresponding to the current time step, reconstructing the clean sample estimation into a complex field form to obtain a source plane complex sound field estimation; introducing an angular spectrum method as a display acoustic propagation model, inputting the source plane complex sound field estimation into the acoustic propagation model for propagation calculation to obtain a reconstructed complex sound field of the target plane. The application realizes stable prediction of a source plane complex amplitude hologram by modeling a conditional distribution in a high-dimensional continuous complex field space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound field reconstruction technology, and more specifically to a method and system for reconstructing complex sound fields based on a physically constrained diffusion model. Background Technology

[0002] The ability to reconstruct target sound fields with high fidelity is of great significance in modern acoustics, supporting a wide range of applications such as particle manipulation, three-dimensional assembly, and advanced biomedical applications. Among many advanced sound field reconstruction techniques, acoustic holography (AH) is a powerful method that can encode complex sound fields into source holograms and reconstruct the target sound field when those holograms are excited. This method achieves high-precision beamforming and full-field quantization, thereby facilitating the accurate reconstruction of the encoded sound field.

[0003] Computer-generated holography (CGH) is a promising interferometric imaging technique for digitally recording and reconstructing sound fields. Iterative methods are widely used in CGH, such as the Iterative Angular Spectrum Approach (IASA) based on the Angular Spectrum Method (ASM) and Iterative Backpropagation (IB). However, these iterative algorithms are computationally complex, and the errors become particularly significant when the target acoustic hologram is complex, while also being time-consuming. In contrast, deep learning shows great potential in learning the intrinsic mathematical structure of the inverse mapping problem of sound field reconstruction. For example, Lin et al. proposed a U-net-based deep learning method for rapidly generating acoustic holograms with optimal amplitude and phase distributions. Lee et al. proposed an autoencoder-based deep learning framework for quickly and accurately generating phase-only holograms.

[0004] Depending on whether amplitude and phase information are involved in the acoustic holographic modulation process, acoustic holograms can generally be divided into several types, including complex holograms, phase-only holograms (POH), and amplitude-only holograms (AOH).

[0005] Among them, POH achieves wavefront control by modulating only the phase, while AOH uses only the amplitude distribution for modulation. Due to differences in hardware implementation difficulty and system cost, existing research mostly focuses on phase-only or amplitude-only modulation methods. However, these limited modulation methods have inherent constraints on wavefront representation capabilities.

[0006] The fixed amplitude of the POH and the discarded phase information of the AOH limit the degree of freedom of holographic modulation, making it difficult to achieve precise control and energy constraint of complex sound fields.

[0007] Under complex targets or strict constraints, such restricted modulation often introduces noise / artifacts (such as speckle and background energy leakage), thereby affecting reconstruction fidelity and sidelobe suppression.

[0008] In contrast, complex amplitude holograms simultaneously encode amplitude and phase distributions, possessing complete wavefront modulation degrees of freedom at the source surface, theoretically providing a larger solution space and stronger expressive power. This form can precisely control the sound pressure amplitude distribution and phase evolution process, demonstrating significant advantages in high-complexity sound field modulation scenarios. However, compared to constrained modulation forms, complex amplitude sound field reconstruction is more complex at the algorithmic level. On the one hand, there is a strong coupling relationship between amplitude and phase, making the inverse mapping process highly nonlinear; on the other hand, complex fields at different source surfaces may correspond to similar target amplitude distributions, leading to multiple solutions and ill-posedness in the problem. Traditional optimization methods based on iterative propagation models can solve this problem, but they typically have high computational complexity, limited convergence speed, and are sensitive to initial conditions.

[0009] Existing research focuses more on the design of POH and AOH propagation models, while efficient generative algorithms for complex amplitude holographic inverse mappings remain relatively limited. Therefore, how to provide a learning framework that can model conditional distributions in a high-dimensional continuous complex field space to achieve stable prediction of source surface complex amplitude holograms is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0010] In view of the above problems, the present invention proposes a complex sound field reconstruction method and system based on a physical constraint diffusion model, so as to overcome the above problems or at least partially solve the above problems.

[0011] To achieve the above objectives, the present invention adopts the following technical solution:

[0012] In a first aspect, the present invention provides a method for reconstructing complex sound fields based on a physically constrained diffusion model, comprising the following steps: The source surface complex sound field is represented in the form of complex sound pressure as the original clean sample of the diffusion model; Gaussian noise is gradually injected into the source surface complex sound field to construct the forward diffusion process, and a conditional denoising network is trained to predict the corresponding noise components under given diffusion time steps and target plane amplitude conditions. The noise component predicted by the conditional denoising network is used to deduce the clean sample estimate corresponding to the current time step. The clean sample estimate is then reconstructed into a complex field form to obtain the source surface complex sound field estimate. An angular spectrum method is introduced as the explicit acoustic propagation model. The source surface complex sound field estimate is input into the acoustic propagation model for propagation calculation to obtain the reconstructed complex sound field of the target plane.

[0013] Furthermore, the process of constructing the original clean sample includes: Read source surface amplitude map Phase Encoding Map Define phase The complex sound field of the source surface is represented as:

[0014] in, Represents the real part of the complex sound field at the source surface. Represents the imaginary part of the complex sound field at the source surface; real part With the imaginary part The channel-dimensional concatenation yields a tensor representation of shape (2, H, W), which serves as the original clean sample for the diffusion model. .

[0015] Furthermore, during the forward diffusion process, noisy samples are constructed at random sampling time steps t. :

[0016] in, , representing random noise sampled from a standard Gaussian distribution. This represents a multidimensional standard normal distribution with a mean of 0 and a covariance matrix of identity. , representing the cumulative retention coefficient from the initial time to step t. This represents the noise variance scheduling parameter for the s-th diffusion step, used to control the intensity of the injected noise in that diffusion step.

[0017] Furthermore, conditional denoising networks... The noise component is used as input to predict the current time step. According to noise components Inversely derive the source surface complex sound field estimate corresponding to the current time step. :

[0018] in, This represents the amplitude diagram of the desired complex sound field on the target plane that needs to be reconstructed.

[0019] Furthermore, during the training phase, the diffusion noise prediction loss is used as the base loss, and physical losses and weakly supervised regularization terms for different samples at different amplitude scales after propagation are introduced as the total loss. Its expression is:

[0020] in, This represents the predicted loss for diffuse noise. Indicates physical loss. Let represent the weakly supervised regularization loss; the expressions for each loss term are as follows:

[0021]

[0022]

[0023] in, This represents the predicted value of the noise component. This represents random noise sampled from a standard Gaussian distribution; A diagram showing the amplitude of the complex sound field on the reconstructed target plane. This represents the amplitude diagram of the desired complex sound field on the target plane that needs to be reconstructed. This indicates normalization processing; This represents a clean sample estimate. This indicates the original clean sample.

[0024] Furthermore, after obtaining the source surface complex sound field estimate, element-wise truncation is performed to restrict all elements of the source surface complex sound field estimate to the interval [-2, 2]. After calculating the gradients of the model parameters through backpropagation, global norm constraints are applied to the gradients of all parameters.

[0025] Where g represents the gradient vector of all trainable parameters of the diffusion model. Represents the gradient of all parameters Norm; using a maximum norm threshold of 1.0, when the parameter gradient... When the norm exceeds the threshold, the gradients of all parameters are scaled proportionally so that the overall norm does not exceed 1.0.

[0026] Furthermore, the process of using the angular spectrum method to estimate the complex sound field at the source surface for propagation calculations includes: A two-dimensional Fourier transform is performed on the source surface complex sound field estimation to obtain a spatial frequency domain representation; When the propagation distance in a uniform medium is z, each frequency component is multiplied by the propagation transfer function in the frequency domain to obtain the frequency domain representation after propagation; A two-dimensional inverse Fourier transform is performed on the propagated frequency domain representation to obtain the reconstructed complex sound field of the target plane.

[0027] Furthermore, the propagation transfer function The expression is:

[0028] in, The imaginary unit is used for estimating the complex sound field at the source surface. Indicates wave number, This represents the distance a sound wave travels from the source surface to the target surface. Indicates the wavelength of the sound wave; and These are the transverse spatial frequency components; The frequency domain representation after propagation is:

[0029] in, This represents the frequency domain representation after propagation. This represents the estimation of the complex sound field at the source surface; A two-dimensional inverse Fourier transform is performed on the propagated frequency domain representation to obtain the reconstructed complex sound field of the target plane. .

[0030] Furthermore, the reconstructed complex sound field of the target plane is represented as:

[0031] The amplitude of the reconstructed complex sound field of the target plane is expressed as:

[0032] in, Let represent the real part of the complex sound field reconstructed on the target plane. denoted by , j represents the imaginary part of the reconstructed complex sound field on the target plane, where j is the imaginary unit.

[0033] Secondly, the present invention provides a complex sound field reconstruction system based on a physically constrained diffusion model, which employs the above-mentioned method, including: The sample construction module is used to represent the source surface complex sound field in complex sound pressure form as the original clean sample of the diffusion model; The noise prediction module is used to progressively inject Gaussian noise into the source surface complex sound field to construct the forward diffusion process, and train the conditional denoising network to predict the corresponding noise components under given diffusion time steps and target plane amplitude measurement conditions. The source surface complex sound field estimation module is used to infer the clean sample estimate corresponding to the current time step from the noise component predicted by the conditional denoising network, and reconstruct the clean sample estimate into a complex field form to obtain the source surface complex sound field estimate. The target complex sound field estimation module is used to introduce the angular spectrum method as the explicit acoustic propagation model. The source surface complex sound field estimation is input into the acoustic propagation model for propagation calculation to obtain the reconstructed complex sound field of the target plane.

[0034] As can be seen from the above technical solution, compared with the prior art, the present invention has the following beneficial effects: 1. Sound field reconstruction, relying solely on amplitude measurement, is essentially an inverse kinematics problem. The amplitude and phase of the source surface sound field often exhibit significant uncertainties and multiple solutions. This invention uses the desired amplitude distribution of the target plane as input to the diffusion model. By learning the data distribution characteristics at different noise scales, the diffusion model can naturally describe the high-dimensional probability distribution of the source surface complex sound field and gradually shrink the solution space during the back-diffusion process. Guided by the amplitude condition of the target plane, the originally highly ill-posed inverse problem is transformed into a series of constrained denoising sub-problems at different noise levels, making the reconstruction process more stable and asymptotic.

[0035] 2. This invention formulates the sound field reconstruction problem as a conditional diffusion modeling process acting on the complex sound field of the source surface. The source surface sound field is represented in the form of complex sound pressure and modeled using a two-channel tensor composed of its real and imaginary parts. This allows the diffusion model to directly act on the complex representation of the sound field without requiring additional discretization or wrapping of the phase. This representation not only maintains the physical integrity of the sound field but also provides a continuous and stable parameter space for subsequent diffusion modeling.

[0036] 3. In the training phase, this invention constructs a forward diffusion process by progressively injecting Gaussian noise into the source surface complex sound field, and trains a conditional denoising network to predict the corresponding noise components under given diffusion time steps and target surface amplitude measurement conditions. This process does not attempt to implicitly learn sound wave propagation relationships, but rather focuses the model on characterizing the statistical prior properties of the source surface sound field in terms of spatial structure, amplitude and phase coupling relationships, thereby providing reasonable probabilistic constraints for sound field reconstruction.

[0037] 4. In the anti-diffusion stage, this invention utilizes the noise components predicted by the network to recover the source surface sound field estimate analytically, and introduces the angular spectrum method as an explicit acoustic propagation model to propagate the reconstructed source surface sound field to the target plane. By constraining the consistency between the propagated sound field amplitude and the actual measurement results, it ensures that the anti-diffusion process always satisfies the physical laws of sound wave propagation. Thus, the statistical prior provided by the diffusion model and the physical propagation mechanism characterized by the angular spectrum method are naturally integrated within the same framework, enabling the model to gradually converge to a physically reasonable sound field solution that is consistent with observations while maintaining solution diversity.

[0038] 5. This invention combines the stepwise denoising mechanism of the diffusion model with the explicit acoustic propagation model, and uses the conditional diffusion model to achieve the inverse mapping learning from the target amplitude map to the source surface complex sound pressure (or equivalent complex phase field). The angular spectrum method is embedded as a differentiable physical propagation module in the training process, thereby forming a closed-loop constraint of data-driven and physically consistent, providing a stable and physically consistent solution to the sound field reconstruction problem under amplitude-limited conditions. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0040] Figure 1 This is a flowchart of the complex sound field reconstruction method based on a physically constrained diffusion model provided in an embodiment of the present invention; Figure 2 This is a diagram illustrating the sound field reconstruction results during the diffusion sampling process provided in this embodiment of the invention; Figure 3 The curves showing the average PSNR variation of the test set under different sampling steps provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the reconstruction effect after adding noise, provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the reconstruction effect after image transformation provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the universality verification results provided in the embodiments of the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] like Figure 1 As shown in the figure, this invention discloses a method for reconstructing complex sound fields based on a physically constrained diffusion model, including the following steps: S1. The source surface complex sound field is represented in the form of complex sound pressure as the original clean sample of the diffusion model; S2. Gaussian noise is gradually injected into the source surface complex sound field to construct the forward diffusion process, and the conditional denoising network is trained to predict the corresponding noise components under the given diffusion time step and target plane amplitude conditions. S3. Use the noise components predicted by the conditional denoising network to deduce the clean sample estimate corresponding to the current time step, reconstruct the clean sample estimate into a complex field form, and obtain the source surface complex sound field estimate. S4. Introduce the angular spectrum method as the explicit acoustic propagation model, input the source surface complex sound field estimate into the acoustic propagation model for propagation calculation, and obtain the reconstructed complex sound field of the target plane.

[0043] Next, the specific implementation methods of each of the above steps will be further explained.

[0044] S1, Source Surface Re-field A two-channel real tensor representation is used. Specifically, the source surface amplitude map is read. Phase Encoding Map Define phase The complex sound field of the source surface is represented as:

[0045] in, Represents the real part of the complex sound field at the source surface. Represents the imaginary part of the complex sound field at the source surface; real part With the imaginary part The channel-dimensional concatenation yields a tensor representation of shape (2, H, W), which serves as the original clean sample for the diffusion model. .

[0046] S2, Linear training is used during the training process. The scheduling strategy has a total diffusion step count of T. During the forward diffusion process, at random sampling time steps... The following formula is used to construct noisy samples. :

[0047] in, , representing random noise sampled from a standard Gaussian distribution. This represents a multidimensional standard normal distribution with a mean of 0 and a covariance matrix of identity. , representing the cumulative retention coefficient from the initial time to step t. This represents the noise variance scheduling parameter for the s-th diffusion step, used to control the intensity of the injected noise in that diffusion step.

[0048] S3, Conditional Denoising Network The noise component is used as input to predict the current time step. According to noise components Inversely derive the source surface complex sound field estimate corresponding to the current time step. :

[0049] in, This represents the amplitude diagram of the desired complex sound field on the target plane to be reconstructed. This yields the conditional amplitude diagram. Predicting the complex sound field from the source surface.

[0050] The predicted Two-channel tensor reconstructed into complex field form That is, the source surface complex sound field estimation.

[0051] To improve the numerical stability of the training process, this paper introduces numerical truncation and gradient norm pruning mechanisms in the model prediction and backpropagation processes, respectively. After obtaining the source surface complex sound field estimate, element-wise truncation is performed to restrict all elements of the source surface complex sound field estimate to the interval [-2,2]. This step is used to suppress abnormally large values ​​generated by the diffusion model in the early stage of training, so as to avoid numerical instability or abnormal gradient amplification problems in subsequent physical propagation and loss calculation.

[0052] After calculating the gradients of the model parameters through backpropagation, a global norm constraint is applied to the gradients of all parameters:

[0053] Where g represents the gradient vector of all trainable parameters of the diffusion model. Represents the gradient of all parameters Norm; using a maximum norm threshold of 1.0, when the parameter gradient... When the norm exceeds the threshold, the gradients of all parameters are scaled proportionally to ensure that the overall norm does not exceed 1.0, thereby effectively preventing the gradient explosion problem.

[0054] Specifically, the conditional denoising network employs a lightweight conditional U-Net structure. The input consists of two channels of a noisy complex field and the desired sound field amplitude distribution on the target plane. (Conditional Amplitude Diagram) A three-channel tensor after concatenation of single-channel features. Time step t is mapped to a temporal embedding vector via sinusoidal position encoding, and then injected additively into convolutional modules at each scale after two MLP layers. The network consists of a symmetrical encoder-decoder structure. The encoder layer contains three downsampling modules with 64, 128, and 256 channels respectively. Each convolutional block consists of two 3×3 convolutions, using GroupNorm and SiLU activation functions. Downsampling uses 2×2 average pooling, and upsampling uses nearest-neighbor interpolation. The intermediate bottleneck layer maintains 256 channels to enhance global feature modeling capabilities. At the decoder, after each upsampling, features are fused with the corresponding encoder layer features via skip connections, then processed by convolutional blocks to recover spatial details. Finally, a 1×1 convolution is used to map the output to a two-channel noise prediction. This structure achieves adaptive denoising at different noise stages while maintaining low computational complexity through multi-scale feature fusion and time-step modulation.

[0055] Specifically, through the synergistic effect of multi-scale feature fusion and temporal step modulation, the network can adaptively adjust its feature extraction and reconstruction strategies at different noise levels: in high-noise stages, it relies more on global structural information for coarse recovery, while in low-noise stages, it focuses on fine reconstruction of local details. This allows the network structure to achieve adaptive denoising at different noise levels while maintaining low computational complexity through multi-scale feature fusion and temporal step modulation.

[0056] S4. The process of using the angular spectrum method to estimate the complex sound field at the source surface and perform propagation calculations includes: 1) The source surface complex sound field estimation is expressed as: ,in, Represents the spatial coordinates of the source surface. The real part of the source surface complex sound field estimation is represented by the equation. Denotes the imaginary part of the source surface complex sound field estimation. The imaginary unit satisfies .

[0057] 2) Perform a two-dimensional Fourier transform on the source surface complex sound field estimate to obtain the spatial frequency domain representation:

[0058] in, and These are the transverse spatial frequency components. The spatial frequency domain representation of the source surface complex sound field estimation obtained by two-dimensional Fourier transform.

[0059] 3) When the propagation distance in a uniform medium is z, each frequency component is multiplied by the propagation transfer function in the frequency domain to obtain the frequency domain representation after propagation; Among them, the propagation transfer function Its expression is:

[0060] Indicates wave number, This represents the distance a sound wave travels from the source surface to the target surface. The wavelength of a sound wave is determined by both the medium parameters and the frequency.

[0061] in, To ensure the physical rationality of the transmission, for satisfying The frequency components, corresponding to the evanescent wave components, exhibit exponential decay during propagation. In the numerical implementation of this invention, this type of component is masked, i.e., its propagation coefficient is set to zero, thus retaining only the propagable frequency components. The specific expression is as follows:

[0062] The propagated frequency domain representation is as follows:

[0063] in, This represents the frequency domain representation after propagation. This represents the estimation of the complex sound field at the source surface; A two-dimensional inverse Fourier transform is performed on the propagated frequency domain representation to obtain the reconstructed complex sound field of the target plane. .

[0064] The reconstructed complex sound field of the target plane is represented as:

[0065] The amplitude of the reconstructed complex sound field of the target plane is expressed as:

[0066] in, Let represent the real part of the complex sound field reconstructed on the target plane. denoted by , j represents the imaginary part of the reconstructed complex sound field on the target plane, where j is the imaginary unit.

[0067] During the training phase, loss is predicted using diffused noise. This basic loss is used to train the network to accurately predict the Gaussian noise introduced during forward diffusion at any diffusion time step t. By minimizing this loss, the network can learn the conditional probability distribution. This gradually transforms random noise into a source surface complex field representation that conforms to the target amplitude constraint during the backdiffusion process. Therefore, It is the core driving force for the model to acquire generative capabilities, and also the probabilistic modeling foundation of the entire physical constraint diffusion framework.

[0068] Based on diffusion training, and considering that different samples may have different amplitude scales after propagation, the training process... and All samples were normalized using a sample-by-sample maximum value method before the physical loss was calculated. This allows us to directly optimize by ensuring the reconstructed amplitude after propagation approaches the target amplitude, thus constructing a backpropagation-enabled physical closed-loop constraint. Simultaneously, to utilize the real source surface complex field information in the dataset and ensure stable training, a weakly supervised regularization term is added. This is used to encourage the generation of complex field distributions to be consistent with the data prior. The final total loss The function expression is:

[0069] in, This represents the predicted loss for diffuse noise. Indicates physical loss. This indicates the loss from weak supervision regularization; This represents the predicted value of the noise component. This represents random noise sampled from a standard Gaussian distribution; A diagram showing the amplitude of the complex sound field on the reconstructed target plane. This represents the amplitude diagram of the desired complex sound field on the target plane that needs to be reconstructed. This indicates normalization processing; This represents a clean sample estimate. This indicates the original clean sample.

[0070] During the inference phase, the diffusion model is shown in the conditional magnitude diagram. Guided by this approach, the complex field representation is initialized from random Gaussian noise, and the source surface complex field is gradually generated through multi-step backdiffusion. Unlike the fixed-terminal-step output, the generation results of each step are retained during the sampling process, and the corresponding reconstructed amplitude map is obtained by propagating the complex field generated in each step using the angular spectrum method (ASM). Then, the target amplitude map is used as the basis for the reconstruction. For reference, PSNR is calculated as the evaluation criterion, and the time step with the optimal PSNR is selected as the final output. This optimal step selection strategy based on stepwise evaluation can achieve more stable reconstruction results with fewer sampling steps, and effectively balance generation quality and inference efficiency, realizing stable and efficient prediction for complex sound field reconstruction tasks.

[0071] In other embodiments, the present invention provides a complex sound field reconstruction system based on a physically constrained diffusion model, which employs the method described above, including: The sample construction module is used to represent the source surface complex sound field in complex sound pressure form as the original clean sample of the diffusion model; The noise prediction module is used to progressively inject Gaussian noise into the source surface complex sound field to construct the forward diffusion process, and train the conditional denoising network to predict the corresponding noise components under given diffusion time steps and target plane amplitude measurement conditions. The source surface complex sound field estimation module is used to infer the clean sample estimate corresponding to the current time step from the noise component predicted by the conditional denoising network, and reconstruct the clean sample estimate into a complex field form to obtain the source surface complex sound field estimate. The target complex sound field estimation module is used to introduce the angular spectrum method as the explicit acoustic propagation model. The source surface complex sound field estimation is input into the acoustic propagation model for propagation calculation to obtain the reconstructed complex sound field of the target plane.

[0072] The performance of the method of the present invention will be further verified by experiments below.

[0073] 1) Dataset Construction This experiment uses the MNIST image dataset to construct the data input for the sound field reconstruction task. 60,000 MNIST handwritten digit images were uniformly scaled to a fixed spatial resolution, and the original grayscale values ​​(0–255) were normalized and mapped to the amplitude distribution of the sound field on the target surface. In this mapping process, grayscale values ​​are first linearly normalized to the [0,1] interval to represent amplitude intensity, and can be further extended to a complex sound field representation, where the amplitude range is set to [0,1] and the phase range is mapped to [0,2π]. In this way, the original two-dimensional image is converted into the sound field distribution of the target surface, thus transforming traditional image data into physical input conditions for the sound field reconstruction problem. Subsequently, the amplitude of the target surface is established using a sound field propagation model. Complex sound field with source surface The mapping relationship between them is used to construct the training method for the diffusion model. - Data pairs. These data pairs are divided into a training set, a validation set, and a test set in an 8:1:1 ratio.

[0074] In addition to the MNIST dataset, this experiment also used the FashionMNIST and Char74K datasets to perform the above image-sound field mapping processing to test the generality of the model.

[0075] The dataset constructed in this invention is a simulated dataset generated based on a high-fidelity solver, and its labels are derived from the output of a validated high-precision model. This type of data construction has a clear theoretical basis in inverse problem and computational imaging research: when the generative model has a sufficient ability to approximate the object's solution, its output can be regarded as a high-quality approximation of the true solution space, thereby providing a stable and consistent training signal for supervised learning.

[0076] Furthermore, the diffusion model proposed in this invention does not simply learn the deterministic mapping relationship of the aforementioned models, but rather models A within the framework of conditional diffusion. E To A S The conditional probability distribution p(A) S | A E This method characterizes the statistical structure of the solution space of a complex sound field through a stepwise denoising process. From a statistical learning perspective, as long as the supervised labels approximate the real physical properties within a controllable error range, the model training can converge to an effective mapping consistent with the laws of physical propagation. Furthermore, due to the introduction of differentiable ASM physical consistency constraints during training, the model optimization objective not only depends on the labels but is also directly constrained by the laws of physical propagation, thereby reducing the bias accumulation that may be caused by label errors. Therefore, the simulated complex sound field data generated based on the high-precision model is theoretically fully reasonable and can be used as reliable supervised training data.

[0077] 2) Determination of evaluation indicators This experiment uses PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index Measure), Global Correct (GC), and Dice Coefficient as evaluation metrics.

[0078] PSNR, or pixel-level error, measures the difference between the reconstructed and real sound fields and is a commonly used quality assessment metric in image reconstruction and signal restoration tasks. It is calculated by comparing the mean square error (MSE) between the reconstructed result and the reference ground truth, reflecting the degree of distortion between the reconstructed and original signals. The formula for calculating PSNR is:

[0079] Where MAX represents the maximum possible value of the signal (e.g., 1 or 255 in a normalized image), and MSE is the mean square error between the reconstructed result and the true value. The higher the PSNR value, the smaller the difference between the reconstructed sound field and the true sound field, and the better the reconstruction quality.

[0080] SSIM is used to evaluate the similarity of the reconstructed sound field to the real sound field in terms of structural information. Compared to PSNR, which only focuses on pixel errors, SSIM pays more attention to the preservation of brightness, contrast, and structural features. SSIM measures structural consistency by comparing the statistical properties of two images in local regions. Its calculation formula is as follows:

[0081] Global Correct is used to evaluate the consistency of the reconstructed sound field across its global distribution, primarily measuring the degree to which the reconstructed result matches the real sound field in terms of overall energy distribution or spatial pattern. This metric is typically obtained by calculating the overall matching ratio between the reconstructed and real results, such as the proportion of pixels in the statistical prediction that match the real labels. Its basic form can be expressed as:

[0082] in This indicates the number of pixels that were correctly predicted. This represents the total number of pixels. A higher GC value indicates that the reconstructed sound field is more consistent with the real sound field in terms of overall distribution.

[0083] The Dice coefficient is a commonly used metric for binary segmentation or evaluating region overlap, measuring the degree of overlap between the predicted and ground truth regions. In sound field reconstruction tasks, the region distribution is typically obtained by thresholding the sound pressure amplitude, and then the overlap between the predicted and ground truth regions is calculated. The Dice coefficient is defined as:

[0084] Where A represents the set of predicted regions, B represents the set of ground truth regions, and |A|B| represents the number of pixels in the overlapping area. The Dice coefficient ranges from 0 to 1, with a value closer to 1 indicating a higher degree of overlap between the predicted and ground truth regions.

[0085] 3) Experimental platform and parameters The experiment was conducted on a Windows 10 operating system-based computing platform. The software environment used Python 3.12 and the PyTorch deep learning framework for model building and training. In terms of hardware configuration, the computing platform was equipped with an AMD Ryzen 77840H processor, an NVIDIA GeForce RTX 4060 graphics processing unit (GPU), and 32 GB of system memory to support the training and inference computations of the diffusion model.

[0086] In the experiment, the model was trained for 500 epochs, with a diffusion timestep of 500. A batch size of 16 was used during training to balance computational efficiency and memory usage. Model parameters were optimized using stochastic gradient descent (SGD). The overall loss function consists of three parts: diffusion loss, physical consistency loss, and regularization term, with weights set as follows: , and = This weighting arrangement makes diffusion learning and physical propagation constraints quite important in the optimization process, while the regularization term is used to provide a certain amount of stabilization.

[0087] 4) Experimental Results ① Validity verification results: The reconstruction results for different reasoning steps are shown as follows Figure 2 As shown, Figure 2 In the middle, the left side shows the desired target amplitude distribution. The right side shows the reconstructed sound field amplitude distribution obtained at different sampling steps (step 1–10) during the diffusion model sampling process. Each row corresponds to a different target sample. Figure 2 The step that yielded the maximum PSNR during sampling is marked, and the corresponding PSNR value is given. It can be observed that the model gradually forms the target structure in the early stages of sampling, but continuing sampling after reaching the optimal step leads to a gradual degradation of the results and a tendency towards noise distribution. The color bars on the right represent the numerical range of phase and amplitude, respectively.

[0088] To verify the effectiveness of the diffusion model proposed in this invention in the sound field reconstruction task, this invention selects multiple samples from the MNIST dataset as the target amplitude distribution. ), and use the trained model to reconstruct the sound field. Figure 3 The results of the reconstructed sound field are shown at different sampling steps during the diffusion sampling process. It can be observed that the model is able to gradually recover the target structure in the initial sampling stage, and reaches the optimal reconstruction effect (marked as the maximum PSNR) at a certain sampling step (Step=6). As sampling continues, the reconstruction result gradually deviates from the optimal solution and tends towards a random noise distribution. This indicates that the method of the present invention can effectively learn the target sound field distribution and gradually approach the true solution during the sampling process, thus verifying the feasibility and effectiveness of the diffusion model in the sound field reconstruction problem.

[0089] To further analyze the sampling behavior of the diffusion model during the inference phase, the average PSNR of the test set was calculated at different sampling steps, such as... Figure 3As shown in the figure, a skip-sampling strategy was employed in the experiment, with the number of sampling steps ranging from 1 to 10. The results show that the model reconstruction quality gradually improves with the increase in the number of sampling steps, reaching the highest average PSNR (23.01 dB) in step 6. As sampling continues, the average PSNR begins to decrease, indicating that continuing iteration after reaching the optimal sampling step will cause the result to gradually deviate from the optimal solution. This phenomenon is consistent with the inverse denoising process of the diffusion model; therefore, in actual inference, selecting the sampling step corresponding to the maximum PSNR as the final output can yield better sound field reconstruction results.

[0090] ② Ablation experiment verification results: To analyze the impact of each component of the loss function on model performance, a systematic ablation experiment was conducted on different weight parameters in the loss function. Since multiple loss terms exist simultaneously, it would be difficult to determine the impact of a single factor on the result if multiple weights were changed at the same time. Therefore, the experiment adopted the method of controlling variables: while keeping other weights fixed, only the weight of a certain term was changed, thereby analyzing the independent effect of that loss term on model performance.

[0091] First, in fixed =1 and In the case of =1e-4, the physical constraint weights Ablation experiments were conducted, and the results are shown in Table 1. By keeping the diffusion loss and regularization term constant, the effect of changes in the strength of the physical constraint term on the model's reconstruction capability can be directly observed. Experimental results show that when... When the weight is too large, the physical constraints occupy too high a proportion in the optimization process, which may limit the expressive power of the diffusion model; when the weight is too small, the physical propagation consistency constraints are insufficient, causing the reconstruction results to deviate from the true sound field distribution, thus causing a performance decrease.

[0092] Table 1 Weights of different physical constraints Impact on reconstruction performance

[0093] In fixed and Under the condition of regularization weights Ablation experiments were conducted, and the results are shown in Table 2. By keeping the diffusion loss and physical constraints constant, the impact of the regularization term on training stability and model generalization ability can be analyzed separately. Experimental results show that moderate regularization can improve model performance when... When regularization is too weak, the model achieves the highest PSNR; however, when regularization is too weak, the constraints on the model parameters are insufficient, which can easily lead to instability or overfitting. When regularization is too strong, it will inhibit the model's learning ability and also reduce reconstruction performance.

[0094] Table 2. Weights for different regularization methods Impact on reconstruction performance

[0095] During the experiment, it was also observed that, regardless of the settings, for Increasing or decreasing the diffusion loss leads to a performance decrease. This indicates that the diffusion loss plays a fundamental and dominant role in the overall optimization; therefore, it was fixed at a certain value in subsequent experiments. The experimental results above show that a proper weight balance needs to be maintained among diffusion loss, physical constraints, and regularization terms in order for the model to achieve good sound field reconstruction performance while ensuring physical consistency.

[0096] ③ Robustness verification: To evaluate the stability and robustness of the method of the present invention under noise interference conditions, a robustness experiment was further conducted. The experiment first involved the target amplitude... The perturbed target distribution is obtained by adding Gaussian noise. This data is then used as input to the model for sound field reconstruction. The reconstruction results are as follows: Figure 4 As shown, (a) represents the original target amplitude. (b) is the target amplitude after adding noise. (c) and (d) are the source surface amplitudes reconstructed using deep learning methods, respectively. Phase of source surface (e) represents the target surface sound field obtained by propagation from the reconstructed source field. It can be observed that, despite the noise disturbance affecting the input target amplitude, the model is still able to recover the source surface amplitude and phase distribution with a clear structure, and reconstruct a sound field morphology that is very consistent with the original target after propagation to the target surface. This indicates that the method of the present invention still has good reconstruction capability and stability under certain noise interference.

[0097] To further evaluate the model's adaptability to changes in target shape, a robustness experiment was designed to test the target amplitude rotation and scaling. In this experiment, the original target amplitude was... Rotation and scaling transformations were performed to obtain a new target distribution, which was then used as input to the model for sound field reconstruction. Experimental results are as follows: Figure 5 As shown, (a) represents the target amplitude after rotation and scaling transformation. (b) and (c) are the source surface amplitudes reconstructed using deep learning methods, respectively. Phase of source surface (d) represents the target surface sound field obtained by propagation from the reconstructed source field. It can be observed that despite the geometric transformation of the input target, the model can still generate a reasonable source surface amplitude and phase distribution, and reconstruct a sound field structure on the target surface that is consistent with the transformed target shape. This result demonstrates that the proposed method has good adaptability to rotation and scale changes in the target distribution, exhibiting a certain degree of geometric robustness.

[0098] To verify the generality of the method of this invention, further cross-dataset experiments were conducted. Specifically, the model was trained only on the MNIST dataset, while in the testing phase, it was evaluated using other datasets not used in the training, including FashionMNIST and Char74k. The experimental results are as follows: Figure 6 As shown, (a) is the target amplitude plot from FashionMNIST and Char74k. (b) and (c) are the source surface amplitudes reconstructed using deep learning methods, respectively. Phase of source surface (d) represents the target surface sound field obtained by propagation from the reconstructed source field. The values ​​labeled in the figure represent the corresponding PSNR. It can be observed that even on unseen data distributions, the model can still generate reasonable source surface amplitude and phase, and reconstruct a sound field morphology consistent with the target structure on the target surface, while maintaining a good PSNR. Calculations show that the model achieves a PSNR of 21.14 on the FashionMNIST dataset and 22.03 on the Char74k dataset. This indicates that the proposed method not only achieves good results on the training data but also has a certain degree of cross-dataset generalization ability, demonstrating good versatility.

[0099] 5) Comparative Experiment To further verify the performance advantages of the method of this invention, comparative experiments were conducted with several representative sound field reconstruction methods, including the traditional iterative method IASA, U-Net based on convolutional neural networks, GAN based on adversarial learning, and AHR-Net, a deep learning method proposed by Ma et al. The comparison results are shown in Table 3, and the evaluation metrics include PSNR, SSIM, Dice Coefficient, and average inference time (Ave Time).

[0100] Table 3 Comparison Results

[0101] From the perspective of reconstruction quality metrics, the method of this invention achieved optimal results across multiple evaluation indicators. Specifically, the PSNR reached 23.01 dB, significantly higher than other methods; in the structural similarity index (SSIM), AcousDiffusion achieved 0.47, outperforming some baseline methods overall; and the Dice coefficient reached 0.96, indicating a higher consistency between the reconstructed sound field and the target distribution in terms of structural morphology. This demonstrates that the diffusion model has a stronger expressive ability in learning complex sound field distributions, thus enabling the generation of higher-quality reconstruction results.

[0102] In terms of computational efficiency, the traditional iterative method IASA requires multiple physical propagation calculations, resulting in significantly longer inference time compared to deep learning-based methods. In contrast, deep learning methods can quickly obtain results with a single forward inference. Our proposed method, using 10 sampling steps, has an average inference time of 0.052 ms, maintaining high computational efficiency while significantly improving reconstruction quality. Considering both reconstruction accuracy and computational efficiency, our method demonstrates superior overall performance in sound field reconstruction tasks.

[0103] This invention addresses the problem of inverse mapping from target amplitude to source surface complex sound field, proposing a physically constrained diffusion framework, AcousDiffusion, to achieve high-precision inverse mapping from the target sound field to the source surface complex amplitude hologram. To address the ill-posedness problem caused by amplitude and phase coupling during sound field reconstruction, this invention combines the probabilistic generation capability of the diffusion model with an explicit acoustic propagation model. By embedding the angle spectral method (ASM) during training to construct differentiable physical constraints, the model learns the statistical prior of the source surface complex field while satisfying the laws of sound wave propagation, thus achieving end-to-end physically consistent sound field reconstruction. Methodologically, this invention utilizes conditional DDPM to establish a hologram from the target amplitude map... To the source surface recovery field The generative model is constructed, and a two-channel complex field representation is used to directly model the complex sound pressure, avoiding the information loss caused by phase discretization. Loss is predicted by jointly optimizing the diffusion noise. Physical consistency loss based on ASM and weakly supervised regularization terms The model achieves an effective balance between generative capabilities and physical constraints. Furthermore, an optimal sampling step selection strategy based on PSNR is introduced during the inference phase, resulting in stable and high-quality reconstruction results with fewer sampling steps, thus balancing reconstruction accuracy and inference efficiency.

[0104] Simulation experiments demonstrate that the method of this invention can achieve high-fidelity reconstruction of complex sound field structures, exhibiting excellent performance in energy focusing, background suppression, and structure preservation, and possessing outstanding robustness and versatility. Compared with existing methods, AcousDiffusion achieves remarkable results in both reconstruction accuracy and efficiency.

[0105] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0106] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for reconstructing complex sound fields based on a physically constrained diffusion model, characterized in that, Includes the following steps: The source surface complex sound field is represented in the form of complex sound pressure as the original clean sample of the diffusion model; Gaussian noise is gradually injected into the source surface complex sound field to construct the forward diffusion process, and a conditional denoising network is trained to predict the corresponding noise components under given diffusion time steps and target plane amplitude conditions. The noise component predicted by the conditional denoising network is used to deduce the clean sample estimate corresponding to the current time step. The clean sample estimate is then reconstructed into a complex field form to obtain the source surface complex sound field estimate. An angular spectrum method is introduced as the explicit acoustic propagation model. The source surface complex sound field estimate is input into the acoustic propagation model for propagation calculation to obtain the reconstructed complex sound field of the target plane.

2. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 1, characterized in that, The process of constructing a clean, original sample includes: Read source surface amplitude map Phase Encoding Map Define phase The complex sound field of the source surface is represented as: in, Represents the real part of the complex sound field at the source surface. Represents the imaginary part of the complex sound field at the source surface; real part With the imaginary part The channel-dimensional concatenation yields a tensor representation of shape (2, H, W), which serves as the original clean sample for the diffusion model. .

3. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 1, characterized in that, During forward diffusion, noisy samples are constructed at random sampling time step t. : in, , representing random noise sampled from a standard Gaussian distribution. This represents a multidimensional standard normal distribution with a mean of 0 and a covariance matrix of identity. , representing the cumulative retention coefficient from the initial time to step t. This represents the noise variance scheduling parameter for the s-th diffusion step, used to control the intensity of the injected noise in that diffusion step.

4. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 1, characterized in that, Conditional denoising networks The noise component is used as input to predict the current time step. According to noise components Inversely derive the source surface complex sound field estimate corresponding to the current time step. : in, This represents the amplitude diagram of the desired complex sound field on the target plane that needs to be reconstructed.

5. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 3, characterized in that, During the training phase, the diffusion noise prediction loss is used as the base loss. Physical losses and weakly supervised regularization terms for different samples at post-propagation amplitude scales are introduced as the total loss. Its expression is: in, This represents the predicted loss for diffuse noise. Indicates physical loss. Let represent the weakly supervised regularization loss; the expressions for each loss term are as follows: in, This represents the predicted value of the noise component. This represents random noise sampled from a standard Gaussian distribution; A diagram showing the amplitude of the complex sound field on the reconstructed target plane. This represents the amplitude diagram of the desired complex sound field on the target plane that needs to be reconstructed. This indicates normalization processing; This represents a clean sample estimate. This indicates the original clean sample.

6. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 1, characterized in that, After obtaining the source surface complex sound field estimate, element-wise truncation is performed to restrict all elements of the source surface complex sound field estimate to the interval [-2, 2]. After calculating the gradients of the model parameters through backpropagation, global norm constraints are applied to the gradients of all parameters. Where g represents the gradient vector of all trainable parameters of the diffusion model. Represents the gradient of all parameters Norm; using a maximum norm threshold of 1.0, when the parameter gradient... When the norm exceeds the threshold, the gradients of all parameters are scaled proportionally so that the overall norm does not exceed 1.

0.

7. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 1, characterized in that, The process of using the angular spectrum method to estimate the complex sound field at the source surface and perform propagation calculations includes: A two-dimensional Fourier transform is performed on the source surface complex sound field estimation to obtain a spatial frequency domain representation; When the propagation distance in a uniform medium is z, each frequency component is multiplied by the propagation transfer function in the frequency domain to obtain the frequency domain representation after propagation; A two-dimensional inverse Fourier transform is performed on the propagated frequency domain representation to obtain the reconstructed complex sound field of the target plane.

8. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 7, characterized in that, Propagation transfer function The expression is: in, The imaginary unit is used for estimating the complex sound field at the source surface. Indicates wave number, This represents the distance a sound wave travels from the source surface to the target surface. Indicates the wavelength of the sound wave; and These are the transverse spatial frequency components; The frequency domain representation after propagation is: in, This represents the frequency domain representation after propagation. This represents the estimation of the complex sound field at the source surface; A two-dimensional inverse Fourier transform is performed on the propagated frequency domain representation to obtain the reconstructed complex sound field of the target plane. .

9. The complex sound field reconstruction method based on a physically constrained diffusion model as described in claim 7, characterized in that, The reconstructed complex sound field of the target plane is represented as: The amplitude of the reconstructed complex sound field of the target plane is expressed as: in, Let represent the real part of the complex sound field reconstructed on the target plane. denoted by , j represents the imaginary part of the reconstructed complex sound field on the target plane, where j is the imaginary unit.

10. A complex sound field reconstruction system based on a physically constrained diffusion model, characterized in that, The complex sound field reconstruction method based on the physical constraint diffusion model as described in any one of claims 1-9 includes: The sample construction module is used to represent the source surface complex sound field in complex sound pressure form as the original clean sample of the diffusion model; The noise prediction module is used to progressively inject Gaussian noise into the source surface complex sound field to construct the forward diffusion process, and train the conditional denoising network to predict the corresponding noise components under given diffusion time steps and target plane amplitude measurement conditions. The source surface complex sound field estimation module is used to infer the clean sample estimate corresponding to the current time step from the noise component predicted by the conditional denoising network, and reconstruct the clean sample estimate into a complex field form to obtain the source surface complex sound field estimate. The target complex sound field estimation module is used to introduce the angular spectrum method as the explicit acoustic propagation model. The source surface complex sound field estimation is input into the acoustic propagation model for propagation calculation to obtain the reconstructed complex sound field of the target plane.