Dynamic scattering medium imaging method and system based on polynomial denoising diffusion probability model
By using the polynomial denoising diffusion probability model and the Brownian bridge diffusion model in dynamic scattering medium imaging, the problem of real-time imaging of dynamic scattering medium in the prior art is solved, and an efficient, accurate and physically interpretable imaging effect is achieved.
Patent Information
- Application Number
- CN202510063273.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to achieve real-time imaging of dynamic scattering media, especially in terms of resource consumption and physical interpretability.
Using a method based on the polynomial denoising diffusion probability model, the scattering phenomenon of photons passing through the fog chamber is modeled through the Brownian bridge diffusion model and diffusion equation, and a special form of polynomial neural network is constructed to improve imaging accuracy and interpretability.
It realizes efficient imaging of dynamic scattering media, improves imaging accuracy and physical interpretability, and reduces resource consumption.
Smart Images

Figure CN119991845A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of optical imaging, and specifically relates to a method and system for imaging a dynamic scattering medium based on a polynomial denoising diffusion probability model. Background Art
[0002] Imaging through dynamic scattering media is one of the most challenging but attractive problems in optics, with applications ranging from biodetection to remote sensing. In early studies, scattered photons were considered to be noise that affects imaging quality, so how to extract ballistic photons and serpentine photons from a large number of scattered photons constituted the main imaging techniques, such as time gates, space gates, coherence gates, and polarization gates. Subsequent researchers found that even scattered photons that have undergone multiple scattering still carry a large amount of target information. Therefore, how to use a large number of scattered photons for computational imaging has become a new research direction. Wavefront shaping and speckle correlation have been proven to be two effective imaging methods. Wavefront shaping adjusts the amplitude and phase of the incident light field to achieve focusing of the outgoing light field. The main technologies related to wavefront shaping include feedback-based wavefront shaping, transfer matrix, and optical phase conjugation. Speckle correlation is based on optical memory effect imaging.
[0003] In recent years, many researchers have applied deep learning to related fields of scattering medium imaging due to its low cost and powerful feature extraction capabilities. The inverse process of light scattering is pathological and difficult to model directly, while deep learning directly treats this inverse process as a black box, using neural networks to mine the intrinsic connections between massive data and learn the mapping relationship between the input and output of scattering media. In 2016, Horisaki et al. realized facial image reconstruction based on deep learning. In 2018, Li et al. realized speckle reconstruction through frosted glass by using convolutional neural networks (CNN). In 2019, a scheme based on hybrid neural networks was proposed to achieve the task of reconstructing the target image behind 3 mm thick white polystyrene. By setting five discrete fat emulsion solutions with concentration gradients, Sun et al. proposed a reconstruction algorithm based on classification-based generative adversarial networks (GANs), which effectively enhanced the generalization of neural networks. Liu et al. discussed in depth how to obtain more realistic data, and adjusted the neural network structure according to experimental results, so that the model has better generalization ability. Chen et al. used a scoring function to predict the intensity distribution of the image as a priori in the iterative process of the diffusion model, which effectively improved the performance of the model.
[0004] At present, most scattering imaging-related tasks are concentrated on stable scattering media, such as frosted glass, polystyrene materials, and scattering plates. Most imaging technologies are difficult to apply to the requirements of real-time imaging of time-varying scattering media such as smoke, clouds, and fog. Deep learning, with its advantages of low cost, imaging quality, and imaging speed, meets the requirements for real-time imaging very well. However, the mainstream neural networks currently used in deep learning are often over-parameterized models (hundreds of millions of parameters), which consumes a huge amount of resources. At the same time, the neural network is a black box model, and end-to-end imaging also makes the imaging process lack its physical interpretability.
[0005] In summary, for the problem of optical imaging of dynamic scattering media, the use of deep learning-based methods still has significant limitations. Summary of the invention
[0006] In order to overcome the above defects, this application proposes a dynamic scattering medium imaging method based on a polynomial denoising diffusion probability model, including:
[0007] The speckle pattern is iteratively sampled based on the Brownian bridge diffusion model to obtain the final target image;
[0008] In the iterative sampling process, the variance table is constructed through the diffusion equation, and the current image is input into the trained polynomial neural network to obtain the noise for the next iteration.
[0009] As an improvement of the above method, the formula for iterative sampling is:
[0010]
[0011] Among them, x t-1 represents the image generated at the t-1th iteration; x T represents the speckle pattern; ∈ θ (x t ,t) represents a polynomial neural network; c xt 、c xT and c ∈t is x in the Brownian bridge diffusion model t 、x T and∈ θ (x t ,t) proportion; represents the variance term in the Brownian bridge diffusion model; z represents the solution of the diffusion equation.
[0012] As an improvement of the above method, the diffusion equation is:
[0013]
[0014] Where D is the diffusion coefficient, r is the scattering distance, and h is the time when the scattering occurs.
[0015] As an improvement of the above method, the construction process of the polynomial neural network includes:
[0016] Use the quadratic activation function instead of the convolutional neural network activation function.
[0017] As an improvement of the above method, the construction process of the polynomial neural network includes:
[0018] A 3×3 convolutional layer, a quadratic activation function, and a layer normalization layer are used to replace the activation function of the convolutional neural network.
[0019] As an improvement of the above method, the convolutional neural network is a U-net neural network.
[0020] The present application also provides a dynamic scattering medium imaging system based on a polynomial denoising diffusion probability model, which is implemented based on the above method. The system includes:
[0021] An iterative sampling module is used to iteratively sample the speckle pattern based on the Brownian bridge diffusion model to obtain the final target image;
[0022] The noise prediction module, during the iterative sampling process, constructs a variance table through the diffusion equation, inputs the current image into the trained polynomial neural network, and obtains the noise for the next iteration.
[0023] Compared with the prior art, the advantages of this application are:
[0024] 1. This application models the scattering phenomenon of photons passing through a fog chamber based on the Brownian bridge diffusion model and the diffusion equation. Combined with the diffusion coefficient of the scattering medium, the model can be adjusted according to the different scattering conditions of the scattering medium. This application combines the diffusion model with scattering imaging, making the diffusion process more consistent with the physical properties of scattering, and making the imaging process more physically explainable rather than a purely random process. At the same time, the introduction of the diffusion coefficient can dynamically adjust the model according to the specific scene, which is more flexible;
[0025] 2. The present invention constructs a neural network by constructing a special form of quadratic activation function to ensure that the output is a polynomial expression of the input, thereby improving the imaging accuracy and further increasing the interpretability of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 The figure shows a schematic diagram of the optical diffusion process; among them, forward diffusion: forward diffusion; reverse diffusion: reverse diffusion;
[0027] Figure 2 The figure shows a noise prediction network schematic diagram; where polynomial function: polynomial function; noisecomponent: noise component;
[0028] Figure 3(a) shows a schematic diagram of a two-layer fully connected network;
[0029] FIG3( b ) shows a network function image obtained by fitting y using a neural network with ReLU as an activation function; wherein, ReLUActivation Function: ReLU activation function;
[0030] FIG3(c) shows a network function image obtained by fitting y with a neural network using a square function as an activation function; wherein, SquareActivation Function: square activation function;
[0031] Figure 4 The figure shows the neural network architecture diagram; Convolution: convolution; Layer normalization: layer normalization;
[0032] Figure 5 Shown is the polynomial base module (BackBone);
[0033] Figure 6 The figure shows the optical path for collecting images of a time-varying scattering medium dataset based on a fog chamber; where Seatteringspeckle: speckle pattern; Reconstructed image: reconstructed target image; Fog generator: fog generator; Fog chamber: fog chamber; Blower: blower; Original target image: original target image; CCD: charge coupled device; DMD: spatial light modulator; Laser: laser; L1, L2: lenses;
[0034] Figure 7 Shown is a transmittance histogram;
[0035] Figure 8 The figure shows the change of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) of PolyNet networks of different layers on the test set over time;
[0036] FIG9( a ) shows a speckle image to be reconstructed in the speckle reconstruction schematic diagram;
[0037] Figure 9(b) shows the reconstruction result of PolyNet in the speckle reconstruction schematic;
[0038] FIG9( c ) shows the target image in the speckle reconstruction schematic diagram;
[0039] Fig.10 Shown are the noise statistics results, where the x-axis represents the pixel value of the speckle and the y-axis represents the number of times the noise occurs;
[0040] FIG11( a ) shows a speckle image to be reconstructed in the schematic diagram of the reconstruction result;
[0041] FIG11( b ) shows the reconstruction result of PDDPM in the reconstruction result diagram;
[0042] FIG11( c ) shows the target image in the schematic diagram of the reconstruction result;
[0043] Fig.12 Shown are histograms of reconstruction results; the left side of the figure shows the PSNR frequency distribution histogram between the speckle image and the target image, and between the reconstructed image and the target image; the right side of the figure shows the SSIM frequency distribution histogram of the speckle and target image, and between the reconstructed image and the target image. DETAILED DESCRIPTION
[0044] The technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0045] In response to the difficulty of dynamic scattering medium imaging, the present application proposes a dynamic scattering medium imaging method and system based on a polynomial denoising diffusion probability model (PDDPM). The present application models the scattering phenomenon of photons passing through a fog chamber based on the Brownian bridge diffusion model and the diffusion equation. The model is combined with the diffusion coefficient of the scattering medium, and the model can be adjusted according to the different scattering conditions of the scattering medium. In addition, the present invention constructs a neural network by constructing a special form of quadratic activation function to ensure that the output is a polynomial expression of the input, thereby improving the imaging accuracy while increasing the interpretability of the network. A new solution is provided for dynamic scattering imaging based on the polynomial denoising diffusion probability model (PDDPM).
[0046] 1. Imaging system based on scattering media
[0047] Imaging through time-varying scattering media is often a nonlinear and reversible problem. Deep learning is based on neural networks to learn the mapping relationship between large amounts of data, avoiding artificially designed strategies. The powerful learning ability of neural networks can well learn the mapping relationship between the imaging system and the scattering effect of the medium, and then mine information from the speckle image to reconstruct the target image. Input target image E out (x out ,y out ) and the output optical speckle Ein (x in ,y in ) can be described using the mathematical model shown in Formula 1:
[0048] E out (x out ,y out )=F(E in (x in ,y in ))#(1)
[0049] Among them, F represents the forward operator, and the input light field is converted into the output light field by the forward operator F. Due to the reversibility of the light path, the inverse operator of scattering can be obtained as:
[0050]
[0051] Among them, F -1 It represents the reverse process of the forward operator F. The key to reconstructing the target image from the optical speckle image is how to accurately model the reverse operator F. -1 With the help of a large dataset of paired optical speckle and target images, an optimization-based scheme is used to fit the inverse operator F using a neural network. -1 .
[0052] 2. Optical diffusion process based on Brownian bridge and diffusion equation
[0053] Many deep learning-based scattering media imaging schemes directly use neural networks to fit F -1 However, due to the black-box nature of neural networks, the reverse process F -1 There is less mining of internal information, and often only focuses on input and output. The diffusion model can show the intermediate steps of the imaging process. For example, for optical speckle reconstruction, the diffusion model can point out which parts of the current speckle are caused by scattering and remove them as noise. Repeating this process can reconstruct the target image. However, the classical diffusion model continuously applies noise that obeys the standard Gaussian distribution to the target image in the forward process, so that the target image can be reconstructed from the noise that obeys the standard Gaussian distribution in the reverse process. Due to the randomness of noise sampling, the diffusion model is often used as a generative model. For the task of optical speckle reconstruction, it is often necessary to reconstruct a specific speckle. Therefore, the input and output of the model are not only paired, but also specific. The process of sampling from the Gaussian distribution of the diffusion model is full of randomness, so the classical diffusion model will no longer meet the requirements. Based on the Brownian bridge, the problem that the diffusion model is difficult to achieve mapping between paired images can be effectively solved.
[0054] The present invention constructs an imaging system of smoke scattering medium based on a time-continuous random process Brownian bridge and a diffusion equation. The Brownian bridge mathematical model is shown in Formula 3.
[0055]
[0056] where t∈[0,T], x0 and x T Represents two fixed endpoints, which respectively represent the target image and the speckle image in this application. t As fixed endpoints x0 and x T A continuous random process between x0 and x T As described in the existing literature, if the variance table of formula 3 is directly used, the model will be untrainable, so a new variance table needs to be rebuilt. This application considers building a variance table through a diffusion equation. For the scattering phenomenon of photons, this application only considers the case of a strong scattering medium, and does not consider absorption for the time being. Therefore, this application uses a diffusion equation to describe the scattering phenomenon of photons. The diffusion equation is shown in formula 4.
[0057]
[0058] Where D represents the diffusion coefficient, r represents the scattering distance, and h represents the time when the scattering occurs. For time-varying scattering media such as smoke, the diffusion coefficient D(r,h) is quite complex, and it is unnecessary to consider the exact D(r,h), so in this application, the diffusion coefficient D(r,h) is regarded as a constant. The Gaussian solution of the above diffusion equation is shown in Equation 5.
[0059]
[0060] It can be seen from Formula 5 that P(r,h)~N(0,2DhI). Combining Formula 3 and Formula 5, the optical diffusion process modeled in this application can be obtained, as shown in Formula 6.
[0061] p(x t |x0,x T )=N(x t |(1-m t )x0+m t x T ,2DhI),h=m t (1-m t ).#(6)
[0062] in, As can be seen from Formula 6, the model of this application can be easily adjusted according to the diffusion coefficient of the real scattering medium. Since it is difficult to measure the diffusion coefficient for time-varying scattering media such as smoke, this application provides an approximate relationship between transmittance and diffusion coefficient based on existing literature as shown in Formula 7.
[0063]
[0064] Wherein, b represents the thickness of the scattering medium, g represents the anisotropy factor, and L represents the transmittance of the scattering medium.
[0065] As shown in Formula 8, the present application establishes a variance table based on the diffusion coefficient of the scattering medium, and introduces the physical information of the scattering medium into the model. The optical diffusion process finally established is shown in Formula 9.
[0066] δ t =2Dh,h=m t (1-m t ).#(8)
[0067] x t =(1-m t )x0+m t x T +δ t ε,∈∈N(0,I)#(9)
[0068] in, N(0,I) represents the standard Gaussian distribution, ∈ represents sampling from the standard Gaussian distribution. When t = 0, x t =x0, x0 is the target image. When t = T, x t =x T , x T is the speckle image. When t∈[0,T], x t Then it is a combination of x0 and x T The optical diffusion process of PDDPM is as follows: Figure 1 As shown in Figure 2, in the forward process, the solution of the diffusion equation is used to replace the noise that obeys the standard Gaussian distribution, making the forward process more consistent with the actual scattering phenomenon. The reverse process uses the polynomial-based U-net (∈ θ ) predicts the noise, so that the predicted noise is a polynomial expression of the image to be processed at the current moment, which enhances the interpretability of the model. The training algorithm of this application is based on the Brownian Bridge Diffusion Model (BBDM) framework, and the training process and sampling process are summarized in Algorithm 1 and Algorithm 2. Where c xt 、c xT and c ∈t From the BBDM model, they are expressed in the formula Medium t 、x T And the polynomial neural network ∈ proposed in this application θ (x t ,t) proportion, is the variance term provided by BBDM, and z is the solution of the diffusion equation proposed in this application, that is, P(r,h) in Formula 5.
[0069]
[0070] 3. Polynomial Denoising Diffusion Probability Model (PDDPM)
[0071] The optical diffusion process of PDDPM requires the use of U-net convolutional neural network to predict the noise that needs to be removed from the image at the current moment when reconstructing the optical speckle in reverse. Figure 2 As shown, however, as a black box model, the relationship between the output and input of U-net is often difficult to know. The present invention builds a polynomial-based neural network to ensure that its output is a polynomial expression of the input, and the noise predicted each time is a polynomial expression of the current image to be processed, thereby enhancing the interpretability of the model noise prediction process.
[0072] In addition, compared with classic activation functions such as ReLU, polynomials have better fitting capabilities. As shown in Figure 3, a simple two-layer fully connected network is used to fit the function y = x 3 +x 2 +x+1. As can be seen from the function graph, the Relu activation function has the property of local linearity. By dividing the function to be fitted into different intervals, different linear functions are used in different intervals to fit the function y, which fully reflects the idea of differentiation. The large number of parameters of the neural network makes this interval division sufficiently detailed and ensures the accuracy of fitting. The solution based on the Relu function ensures that the neural network has good sparsity, making the calculation simple and not easy to overfit. However, this linear division is difficult to explain from the perspective of the function, because when there are enough divided intervals, it is not known how many intervals the neural network has divided and how to divide them, so it will become infeasible to explore each interval. Using the square as the activation function can avoid this. No matter how many layers the neural network is nested, its output is a polynomial expression of the input. At the same time, the square activation function has stronger nonlinear expression ability than Relu. Replacing the activation function of the neural network in Figure 3(a) with the square function, the result is shown in Figure 3(c), which has a better fitting function y than the Relu activation function shown in Figure 3(b). However, if the square function is directly used as the activation function of the neural network, it will make the neural network difficult to optimize. Due to the properties of the square function, gradient explosion and gradient disappearance are very likely to occur during the training process of the neural network. However, this application provides a method that combines the square function with the residual structure to ensure smooth optimization of the neural network. At the same time, the residual structure is conducive to avoiding the phenomenon that polynomial differences are prone to overfitting.
[0073] The Weierstrass approximation theorem was proved by German mathematician Weierstrass in 1885 and is the starting point of the later function approximation and interpolation theory. The Weierstrass approximation theorem shows that the polynomial set exist In other words, for any ε>0 and any There exists a polynomial P such that:
[0074] sup|f(x)-P(x)|<∈,x∈[0,1].#(10)
[0075] The present invention constructs a polynomial neural network (PolyNet) based on Weierstrass approximation theorem, and its network structure diagram is as follows: Figure 4 As shown in the figure, each basic module consists of a 3×3 convolution, a quadratic activation function, and a layer normalization.
[0076] The convolution module and layer normalization do not provide nonlinear expression capabilities. Since the quadratic activation function is prone to gradient vanishing and gradient explosion, the PolyNet network becomes difficult to optimize. Layer Normalization (LN) can ensure a more stable training process. The LN calculation formula is shown in Equation 11.
[0077]
[0078] where x and Represent the input and output respectively, H represents the total number of pixels of the input image, a i represents a pixel of the input image, μ and σ represent the mean and standard deviation of the input image, respectively, and γ and β represent the adaptive gain and bias, respectively.
[0079] In fact, the quadratic activation function is to multiply the input matrix by itself, which allows the entire PolyNet to express the operations performed by the entire network through a simple formula, as shown in Formula 12.
[0080]
[0081] Where I represents the input speckle image, K i Represents the convolution kernel of the i-th layer network.
[0082] The nonlinear expression ability of the PolyNet network mainly comes from the quadratic activation function. By superimposing multiple layers of the network, a high-order polynomial can be quickly obtained. The order of the high-order polynomial is exponentially related to the number of layers of the network, as shown in Formula 13.
[0083] Order = 2L #(13)
[0084] Wherein, L represents the number of layers of PolyNet. At the same time, in addition to being the input of the next layer, the output of each layer is retained to the last layer to be combined into the final polynomial, which is similar to the residual network that can obtain a more powerful fitting ability by stacking more layers without worrying about overfitting. The present invention provides five specific forms of PolyNet networks, and their network architectures and parameters are shown in Table 1. Each layer contains a basic block, and a network with stronger fitting ability can be obtained by simple stacking, which has good scalability, and its excellent fitting ability is verified in the experimental part.
[0085] Table 1 Network architecture
[0086]
[0087] Based on the PolyNet network, in order to improve the quality of the image, the present invention also proposes a polynomial denoising diffusion probability model (PDDPM). Figure 5 As shown, the model proposes the BackBone polynomial module, which inherits the idea of PolyNet, ensures that the output is a polynomial expression of the input, and provides good portability. Therefore, it only takes two steps to replace an ordinary neural network with a polynomial neural network. First, remove all activation functions of the original neural network, which will ensure that the neural network is a linear model without nonlinear expression capabilities. Secondly, insert the BackBone proposed in this application into the neural network, and the model can be constructed as a polynomial neural network. BackBone can include a quadratic activation function, and can also include a 3×3 convolutional layer, a quadratic activation function, and a layer normalization layer. In view of the good performance of the U-net model in various tasks in image-related fields, the noise prediction part of the PDDPM model of this application is implemented using U-net based on the polynomial BackBone.
[0088] The technical effect of this application is demonstrated through experiments.
[0089] 1. Experimental setup
[0090] Data acquisition optical path Figure 6As shown. The laser with a wavelength of 532nm is expanded by an optical system composed of lenses L1 and L2 and then irradiated onto a digital micromirror device (DMD, number of pixels: 1920×1080, pixel pitch: 7.56μm). DMD is used to display the original target image. The target image is selected from the MNIST (Modified National Institute of Standards and Technology) data set. The target image passes through a fog chamber filled with water mist to form speckles. In order to make the scattering medium (fog) change more drastically, a fan that can rotate 360° is started when the fog generator enters the fog chamber. The speckles after passing through the fog chamber are collected by a charge coupled device (CCD). The present invention collects a total of 7033 pairs of images, including a test set (702 pairs), a validation set (705 pairs), and a training set (5626 pairs). In order to train the model, the network weights are updated with the L1 loss between the model output and the target. In order to facilitate the calculation of the degradation degree, only one pair of images is processed at a time. Adam is used as the default optimizer and trained for 10 cycles (epochs). The model was implemented using Pytorch (2.1.1+cu121) in Ubuntu 20.04 and accelerated by GeForce RTX 4090.
[0091] 2. Fog chamber transmittance measurement
[0092] As shown in Formula 7, the diffusion coefficient of the scattering medium can be obtained according to the transmittance. In order to quantitatively measure the transmittance of the fog chamber, the present application takes the ratio of the sum of the pixels S1 of the speckle image passing through the fog chamber to the sum of the pixel values S2 of the corresponding target image as the transmittance of the fog chamber, as shown in Formula 14.
[0093]
[0094] This application measures the transmittance of 500 groups of handwritten numbers "8" as follows Figure 7 As shown. Figure 7 It can be seen that the transmittance of the fog chamber is mainly concentrated between 0.04 and 0.07, which shows that the water mist concentration in the fog chamber is relatively high and the scattering is serious, resulting in serious loss of image information.
[0095] According to the approximate relationship between the diffusion coefficient and the transmittance in equation 7, the diffusion coefficient can be calculated according to the measured transmittance of the fog chamber as shown in Table 2. Only a few representative transmittances are selected for calculation.
[0096] Table 2 Relationship between transmittance and diffusion coefficient
[0097]
[0098] 3. Polynomial-based neural network (PolyNet network) performance verification experiment
[0099] This section uses the PolyNet network as an example to conduct experiments to verify the feasibility of polynomial-based neural networks. It shows the performance of five forms of PolyNet on the test set. Here, PolyNet is directly used to fit the inverse process of scattering F. -1 The experimental results are as follows Figure 8 As shown. From top to bottom, the test results of PolyNet networks with 10 layers to 6 layers are shown. The blue curve and the yellow curve show the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) test results of PolyNet on the test set, respectively. The vertical axis is the PSNR and SSIM values, and the horizontal axis is epoch. Each form of PolyNet was trained for 3000 epochs on the training set, and the model was saved every 500 epochs. From the experimental results, it can be seen that since the optimization process of the model is affected by various environmental factors and initialization parameters, there are slight fluctuations in the test results of the model. For example, at the 1500th epoch, the PSNR index of the 9-layer PolyNet network on the test set is lower than that of the 8-layer PolyNet network. However, as the number of network layers increases, from the overall test results, the fitting ability of the PolyNet network is gradually increasing, and the 10-layer PolyNet can already achieve good reconstruction results.
[0100] Some speckle reconstruction images are shown in Figure 2. Figure 9(a)-Figure 9(c) As shown in Figure 9(a) represents the speckle image collected in the experiment, Figure 9(b) represents the image reconstructed by the PolyNet network, and Figure 9(c) represents the target image. It can be seen from the figure that the PolyNet network can reconstruct the target image, but there are white noise points.
[0101] 4. Polynomial Denoising Diffusion Probabilistic Model (PDDPM)
[0102] The previous experiment verified the feasibility of the polynomial neural network based on PolyNet. Although PolyNet achieved good results in PSNR and SSIM on the test set, the reconstructed image had abnormal white noise to the naked eye. This was caused by the oscillation around individual data points when fitting the data with the polynomial, which was similar to the Runge phenomenon. This experiment collected 1,000 pairs of speckle images and reconstructed images and counted the occurrence of noise. Fig.10 As shown in the figure, the noise points mainly appear near the pixel value of 0, which means that PolyNet has an oscillation phenomenon near the pixel value of 0, which is similar to the oscillation phenomenon of the polynomial at the endpoint.
[0103] In view of the oscillation phenomenon that may occur when fitting data with high-order polynomials, this application proposes to integrate it into the diffusion model to solve this phenomenon. The idea of residual and regularization included in the gradual denoising of the diffusion model can effectively suppress the occurrence of polynomial oscillation phenomenon. At the same time, compared with directly using polynomials to fit the data distribution of the target image, fitting the Gaussian distribution is easier. In the field of image-related, the U-net model has been proven to have good performance in various tasks. Therefore, the PDDPM model proposed in this application is implemented using U-net based on polynomial BackBone. As mentioned earlier, this still ensures that the output is a polynomial expression of the input, but the model architecture of U-net is more common. At the same time, the PDDPM model uses a polynomial-based U-net network to predict noise instead of directly fitting the mapping relationship between speckle and target image. This effectively avoids the influence of Runge phenomenon on the reconstruction results. The reconstruction results are shown in Figures 11(a), 11(b) and 11(c). The frequency distribution histograms of PSNR and SSIM of PDDPM on the test set are shown in Figures 11(a), 11(b) and 11(c). Fig.12 The reconstruction results of different models are shown in Table 3. The experimental results show that the PSNR and SSIM evaluation indicators of PDDPM on the test set have reached the optimal level.
[0104] Table 3 Comparison of reconstruction results
[0105]
[0106]
[0107] The present application also provides a dynamic scattering medium imaging system based on a polynomial denoising diffusion probability model, which is implemented based on the above method. The system includes:
[0108] An iterative sampling module is used to iteratively sample the speckle pattern based on the Brownian bridge diffusion model to obtain the final target image;
[0109] The noise prediction module, during the iterative sampling process, constructs a variance table through the diffusion equation, inputs the current image into the trained polynomial neural network, and obtains the noise for the next iteration.
[0110] The present application may also provide a computer device, comprising: at least one processor, a memory, at least one network interface and a user interface. The various components in the device are coupled together through a bus system. It is understood that the bus system is used to achieve connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus.
[0111] The user interface may include a display, a keyboard or a pointing device, such as a mouse, a trackball, a touch pad or a touch screen.
[0112] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0113] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and applications.
[0114] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present disclosure can be included in the application.
[0115] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in an application program, and is used to:
[0116] Execute the steps of the above method.
[0117] The above method can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The above-disclosed methods, steps and logic block diagrams can be implemented or executed. The general processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the above-disclosed method can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0118] It is understood that the embodiments described in the present application can be implemented by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application or a combination thereof.
[0119] For software implementation, the technology of the present application can be implemented by executing the functional modules (such as procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0120] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application is described in detail with reference to the embodiments, a person skilled in the art should understand that any modification or equivalent replacement of the technical solution of the present application does not depart from the spirit and scope of the technical solution of the present application and should be included in the scope of the claims of the present application.
Claims
1. A method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model, comprising: The speckle pattern is iteratively sampled based on the Brownian bridge diffusion model to obtain the final target image; In the iterative sampling process, the variance table is constructed through the diffusion equation, and the current image is input into the trained polynomial neural network to obtain the noise for the next iteration.
2. The method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model according to claim 1, characterized in that: The formula for iterative sampling is: Among them, x t-1 represents the image generated at the t-1th iteration; x T represents the speckle pattern; ∈ θ (x t , t ) represents a polynomial neural network; c xt 、c xT and c ∈t is x in the Brownian bridge diffusion model t 、x T and∈ θ (x t ,t) proportion; represents the variance term in the Brownian bridge diffusion model; z represents the solution of the diffusion equation.
3. The method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model according to claim 2, characterized in that: The diffusion equation is: Where D is the diffusion coefficient, r is the scattering distance, and h is the time when the scattering occurs.
4. The method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model according to claim 2, characterized in that: The construction process of the polynomial neural network includes: Use the quadratic activation function instead of the convolutional neural network activation function.
5. The method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model according to claim 2, characterized in that: The construction process of the polynomial neural network includes: A 3×3 convolutional layer, a quadratic activation function, and a layer normalization layer are used to replace the activation function of the convolutional neural network.
6. The method for dynamic scattering medium imaging based on a polynomial denoising diffusion probability model according to claim 4 or 5, characterized in that: The convolutional neural network is a U-net neural network.
7. A dynamic scattering medium imaging system based on a polynomial denoising diffusion probability model, implemented based on the method of any one of claims 1 to 6, characterized in that: The system comprises: An iterative sampling module, used for iteratively sampling the speckle pattern based on the Brownian bridge diffusion model to obtain a final target image; and The noise prediction module, during the iterative sampling process, constructs a variance table through the diffusion equation, inputs the current image into the trained polynomial neural network, and obtains the noise for the next iteration.