Stimulated Raman microscopic image reconstruction method and system based on diffusion model
By constructing a stimulated Raman microscopic image reconstruction method based on diffusion model, and using the PSF microscopic image dataset training model of the stimulated Raman scattering imaging system, the problem of poor image reconstruction quality in the prior art is solved, and the dual optimization of image detail enhancement and noise suppression is achieved, which improves reconstruction accuracy and adaptability.
Patent Information
- Application Number
- CN202510275705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the stimulated Raman microscopic image reconstruction method is difficult to balance between denoising, deblurring and maintaining image details, resulting in poor image quality.
Using a diffusion model-based image reconstruction method, the diffusion model is trained by constructing the PSF microscopic image dataset of the stimulated Raman scattering imaging system, and using the LR encoder, conditional noise predictor and self-supervised denoiser, high-quality microscopic images are gradually restored.
It significantly improves the quality of image reconstruction, enhances image detail and noise suppression capabilities, improves reconstruction accuracy and adaptability, and ensures the adaptability of the model to specific imaging systems.
Smart Images

Figure CN120259119A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing, and more specifically, relates to a stimulated Raman microscopic image reconstruction method and system based on a diffusion model. Background Art
[0002] As a label-free chemical imaging tool, stimulated Raman scattering (SRS) microscopy has important applications in biomedical research. However, its imaging process is often affected by noise, optical blur, and background interference, resulting in a decline in image quality. Traditional image reconstruction methods often struggle to strike a balance between denoising, deblurring, and preserving image details. This application utilizes the step-by-step denoising characteristics of the diffusion model to design a generative reconstruction framework for SRS microscopic images. This framework simulates the degradation process of SRS images and uses the diffusion model to gradually recover high-quality microscopic images from noise and blur.
[0003] Traditional reconstruction algorithms start from the spatial convolution of the true image of the SRS system and the point spread function (PSF) of the SRS system, and use the deconvolution method to reconstruct microscopic images. In recent years, with the development of deep learning, models based on generative adversarial networks (GANs) have also been widely used in image super-resolution and reconstruction tasks. However, when these methods in the prior art perform image reconstruction, they still suffer from problems such as noise, model accuracy, and model generation effects, resulting in poor image quality.
[0004] Therefore, how to improve the quality of the reconstructed image and the reconstruction efficiency is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the purpose of this application is to provide a stimulated Raman microscopic image reconstruction method and system based on a diffusion model, aiming to solve the problem of poor image reconstruction quality in the prior art.
[0006] To achieve the above objective, in a first aspect, this application provides a stimulated Raman microscopic image reconstruction method based on a diffusion model, including: Obtain the LR image to be reconstructed; Input the LR image into the trained diffusion model to obtain a high-frequency image; Perform denoising processing on the LR image to obtain a denoised image; Add the denoised image and the high-frequency image to obtain a reconstructed image; Wherein, the diffusion model is obtained by constructing a data set based on the PSF microscopic image of the stimulated Raman scattering imaging system and performing noise prediction training on the data set.
[0007] Optionally, the method for constructing the data set includes: Select high-resolution images from the target dataset as the reference images; Perform two-dimensional spatial convolution processing on the reference image and the PSF microscopic image to obtain a degraded blurred image; Add noise to the blurred image to obtain a low-resolution image; Form an HR-LR training sample pair with the reference image and the corresponding low-resolution image, and construct the dataset according to the training sample pair.
[0008] Optionally, the method for obtaining the PSF microscopic image includes: Use fluorescent microspheres of the target size as the imaging sample, and image multiple fluorescent microspheres respectively through a stimulated Raman scattering imaging system to obtain multiple PSF original images; Perform cumulative averaging processing on multiple PSF original images to generate a PSF microscopic image after eliminating random noise.
[0009] Optionally, the diffusion model includes an LR encoder, a conditional noise predictor, and a self-supervised denoiser; The LR encoder is used to encode according to the LR image and the current diffusion time step to obtain image features; The conditional noise predictor is used to predict the noise in the diffusion process based on the image features and the diffusion time step; The self-supervised denoiser is used to add the predicted noise to the LR image to generate a reconstructed high signal-to-noise ratio reconstructed image.
[0010] Optionally, the LR encoder includes a spatial attention module and a channel attention module; The channel attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a first pooling result; Input the first pooling result into a shared multi-layer perceptron respectively, generate channel weights through a Sigmoid activation function, and multiply the channel weights by the LR image to obtain a channel-weighted feature map; The spatial attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a second pooling result; Concatenate the second pooling result along the channel dimension, compress it through a convolutional layer, and generate spatial weights through a Sigmoid activation function, and multiply the spatial weights by the LR image to obtain a spatial-weighted feature map.
[0011] Optionally, the conditional noise predictor specifically includes: A convolutional block for converting image features into hidden states through two-dimensional convolution and an activation function; An information fusion layer for fusing the image features with the hidden state output by the convolutional block to generate a fused hidden state; A time step encoding layer for converting the diffusion time step into a time step with position information through a position encoder and embedding the time step with position information into the image features of the fused hidden state; A downsampling residual layer for performing feature extraction and dimensionality reduction on the fused hidden state and the encoded time step through a residual block; An intermediate feature extraction layer for further extracting and fusing features; An upsampling restoration layer for gradually restoring the resolution of the feature map and reducing the number of channels through upsampling to generate a high-resolution feature map; A noise prediction layer for predicting the noise added in the current diffusion step based on the feature map.
[0012] Optionally, the training method of the diffusion model includes: Selecting HR-LR training sample pairs and taking the difference between the high-resolution image and the corresponding low-resolution image as the target residual; Performing feature encoding on the low-resolution image through a low-resolution image encoder to obtain the encoded low-resolution features; Randomly selecting a diffusion time step, randomly generating noise from a Gaussian distribution, and adding noise to the target residual according to the diffusion time step to generate noisy intermediate data; Inputting the noisy intermediate data, the diffusion time step, and the encoded low-resolution features into a conditional noise predictor to obtain the predicted noise; Calculating the mean square error between the predicted noise and the actually added noise, and optimizing the network parameters through error backpropagation, repeating the iteration until the model converges.
[0013] Optionally, inputting the LR image into the trained diffusion model to obtain a high-frequency image includes: Inputting the low-resolution image to be reconstructed into the LR encoder to extract encoded features; Randomly generating an initial noise image from a Gaussian distribution and performing iterative denoising in descending order from the maximum to the minimum according to the diffusion time step until the time step reaches the minimum value, and outputting the final high-frequency residual image.
[0014] This application provides a stimulated Raman microscopic image reconstruction system based on a diffusion model, including: An acquisition module for acquiring the LR image to be reconstructed; An encoding module for inputting the LR image into the trained diffusion model to obtain a high-frequency image; A denoising module for denoising the LR image to obtain a denoised image; A reconstruction module for adding the denoised image and the high-frequency image to obtain a reconstructed image; Wherein, the diffusion model is obtained by constructing a data set based on the PSF microscopic image of the stimulated Raman scattering imaging system and performing noise prediction training on the data set.
[0015] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0017] In a fifth aspect, the present application provides a computer program product, and when the computer program product runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0018] It can be understood that the beneficial effects of the above second to fifth aspects can refer to the relevant descriptions in the above first aspect and will not be elaborated here.
[0019] Generally speaking, compared with the prior art through the above technical solutions conceived by the present application, the following beneficial effects are obtained: (1) The present application inputs the low-resolution LR image to be reconstructed into the trained diffusion model to generate a high-frequency image to supplement detailed information, and at the same time denoises the LR image to eliminate noise interference. Finally, the denoised image and the high-frequency image are added together, realizing the dual optimization of image detail enhancement and noise suppression, and significantly improving the quality of the reconstructed image. In addition, the diffusion model is trained based on the PSF microscopic image data set of the stimulated Raman scattering imaging system, ensuring the adaptability of the model to a specific imaging system and further improving the reconstruction accuracy.
[0020] (2) The diffusion model of the present application is trained based on the PSF microscopic image data set of the stimulated Raman scattering imaging system. By simulating the degradation process of the real imaging system, such as blurring and noise addition, the adaptability of the model to a specific imaging system is ensured. Through targeted training, the model can more accurately restore image details and improve the reconstruction accuracy.
[0021] (3) The LR encoder of this application introduces a spatial attention module and a channel attention module, which weight the image features from the spatial and channel dimensions respectively, enhancing the model's ability to extract key features and enabling more effective utilization of useful information in low-resolution images, further improving the reconstruction effect.
[0022] (4) The conditional noise predictor of this application accurately predicts the noise added during the diffusion process through operations such as multi-level feature fusion, time step encoding, and residual dimensionality reduction. Combined with the self-supervised denoiser, the model can gradually remove the noise and restore the high signal-to-noise ratio image, ensuring the stability and reliability of the reconstruction process. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is one of the schematic flowcharts of the stimulated Raman microscopic image reconstruction method based on the diffusion model provided by the embodiments of this application; Figure 2 is a schematic diagram of the SRS system PSF results collected by the fluorescence microsphere method in the embodiments of this application; Figure 3 is a schematic diagram of constructing the LR-SR dataset according to the SRS system degradation principle in the embodiments of this application; Figure 4 is a schematic diagram of the LR encoder in the embodiments of this application; Figure 5 is a network schematic diagram of the conditional noise predictor in the model of the embodiments of this application; Figure 6 is a schematic flowchart of the model training in the embodiments of this application; Figure 7 is a schematic flowchart of the model inference in the embodiments of this application; Figure 8 is another schematic flowchart of the stimulated Raman microscopic image reconstruction method based on the diffusion model provided by the embodiments of this application; Figure 9 is a schematic structural diagram of the stimulated Raman microscopic image reconstruction device based on the diffusion model provided by the embodiments of this application; Figure 10 is a schematic structural diagram of the electronic device provided by the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0025] As used herein, the term "and / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. As used herein, the symbol " / " indicates that the associated objects are in an "or" relationship. For example, A / B means A or B.
[0026] The terms "first", "second", etc. in the description and claims of this application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0027] In the embodiments of this application, words such as "exemplary" or "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0028] In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0029] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application.
[0030] Referring to Figure 1 , this application provides a stimulated Raman microscopic image reconstruction method based on a diffusion model, including: S101. Obtain the LR image to be reconstructed; S102. Input the LR image into the trained diffusion model to obtain a high-frequency image; S103. Denoise the LR image to obtain a denoised image; S104. Add the denoised image and the high-frequency image to obtain a reconstructed image; Wherein, the diffusion model is constructed by using the PSF microscopic images of the stimulated Raman scattering imaging system to construct a data set, and the data set is trained for noise prediction.
[0031] Specifically, the embodiments of this application are for the process of image reconstruction.
[0032] First, obtain the LR image to be reconstructed through step S101, At this stage, it is first necessary to obtain the low-resolution (LR) images to be reconstructed. These images are usually acquired through a certain imaging device or method. However, due to resolution limitations, the images may lack details or clarity. By selecting appropriate LR images as input, the effectiveness of subsequent processing can be ensured.
[0033] Secondly, input the LR image into the trained diffusion model through step S102 to obtain a high-frequency image.
[0034] Input the LR image into a trained diffusion model. This model is trained using a dataset constructed from point spread function (PSF) microscopic images of a stimulated Raman scattering imaging system and can effectively process spatial information and noise. In this step, the diffusion model uses the features it has learned to convert the low-frequency information in the LR image into high-frequency information, thereby generating a high-frequency image containing more details. This process provides important high-frequency features for reconstructing the image and helps improve the clarity of the final image.
[0035] Furthermore, denoise the LR image through step S103 to obtain a denoised image.
[0036] At this stage, denoise the input LR image, aiming to reduce the random noise in the image and extract a cleaner image signal. This process usually employs various denoising algorithms, such as median filtering, Wiener filtering, or deep learning denoising techniques, with the aim of improving the image quality and ensuring that subsequent steps are not interfered by noise. After denoising, the generated image is called a denoised image, which can significantly reduce the error in the image reconstruction process.
[0037] Finally, add the denoised image and the high-frequency image through step S104 to obtain a reconstructed image The last step is to add the image after denoising processing to the high-frequency image obtained from the diffusion model. This addition process mainly combines the clear basis of the denoised image with the detailed part of the high-frequency image to form a more complete and higher-quality reconstructed image. This process ensures the balance between the details and the overall quality of the image, and the finally output reconstructed image contains high visual clarity and details. The embodiments of this application aim to improve the image quality through the trained diffusion model and denoising techniques and achieve a higher level of image reconstruction effect.
[0038] Optionally, the method for constructing the dataset includes: Select high-resolution images from the target dataset as reference images; Perform two-dimensional spatial convolution processing on the reference image and the PSF microscopic image to obtain a degraded blurred image; Add noise to the blurred image to obtain a low-resolution image; Form an HR-LR training sample pair from the reference image and the corresponding low-resolution image, and construct the dataset according to the training sample pair.
[0039] Further, the method for obtaining the PSF microscopic image includes: Use fluorescent microspheres of a target size as an imaging sample, and image multiple fluorescent microspheres respectively through a stimulated Raman scattering imaging system to obtain multiple original PSF images; Perform cumulative averaging processing on multiple original PSF images to generate a PSF microscopic image after eliminating random noise.
[0040] Specifically, it should be noted first that the stimulated Raman scattering imaging system (Stimulated Raman Scattering, SRS) in the embodiments of the present application is a non-linear optical imaging technology, mainly used for label-free chemical imaging or biological imaging. The microscopic images of the SRS system may be affected by various degradation factors, and these factors can be described by a degradation model. The following is the degradation model of the microscopic image of the SRS system and its main components:
[0041] Among them, is the observed degraded image, is the ideal, non-degraded image. is the degradation function, which describes the process of image degradation. is the additive noise.
[0042] The optical resolution of the SRS system is limited by the point spread function (PSF) in the same way as the traditional optical microscopic system. This is determined by the diffraction limit of the optical system and is an inherent property of the optical system. During the system imaging process, the convolution of the real sample with the PSF of the SRS system will cause image blurring and reduce the spatial resolution. Therefore, many algorithms start from the perspective of deconvolution, and the algorithms for image restoration are used to eliminate or reduce image degradation. The degradation function can be expressed as the spatial convolution of the ideal image and the PSF:
[0043] Common noises in SRS images include: detector noise, thermal noise, shot noise, etc. Photon noise is Poisson noise caused by the randomness of photon counting. Detector noise is additive Gaussian noise introduced by the detector itself when receiving photons. Thermal noise is noise caused by system temperature fluctuations. However, shot noise is the main factor affecting SRS. Shot noise is caused by the randomness of photon counting and follows a Poisson distribution. It is the inherent noise of the laser light source and is proportional to the square root of the laser power. In an SRS system, the high power and modulation frequency (usually in megahertz) of the laser will amplify the impact of shot noise. These noises are usually additive and can be approximated by Gaussian noise. The noise signal and the background signal can be separately modeled as an additive component:
[0044] During a long imaging process, the sample may drift or vibrate, which also causes image blurring and ghosting. Minimize the imaging time as much as possible when acquiring images to reduce sample drift.
[0045] Refer to Figure 2 , and the PSF of the SRS system is collected by using the fluorescent microsphere method. Usually, the size of the fluorescent microspheres should be as small as possible to reduce the impact of their size on PSF measurement. Under certain conditions, multiple small spheres are imaged separately, and multiple PSF images are obtained and subjected to cumulative averaging processing to obtain the PSF of the SRS system.
[0046] Use a target dataset, such as high-resolution images in the DIV2K dataset, for degradation processing, that is, perform two-dimensional convolution with Figure 2 , and add noise with an amplitude of 5% of the maximum intensity of the signal to obtain the degraded low-resolution image LR, which forms an HR-LR image pair with the original image and is used as the training set. The degradation process is as shown in Figure 3 .
[0047] It should be added that the DIV2K dataset in the embodiments of this application is a high-quality dataset widely used in image super-resolution (SR) research. The "DIV" in the dataset name represents "DiverseImage Vision", and "M2K" indicates that the dataset contains images with a resolution of 2K. This dataset consists of multiple high-resolution (HR) images and corresponding low-resolution (LR) images and can be used to train, validate, and test different super-resolution models.
[0048] Optionally, the diffusion model includes an LR encoder, a conditional noise predictor, and a self-supervised denoiser; The LR encoder is used to encode according to the LR image and the current diffusion time step to obtain image features; The conditional noise predictor is used to predict the noise in the diffusion process based on the image features and the diffusion time step; The self-supervised denoiser is used to add the predicted noise to the LR image to generate a reconstructed high signal-to-noise ratio reconstructed image.
[0049] Optionally, the LR encoder includes a spatial attention module and a channel attention module; The channel attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a first pooling result; Input the first pooling result into a shared multi-layer perceptron respectively, generate channel weights through the Sigmoid activation function, and multiply the channel weights by the LR image to obtain a channel-weighted feature map; The spatial attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a second pooling result; Concatenate the second pooling result along the channel dimension, compress it through a convolutional layer, generate spatial weights through the Sigmoid activation function, and multiply the spatial weights by the LR image to obtain a spatially weighted feature map.
[0050] Specifically, the main role of the LR encoder in the embodiments of the present application is to encode LR information, and the encoded LR information is added to each reverse diffusion process to guide the generation in the corresponding HR space.
[0051] The LR encoder is a nested residual dense block (RRDB) module, which is an encoding module for image super-resolution, adopting a residual-non-residual structure and multiple dense skip connections, and deleting the batch normalization layer.
[0052] It should be noted that referring to Figure 4 , Figure 4 is a schematic diagram of the LR encoder in the embodiments of the present application. In order to enhance the acquisition of hidden information in the LR image in the embodiments of the present application, a spatial attention module and a channel attention module are added on the basis of the RRDB, and the performance of extracting image hidden information is improved with a small increase in computational complexity. 5 RRDBs with added spatial attention module and channel attention module are connected in series as the LR encoder. Channel attention is calculated by calculating, through global max pooling on the channel dimension of the input feature map F , two feature maps of H×W×1 are obtained, and then fed into a shared multi-layer perceptron (MLP). The weights are obtained by adding the two, and then the Sigmoid activation function is used to map it to the range of [0,1] and multiply it with the original feature map F to obtain a feature map with channel weights. The spatial attention is calculated through . After performing global max pooling and global average pooling in the channel dimension, the feature maps are concatenated according to the channel dimension, and then the concatenated feature maps are subjected to a convolution operation. After passing through the Sigmoid activation function, it is multiplied by the original feature map F to obtain a feature map with spatial weights.
[0053] Optionally, the conditional noise predictor specifically includes: A convolutional block for converting image features into hidden states through two-dimensional convolution and an activation function; An information fusion layer for fusing image features with the hidden states output by the convolutional block to generate fused hidden states; A time step encoding layer for converting the diffusion time step into a time step with position information through a position encoder and embedding the time step with position information into the image features of the fused hidden states; A downsampling residual layer for extracting features and reducing the dimension of the fused hidden states and the encoded time step through a residual block; An intermediate feature extraction layer for further extracting and fusing features; An upsampling recovery layer for gradually recovering the resolution of the feature map and reducing the number of channels through upsampling to generate a high-resolution feature map; A noise prediction layer for predicting the noise added in the current diffusion step based on the feature map.
[0054] Specifically, referring to Figure 5 , the main role of the noise predictor is to predict the noise added at each diffusion time step under the condition of LR image information. The network needs to continuously change the parameters to gradually approximate the difference between the actually added noise and the noise predicted by the network. The network uses UNet as the main body, and the diffusion time step , and takes the output features of the LR encoder as the input.
[0055] The specific process is as follows: First, the image features x t after diffusion for T - t are converted into hidden states through a 2D convolutional block composed of a 2D convolutional layer and a Mish activation. Then, the LR information is fused with the hidden states output by the next 2D convolutional block. The time step t is converted into a time step t with position information using the position encoder of the Transformer e, and embed the transformed x t into it. After these operations, the final output hidden state of the 2D convolutional block and t e are input into the downsampled residual block. Then, through the intermediate layer, which contains multiple convolutional layers and activation functions, for further feature extraction and fusion. Finally, the resolution of the feature map is gradually restored through upsampling, the number of channels is reduced, and a higher-resolution feature map is generated. Finally, the noise added in the t-th diffusion step is predicted and used to restore x t-1 .
[0056] Optionally, the training method of the diffusion model includes: Select HR-LR training sample pairs, and use the difference between the high-resolution image and the corresponding low-resolution image as the target residual; Perform feature encoding on the low-resolution image through the low-resolution image encoder to obtain the encoded low-resolution features; Randomly select the diffusion time step, randomly generate noise from the Gaussian distribution, and add noise to the target residual according to the diffusion time step to generate the noisy intermediate data; Input the noisy intermediate data, the diffusion time step, and the encoded low-resolution features into the conditional noise predictor to obtain the predicted noise; Calculate the mean square error between the predicted noise and the actually added noise, and optimize the network parameters through error backpropagation, repeating the iteration until the model converges.
[0057] Specifically, referring to Figure 6 , in the training stage, the constructed LR-HR image pairs are used to train the network with a total of T diffusion steps. At the beginning of training, the conditional noise predictor is randomly initialized first, sample LR-HR image pairs from the training set, and calculate the residual image x r . The LR image is encoded by the LR encoder into x T , which is input into the noise predictor together with the diffusion step t. Then, we sample from the standard Gaussian distribution, where t belongs to the set of integers . In the forward process of the network, for each data x0, randomly sample the time step t, and generate the noisy data x t according to the forward process, and use the neural network to predict the noise. The loss function is defined as the mean square error between the predicted noise and the true noise. Then update the parameters of the network, repeat the above steps until the model converges.
[0058] Optionally, input the LR image into the trained diffusion model to obtain the high-frequency image, including: Input the low-resolution image to be reconstructed into the LR encoder to extract the encoded features; Randomly generate an initial noise image from a Gaussian distribution, and perform iterative denoising by gradually decreasing from the maximum to the minimum according to the diffusion time step until the time step reaches the minimum, and then output the final high-frequency residual image.
[0059] Refer to Figure 7 , during the inference process of the diffusion model, first take the LR image as the input, sample from the standard Gaussian distribution, and encode the LR image only once through the LR encoder before the start of the iterative process and apply it in each iteration. In each iteration, the trained network is used to predict a noise, and then the noise is removed from the image. By continuously iterating, a residual image with different noise levels is output, and the noise level gradually decreases as the time step t decreases. When t > 1, the model continuously samples from the standard Gaussian distribution and calculates the next time step x t-1 . When the time step t = 1, the obtained x0 is the final high-frequency prediction image, which contains rich detailed information predicted by the network based on the LR image. The network finally outputs x0, which is the high-frequency detail part of the image predicted based on the LR, while the low-frequency part of the reconstructed image is provided by the LR image. To improve the overall signal-to-noise ratio of the reconstructed image, the LR image is denoised to obtain the image LR denosied , and add it to x0 as the final reconstruction result x sr . The output image of the model inference is a high-frequency image predicted based on the LR degraded image as the latent variable input. It needs to be added to the original low-frequency image to form a complete reconstructed image. Therefore, the noise of the LR degraded image needs to be removed to ensure that the synthesized SR image contains as little noise as possible, and improve the signal-to-noise ratio and image details of the output.
[0060] Refer to Figure 8 , the complete steps of the embodiment of the present application are as follows: Collect the PSF of the SRS system and construct a low-resolution LR - super-resolution SR dataset; Use the LR - SR dataset to train the model; Use the model to infer and obtain a high-frequency image; Add the denoised LR to the high-frequency image to obtain the reconstructed image.
[0061] Refer to Figure 9 , the present application provides a stimulated Raman microscopic image reconstruction system based on a diffusion model, including: An acquisition module 910, configured to acquire the LR image to be reconstructed; An encoding module 920, configured to input the LR image into the trained diffusion model to obtain a high-frequency image; A denoising module 930, configured to perform denoising processing on the LR image to obtain a denoised image; A reconstruction module, configured to add the denoised image and the high-frequency image to obtain a reconstructed image; Wherein, the diffusion model is obtained by constructing a dataset based on the PSF microscopic images of the stimulated Raman scattering imaging system and performing noise prediction training on the dataset.
[0062] Optionally, the method for constructing the dataset includes: Selecting a high-resolution image from the target dataset as a reference image; Performing two-dimensional spatial convolution processing on the reference image and the PSF microscopic image to obtain a degraded blurred image; Adding noise to the blurred image to obtain a low-resolution image; Combining the reference image and the corresponding low-resolution image into an HR-LR training sample pair, and constructing the dataset according to the training sample pair.
[0063] Optionally, the method for obtaining the PSF microscopic image includes: Taking fluorescent microspheres of a target size as imaging samples, and respectively imaging a plurality of fluorescent microspheres through a stimulated Raman scattering imaging system to obtain a plurality of PSF original images; Performing cumulative average processing on the plurality of PSF original images to generate a PSF microscopic image after eliminating random noise.
[0064] Optionally, the diffusion model includes an LR encoder, a conditional noise predictor, and a self-supervised denoiser; The LR encoder is configured to encode according to the LR image and the current diffusion time step to obtain image features; The conditional noise predictor is configured to predict the noise in the diffusion process based on the image features and the diffusion time step; The self-supervised denoiser is configured to add the predicted noise to the LR image to generate a reconstructed high signal-to-noise ratio reconstructed image.
[0065] Optionally, the LR encoder includes a spatial attention module and a channel attention module; The channel attention module is specifically configured to: Performing global max pooling and global average pooling on the LR image to obtain a first pooling result; Inputting the first pooling result into a shared multi-layer perceptron respectively, generating channel weights through a Sigmoid activation function, and multiplying the channel weights by the LR image to obtain a channel-weighted feature map; The spatial attention module is specifically configured to: Performing global max pooling and global average pooling on the LR image to obtain a second pooling result; Concatenate the second pooling result along the channel dimension, compress it through a convolutional layer, and generate a spatial weight through a Sigmoid activation function. Multiply the spatial weight by the LR image to obtain a spatially weighted feature map.
[0066] Optionally, the conditional noise predictor specifically includes: A convolutional block for converting image features into a hidden state through two-dimensional convolution and an activation function; An information fusion layer for fusing the image features with the hidden state output by the convolutional block to generate a fused hidden state; A time step encoding layer for converting the diffusion time step into a time step with position information through a position encoder and embedding the time step with position information into the image features of the fused hidden state; A downsampling residual layer for performing feature extraction and dimensionality reduction on the fused hidden state and the encoded time step through a residual block; An intermediate feature extraction layer for further extracting and fusing features; An upsampling recovery layer for gradually restoring the resolution of the feature map and reducing the number of channels through upsampling to generate a high-resolution feature map; A noise prediction layer for predicting the noise added in the current diffusion step based on the feature map.
[0067] Optionally, the training method of the diffusion model includes: Select HR-LR training sample pairs, and use the difference between the high-resolution image and the corresponding low-resolution image as the target residual; Perform feature encoding on the low-resolution image through a low-resolution image encoder to obtain encoded low-resolution features; Randomly select a diffusion time step, randomly generate noise from a Gaussian distribution, and add noise to the target residual according to the diffusion time step to generate noisy intermediate data; Input the noisy intermediate data, the diffusion time step, and the encoded low-resolution features into the conditional noise predictor to obtain a predicted noise; Calculate the mean square error between the predicted noise and the actually added noise, and optimize the network parameters through error backpropagation. Repeat the iteration until the model converges.
[0068] Optionally, inputting the LR image into the trained diffusion model to obtain a high-frequency image includes: Input the low-resolution image to be reconstructed into the LR encoder to extract encoded features; Randomly generate an initial noise image from a Gaussian distribution, and perform iterative denoising from the maximum value to the minimum value according to the diffusion time step until the time step reaches the minimum value, and then output the final high-frequency residual image.
[0069] It can be understood that for the detailed function implementation of each of the above units / modules, reference can be made to the introduction in the foregoing method embodiments, and details are not described herein.
[0070] It should be understood that the above device is used to execute the method in the above embodiments. For the corresponding program modules in the device, their implementation principles and technical effects are similar to those described in the above method. The working process of the device can refer to the corresponding process in the above method, and details are not described herein.
[0071] Referring to Figure 10 , based on the method in the above embodiments, an embodiment of the present application provides an electronic device, which may include: a processor 110, a communications interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communications interface 120, and the memory 130 communicate with each other through the communication bus 140. The processor 110 can call the logical instructions in the memory 130 to execute the method in the above embodiments.
[0072] In addition, when the logical instructions in the above memory 830 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.
[0073] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0074] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0075] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0076] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in the ASIC.
[0077] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0078] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for convenience of description and are not used to limit the scope of the embodiments of the present application.
[0079] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A stimulated Raman microscopic image reconstruction method based on a diffusion model, characterized in that, Including: Obtain the LR image to be reconstructed; Input the LR image into the trained diffusion model to obtain a high-frequency image; Denoise the LR image to obtain a denoised image; Add the denoised image and the high-frequency image to obtain a reconstructed image; Among them, the diffusion model is constructed by using the PSF microscopic images of the stimulated Raman scattering imaging system to construct a dataset, and the dataset is trained for noise prediction.
2. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 1, wherein The method for constructing the dataset includes: Select a high-resolution image from the target dataset as the reference image; Perform two-dimensional spatial convolution processing on the reference image and the PSF microscopic image to obtain a degraded blurred image; Add noise to the blurred image to obtain a low-resolution image; Form an HR-LR training sample pair with the reference image and the corresponding low-resolution image, and construct the dataset according to the training sample pair.
3. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 2, wherein The method for obtaining the PSF microscopic image includes: Use fluorescent microspheres of the target size as imaging samples, and image multiple fluorescent microspheres respectively through the stimulated Raman scattering imaging system to obtain multiple PSF original images; Perform cumulative averaging processing on multiple PSF original images to generate a PSF microscopic image after eliminating random noise.
4. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 1, wherein The diffusion model includes an LR encoder, a conditional noise predictor, and a self-supervised denoiser; The LR encoder is used to encode according to the LR image and the current diffusion time step to obtain image features; The conditional noise predictor is used to predict the noise in the diffusion process based on the image features and the diffusion time step; The self-supervised denoiser is used to add the predicted noise to the LR image to generate a reconstructed high signal-to-noise ratio reconstructed image.
5. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 4, wherein, The LR encoder includes a spatial attention module and a channel attention module; The channel attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a first pooling result; Input the first pooling result into a shared multi-layer perceptron respectively, generate channel weights through the Sigmoid activation function, and multiply the channel weights by the LR image to obtain a channel-weighted feature map; The spatial attention module is specifically used for: Perform global max pooling and global average pooling on the LR image to obtain a second pooling result; Concatenate the second pooling result along the channel dimension, generate spatial weights through convolution layer compression and then through the Sigmoid activation function, and multiply the spatial weights by the LR image to obtain a spatial-weighted feature map.
6. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 4, wherein, The conditional noise predictor specifically includes: A convolutional block for converting image features into hidden states through two-dimensional convolution and activation functions; An information fusion layer for fusing the image features with the hidden states output by the convolutional block to generate fused hidden states; A time step encoding layer for converting the diffusion time step into a time step with position information through a position encoder, and embedding the time step with position information into the image features of the fused hidden states; A downsampling residual layer for performing feature extraction and dimensionality reduction on the fused hidden states and the encoded time step through a residual block; An intermediate feature extraction layer for further extracting and fusing features; An upsampling restoration layer for gradually restoring the resolution of the feature map and reducing the number of channels through upsampling to generate a high-resolution feature map; A noise prediction layer for predicting the noise added in the current diffusion step based on the feature map.
7. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 2, characterized in that, The training method of the diffusion model includes: Selecting HR-LR training sample pairs, and taking the difference between the high-resolution image and the corresponding low-resolution image as the target residual; Performing feature encoding on the low-resolution image through a low-resolution image encoder to obtain the encoded low-resolution features; Randomly selecting a diffusion time step, randomly generating noise from a Gaussian distribution, and adding noise to the target residual according to the diffusion time step to generate noisy intermediate data; Inputting the noisy intermediate data, the diffusion time step, and the encoded low-resolution features into a conditional noise predictor to obtain predicted noise; Calculating the mean square error between the predicted noise and the actually added noise, and optimizing the network parameters through error backpropagation, repeating the iteration until the model converges.
8. The stimulated Raman microscopic image reconstruction method based on a diffusion model according to claim 1, characterized in that Inputting the LR image into the trained diffusion model to obtain a high-frequency image, including: Inputting the low-resolution image to be reconstructed into the LR encoder to extract encoded features; Randomly generating an initial noise image from a Gaussian distribution, and performing iterative denoising from the maximum value to the minimum value step by step according to the diffusion time step until the time step reaches the minimum value, and outputting the final high-frequency residual image.
9. A stimulated Raman microscopic image reconstruction system based on a diffusion model, characterized in that, Including: An acquisition module for acquiring the LR image to be reconstructed; An encoding module for inputting the LR image into the trained diffusion model to obtain a high-frequency image; A denoising module for denoising the LR image to obtain a denoised image; A reconstruction module for adding the denoised image and the high-frequency image to obtain a reconstructed image; Wherein, the diffusion model is constructed by using the PSF microscopic images of the stimulated Raman scattering imaging system to construct a data set, and the data set is trained for noise prediction.
10. An electronic device, characterized in that, Including: At least one memory for storing a computer program; At least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-8.
Citation Information
Cited By
Super-resolution fluorescence image reconstruction method based on diffusion model and related equipment
CN120598780A
A super-resolution fluorescence image reconstruction method based on a diffusion model and related equipment
CN120598780B