Microscopic system out-of-focus identification and restoration method
By constructing a two-stage cycle image defocus recognition and repair neural network model, combining defocus parameter prediction and image repair network, the image blur recognition and repair problems of microscopic imaging system during large-scale defocusing is solved, and higher adaptability and interpretability are achieved.
Patent Information
- Application Number
- CN202510281003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
When existing microscopic imaging technologies face large-scale defocusing, it is difficult to effectively identify and repair image blur, resulting in narrow adaptation range, misjudgment and confusion, and the image repair process lacks interpretability.
Using a fusion method based on the generative adversarial network and the focal surface recognition neural network of the microscopic system, a two-stage cycle image defocus recognition and repair neural network model is constructed. Through the combination of defocus parameter prediction and image repair network, the point diffusion function is used to perform convolution and deconvolution operations to achieve clear image processing.
It improves the adaptability of microscopic imaging systems to large-scale defocus blur, can stably analyze the causes of defocusing, explicitly describe target defocus blur, reduce artifacts, and improve the accuracy and interpretability of image repair.
Smart Images

Figure CN120356208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent image processing, and more specifically, to a method for defocus recognition and repair of a microscopic system. Background Art
[0002] The defocus processing algorithms for microscopic images mainly adopt the following two deep learning-based defocus processing solutions:
[0003] 1) Depth of field restoration based on defocus parameter judgment
[0004] For the defocus blur existing in the image, such algorithms use a deep learning model to identify the defocus information, and complete the defocus removal calculation of the image by using a post-processing method according to the defocus information. This type of deep learning model focuses on the inference of the physical parameters of the microscopic system itself or the degree of blur, and inversely calculates the real in-focus image according to the predicted blur coefficient. The key lies in the understanding of the prior defocus distortion of the microscopic system. In this direction, a traditional inference prediction model is generally used as a feature extractor, and a description head of physical parameters is used to understand the system.
[0005] The neural network structures applied to this type of model mostly originate from traditional deep learning classification networks such as VGG and ResNet. Most of the early related neural network systems were used for a single autofocus function, classifying the defocus of the image according to refined discrete partitions, and guiding the microscopic system to move the corresponding distance after the judgment. To obtain higher defocus parameter judgment accuracy, these models also tried to incorporate means such as frequency domain signal input, off-axis illumination light source, and fixed vertical distance image difference to complete further focusing classification. Another system focuses on the prediction of the overall defocus and distortion parameters of the microscopic system, and is used to generate a point spread function (PSF) convolution kernel for deconvolution processing on software. In this case, the defocus parameter judgment is no longer limited to the discrete focal length level, but directly outputs continuous parameter values according to the fully connected layer of the neural network classifier or outputs continuous parameter values through a multi-layer perceptron (MLP) to simulate numerical outputs proportional to the defocus distance, astigmatism, coma, spherical aberration, etc., such as Zernike polynomial coefficients. The frequency domain image can be further phase-modulated, etc. to better emphasize the defocus distance information; and at the output level, since the images in a large imaging field of view often do not uniformly obey the defocus of the paraxial PSF, there are also studies on developing a sliding window method to separately judge the defocus parameters of different field of view regions and then perform average calculation to obtain the final repair result.
[0006] Most of the above methods require the cooperation of hardware system control to refocus, and the depth of field range of a single image is still narrow; or traditional deconvolution methods are needed to perform image restoration and perform deconvolution according to the predicted defocus parameters. However, the iteration of the deconvolution algorithm is not the inverse operation of the convolution itself, and it is impossible to completely restore the signals lost or aliased during the convolution process. Moreover, the comparison between the restored image and the real in-focus image is only used as an evaluation indicator for the quality of the defocus coefficient prediction, and does not participate in the constraints of image restoration. Therefore, the depth of field restoration of this type of method often cannot cope with more complex textures and degradation information, and still retains a lot of blur and artifacts.
[0007] 2) Defocus Restoration Based on Generative Adversarial Network (GAN)
[0008] This type of defocus restoration network is different from the above methods. It does not estimate the defocus parameters or only implicitly involves the estimation of defocus restoration in the network. It directly restores the blurred image to a real in-focus clear image through feature extraction of the neural network in an end-to-end manner. Most of these traditional methods can only focus on a single specific image degradation such as blurring, scaling, etc., and cannot take into account the various reasons that cause the blurring of small objects in the background in complex environments. The feature description of the image is also limited by the generator structure used.
[0009] Most of the relatively successful image super-resolution processing methods are derived from the SRGAN (Super-Resolution Generative Adversarial Network) neural network framework. The generator of this framework adopts a ResNet-based structure, ensuring the accuracy of feature map extraction and description. This method first used the concept of perceptual loss, and optimized the image content of the generator by the mean error between the feature map content extracted by the pre-trained VGG network and the generated image, ensuring the credibility of the generated high-resolution image. The ESRGAN (Enhanced Super Resolution Generative Adversarial Networks) network improves the generator on the basis of SRGAN, and introduces the RRDB (Residual-in-Residual Dense Block) structure into the generator, introducing cross-layer connections during feature extraction, which is beneficial to the complete extraction of context-related feature information and improves training stability. However, this network only considers the case of simple image downscaling. Real-ESRGAN uses two ESRGAN generators in the generator and trains them by the exponential moving average method. In the discriminator, a simple Unet structure is used to replace the original VGG discriminator, and multiple degradation modes are fused during the training dataset. Another example is the DLAM model, which focuses on the super-resolution of medical microscopic images and adds a self-attention mechanism to the RRDB module to make the detailed features more prominent. The DeblurGAN model uses the FPN architecture to achieve the function of connecting the context, making it easier to implement object detection. However, it mainly focuses on the removal of motion blur rather than the improvement of image scaling resolution. In addition, in recent years, large language models based on Diffusion and Transformer, such as UniFMIR, which are trained based on big data, have also started to handle image restoration tasks for various microscopic systems, such as super-resolution and 3D denoising.
[0010] The above systems have been widely used in biomedical microscopy and are used in conjunction with multi-modal fluorescence imaging microscopes to complete multi-modal imaging work with enhanced spatial resolution and running time advantages. In addition, in recent years, there has been a growing body of research using prior physical knowledge, such as ray tracing of optical systems and implicit inference of specific imaging plane distortion, to more stably restore images within a specific range. Further research on GANs is also optimizing image restoration. For example, the use of task-GAN, which uses additional tasks such as image semantic segmentation, object detection, and instance segmentation to freeze and iteratively train with the image restoration network, thus achieving the result of further supervising image restoration using the semantic judgment of the added tasks.
[0011] However, the above image restoration techniques also have defects. Even if prior knowledge is introduced to correct image blurring or distortion, their work is implicit and there is no targeted description of the cause of blurring in terms of interpretability. Using the same set of parameters to describe and restore the defocus of the system, when the defocus range is too large, the styles of blurred images are not unified, which is very likely to cause misjudgment and confusion in system restoration. Therefore, the above systems often face problems such as a narrow adaptation range, inferring more detailed high-frequency information based on a small amount of information, and causing a large number of redundant artifacts.
[0012] During the defocus removal process of the microfluidic system, the system needs to face a defocus range of ±40 - 50 μm, far exceeding the style conversion range that traditional end-to-end mapping can handle. Therefore, a new structure needs to be designed to understand the image blurring caused by different defocus ranges and map a large range of blurring situations to a clear in-focus image. Summary of the Invention
[0013] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention proposes a method for defocus recognition and restoration of a microscopic system based on the fusion of a generative adversarial model and a microscopic system focal plane recognition neural network. It is a new software processing solution for panoramic depth imaging. Through a neural network model algorithm for two-stage cyclic image defocus recognition and restoration, it completes the clarification of defocus images and solves the problem of the mismatch between the effective depth of field of traditional microscopic equipment and the flow channel depth in large-range detection systems such as flow cytometry systems.
[0014] The technical solution of the present invention is specifically introduced as follows.
[0015] The present invention provides a method for defocus recognition and restoration of a microscopic system, including the following steps:
[0016] 1) Preparation of a microscopic defocus image dataset and degradation modeling
[0017] Prepare two microscopic defocus image datasets for training. One dataset is a simulated dataset, which is a blurred dataset with randomly defocused distances based on in-focus particle data of a flow cytometer. The other dataset is a real-shot dataset, which is a real-shot blurred particle dataset taken in a microscopic imaging system;
[0018] The microscopic defocus image datasets used for training are all pre-processed including background removal before training;
[0019] 2) Construction and training of a microscopic image defocus prediction and restoration network
[0020] The microscopic image defocus prediction and restoration network includes a defocus parameter prediction network and an image restoration network, and is connected by convolution and deconvolution operations based on the point spread function (PSF) blur kernel to form a cyclic structure; the defocus parameter prediction network is used to predict the defocus parameters of a blurred image based on multi-label parameter inference and prior knowledge of the microscopic system, and the image restoration network restores the blurred image based on a generative adversarial network; the defocus parameter prediction network includes an image feature extractor and a multi-layer perceptron (MLP). After the input of the defocus parameter prediction network passes through the image feature extractor, it outputs continuous defocus parameter prediction values through the multi-layer perceptron (MLP) to simulate the Zernike polynomial coefficients. The defocus parameters include defocus distance, astigmatism, spherical aberration, coma, and tilt; the prediction result of the defocus parameter prediction network is transformed into a PSF convolution kernel tensor of the wavefront function, and deconvolution operation is performed on the image, which is used as the input of the image restoration network.
[0021] A simulation dataset is used to train the constructed microscopic image defocus prediction and restoration network. First, the defocus parameter prediction network and the image restoration network are trained independently, and then the defocus parameter prediction network and the image restoration network are connected by convolution and deconvolution operations based on the PSF blur kernel to form a cyclic structure for joint cyclic training.
[0022] 3) Model fine-tuning
[0023] According to the model of the microscopic imaging system to be trained, update the parameters used to calculate the model wavefront function in the system, and use the actual shooting dataset to fine-tune the model based on the existing model with a reduced learning rate.
[0024] 4) Testing
[0025] After preprocessing the background removal of the defocus microscopic data to be detected, input it into the trained microscopic image defocus prediction and restoration network to obtain the prediction result of the target defocus parameters and the restoration result of the blurred image.
[0026] In the present invention, in step 1), the simulation dataset is based on Zemax simulation software, with a step size of 2 μm, simulating the parameters of the 4th-order Zernike polynomial within the range of ±50 μm on the z-axis of the microscopic system, fitting the relationship function with the defocus distance and then randomly sampling, and converting it into a PSF convolution kernel; the actual shooting dataset is an actual shooting blurred particle dataset taken at a step size of 2 μm in the microscopic imaging system.
[0027] In the present invention, in step 1), the image preprocessing method is as follows: the particles are cropped from the original background according to the adjacent rectangle and randomly rotated, the background brightness is uniformly adjusted to a fixed value, and then they are tiled in sequence on a unified background image with a size of 256 * 256 pixels and added with random Gaussian white noise with an intensity variance of 1 and the background brightness is unified.
[0028] In the present invention, in step 2), the defocus parameter prediction network uses ResNeSt50 as the backbone of the feature extractor, and the input end adopts a multi-channel input of the blurred image RGB input and its frequency domain amplitude spectrum and phase spectrum; the loss function for training the defocus parameter prediction network is l2loss, that is, the loss function for calculating the weighted mean square error, which emphasizes the absolute value of the defocus distance, takes into account the astigmatism, spherical aberration, coma, and tilt losses, and is premised on the judgment of whether the image is a simple background image.
[0029] In the present invention, in step 2), an optimized Real-ESRGAN model is adopted in the image restoration network, and a residual dense connection RRDB module with a self-attention mechanism is added to its generator; the loss function used when training the image restoration network is a function combining three parts: pixel difference, perceptual loss, and GAN loss; among them, the pixel difference is the average pixel value difference, structural similarity difference, and artifact constraint in the previous and subsequent training between the generated image of the generative adversarial network and the Ground truth; the perceptual loss is the mean error between the feature map content extracted by the pre-trained VGG network and the generated image of the generative adversarial network; the GAN loss is the MSE difference loss between the discriminator's judgment of the true and false image labels and the actual labels.
[0030] In the present invention, in step 2), the specific method for connecting the convolution and deconvolution operations based on the point spread function PSF blur kernel to form a loop structure includes: the network-generated image G of the image restoration network F After deconvolution, it is convolved again with the point spread function PSF inferred by the defocus parameter prediction network to obtain the loop detection blurred image L R ; the real in-focus image GT ((GroundTruth) is also convolved with the point spread function PSF inferred by the defocus parameter prediction network after deconvolution to generate the pseudo-blurred image L F , the pseudo-blurred image L F After deconvolution again, the steps of generative adversarial restoration are performed and compared with the real in-focus image GT to complete the loop feedforward process of the network.
[0031] In the present invention, during joint loop training, the loop loss is composed of the pixel MSE loss and SSIM loss of the image pair, which are respectively the losses between the loop blurred image L F formed by the PSF recognized by the convolution of the network-generated image G R and the original blurred image L, the losses between the pseudo-blurred image L F generated by the real in-focus GT image through the recognized PSF convolution and the original blurred image L F , and the losses between the loop high-quality image G R formed by the same RL deconvolution and image restoration steps of the pseudo-blurred image L and the real in-focus image GT.
[0032] In the present invention, in step 3), the microscopic system parameters are directly modified as training parameters, and based on the data pairs of in-focus and out-of-focus images and the out-of-focus distance labels in this system, a new data set is simulated or collected to fine-tune the model; wherein, the microscopic system parameters include the imaging center wavelength, the camera pixel size, the physical aperture, the magnification, and the PSF convolution kernel sampling size. Further, during fine-tuning, in addition to the actual shooting data set, data sets for other systems can also be used according to the differences in the systems.
[0033] In the present invention, in step 4), a step-by-step analysis and restoration are performed based on the field of view according to a sliding window with a size of 256 * 256 pixels; the final image restoration result is weighted and averaged according to the image results restored by each sliding window, or the restored patches of particles in different orientations are tiled onto the same large image to complete the final restoration of the image.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1) Compared with the traditional end-to-end out-of-focus restoration generative network, the present invention connects a convolution and deconvolution operation based on the point spread function (PSF) blur kernel to form a loop structure. The system directly blurs the ground truth in-focus image using a PSF convolution kernel generated based on Zernike polynomials, transforming the inference of a single blur mode in the traditional loop generative adversarial method into an inference behavior of identifying multiple blur methods, thus expanding the inference range. At the same time, in the process of converting a blurred image into a clear image, the present invention performs deconvolution processing on the image based on the identified blur parameters, which helps to partially remove the influence of out-of-focus, and the GAN repairs the distortion in a smaller range after processing, improving the adaptability of this part. Therefore, the method of the present invention can adapt to a wider range of out-of-focus blurs and is more sensitive to the blurred particle contours in the flow imaging system.
[0036] 2) Compared with the traditional generative network restoration, the present invention combines an out-of-focus prediction network, no longer implicitly simulating the unknown image blur process, but explicitly outputting the out-of-focus image system blur coefficient, i.e., the 4th-order Zernike polynomial coefficient, which enhances the interpretability and can stably and explicitly analyze the 4th-order Zernike polynomial coefficient of the out-of-focus image system to describe the cause of the target out-of-focus blur.
[0037] 3) The present invention optimizes the loss functions such as pixel difference, perceptual loss, and GAN loss for the network in the out-of-focus image restoration, and changes the training conditions, which can more effectively suppress the possible artifacts in the out-of-focus image restoration, making the image restoration result closer to the real in-focus image, and can adapt to other microscopic systems through the correction and fine-tuning of the microscopic system coefficients, having a certain degree of versatility and portability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 : Overall structure diagram of the defocus prediction and restoration network for microscopic images.
[0039] Figure 2 : Structure diagram of the defocus parameter prediction network.
[0040] Figure 3 : Structure diagram of the image restoration network, (a) Schematic diagram of the overall GAN network; (b) Schematic diagram of the RRDAB structure.
[0041] Figure 4 : Schematic diagram of the network loss function.
[0042] Figure 5 : Verification result diagram of the system on FlowCAM data, (a) Schematic diagram of particle image restoration; (b) Edge gradient diagrams of particles before and after restoration in Con mode; (c) Edge gradient diagrams of particles before and after restoration in PS&MO mode; (d) Statistical comparison of particle sizes before and after restoration of 10-μm particles and size difference before and after restoration. Specific implementation manner
[0043] The technical solution of the present invention will be elaborated in detail below in conjunction with the accompanying drawings and embodiments.
[0044] Particle imaging in systems such as flow cytometry imaging usually faces problems such as blurred distortion and errors in morphological statistics such as size and contour, and it is necessary to develop an algorithm-based depth of field expansion scheme optimized by deep learning technology. Traditional depth of field expansion methods need to cooperate with microscopic equipment for re-focusing calculation, or have problems such as insufficient versatility, weak interpretability, and obvious artifacts in the processed images, and it is necessary to develop a method for single-shot imaging clarity of the entire flow channel depth of the system.
[0045] The present invention provides a method for defocus recognition and restoration of a microscopic system, including the following steps:
[0046] Step 1, preparation of a microscopic defocus image dataset and degradation modeling
[0047] The main object of network processing is the panoramic deep defocus particles in various streaming imaging systems. However, due to the limitations of sampling time and sample types, the actual shot dataset cannot be complete. The simulated dataset based on the defocus parameters of the microscopic system is generated from the in-focus 2D images captured by commercial streaming systems. However, the particulate matter often has different morphologies within the 3D space range, and the simulated data may ignore the morphological changes of the particles on the z-axis. Therefore, a training method using two datasets is adopted. Based on the Zemax simulation software, with a step size of 2μm, the parameters of the 4th-order Zernike polynomial within the range of ±50μm on the z-axis of the microscopic system are simulated. After fitting the relationship function with the defocus distance, random sampling is performed and converted into a PSF convolution kernel. Based on the in-focus particle data of the streaming imager, a blurred dataset with random defocus distances is simulated for large-scale morphological restoration training. Subsequently, the network is further fine-tuned according to the actual shot blurred particle dataset captured at a step size of 2μm in the microscopic imaging system to complete the training. The dataset needs to undergo preprocessing steps such as background removal.
[0048] Since the field of view in the actual shot dataset far exceeds the range that the paraxial PSF can withstand, according to three different image plane regions: paraxial, approximately 500μm from the axis, and approximately 707.11μm from the axis, the corresponding 4th-order Zernike polynomial coefficients are fitted in the Zemax software respectively, which serves as the basis for different region labels in the image plane field of view. The simulated dataset also generates blurred images at different positions according to these three regions respectively.
[0049] The following method is used for image preprocessing: The particles are randomly rotated after being cropped from the original background according to the adjacent rectangle. The background brightness is uniformly adjusted to a fixed value, and then they are tiled in order on a unified background image with a size of 256*256 pixels and added with random Gaussian white noise with an intensity variance of 1, and the background brightness is unified. Multiple particles from the same type can be tiled independently or merged onto the same background to save space.
[0050] Step 2: Design and training of the network model, including the following sub-steps:
[0051] Step 2.1: Design and training of the defocus parameter prediction network
[0052] The first part of the network is the defocus parameter prediction network, which is used to predict the defocus parameters of a blurred image, including defocus distance, astigmatism, spherical aberration, coma, tilt, etc. This part constructs a prediction and inference network for Zernike coefficients based on a ResNest50 feature extractor (backbone) and a multi-layer perceptron head (head), and first conducts independent training based on the dataset in step one to form a relatively stable first-step model. The judgment result is compared with the corresponding Zernike coefficient label, and the loss function is the l2loss, that is, the loss function for calculating the weighted mean square error. The judgment result of this step of the network is Fourier-transformed into a PSF convolution kernel tensor through the wavefront function tensor and then deconvolved to obtain the input of step 2.2.
[0053] The defocus parameter prediction network uses ResNeSt50 as the backbone of the feature extractor. The input end adopts a 5-channel input that concatenates the original RGB input with its frequency domain amplitude spectrum and phase spectrum. The output head uses a 3-layer multi-layer perceptron (MLP) to output continuous parameter prediction results instead of discrete label classification. Its loss function uses the mean square error value that emphasizes the absolute value of the defocus distance and takes into account the losses of other parameters, and takes the judgment of whether the image is a simple background image as a prerequisite.
[0054] Step 2.2, Design and training of the image restoration network
[0055] This part of the network is optimized based on the Real-ESRGAN generative adversarial model, adding a loss function for further constraint processing of artifacts, and adding an attention mechanism to the convolutional structure, so as to further emphasize and preserve the texture details of the image. During training, the input uses the blurred image deconvolved according to the real defocus parameter label to compare with the real in-focus image to train the restoration performance of the network. At this stage, the network also conducts independent training first, and after forming a relatively stable model, it enters step 2.3.
[0056] The generator of the image restoration network uses a single deep learning feature extraction network. The basic framework is the RRDB module with self-attention mechanism, which is used to better extract feature information from feature maps of different dimensions; its discriminator is a deep learning classifier using the UNet architecture; the loss function used when training the generative network is a function that combines pixel difference, perceptual loss, and GAN loss. Among them, the pixel difference is the average pixel value difference, structural similarity difference between the generated result and the Ground truth, and the artifact constraint in two consecutive trainings; the perceptual loss is the mean error between the feature map content extracted by the pre-trained VGG network and the generated image; the GAN loss is the MSE difference loss between the discriminator's judgment of the true and false image labels and the actual labels.
[0057] The deconvolution structure used for deconvolution processing of blurred images based on true defocus parameter labels performs an image negation operation first, compared with traditional Lucy-Richardson deconvolution iteration, to emphasize the contours of particles darker than the background in bright-field images. This deconvolution is preset to perform 30 rounds of iteration.
[0058] Step 2.3, Joint cyclic training
[0059] Here, the above two parts of the network are combined through a cyclic structure as shown in Figure 1 for joint training to achieve the effect of further optimizing and stabilizing the inference results. The connection between the two modules here needs to be realized through additional convolution and deconvolution operations. The image restoration result is convolved with the PSF inferred by the network again for cyclic verification of the two models. In addition, the true in-focus image is also convolved with the predicted PSF and compared with the original blurred image. The blurred image generated through this step undergoes deconvolution and generative adversarial restoration steps again and is compared with the in-focus image to complete the cyclic feedforward process of the network. During this cyclic process, the MSE and image structure similarity (SSIM) losses between several groups of generated images are calculated to further ensure the supervision and constraint of system image restoration.
[0060] Step Three: Model fine-tuning based on a specific system and dataset
[0061] According to the model of the microscopic imaging system to be trained, microscopic parameters such as the imaging center wavelength, camera pixel size, physical aperture, magnification, and the sampling size of the required PSF convolution kernel used to calculate the model wavefront function in the system need to be updated. Based on the modification of the internal parameters of the system and the data pairs of some in-focus - defocus images, defocus distances, etc. in this system, a new dataset is simulated or collected, and further fine-tuning is performed based on the existing model by reducing the learning rate.
[0062] The objective of the present invention is to develop a defocus processing algorithm with universality that can be applied in multiple microscopic imaging systems. The new dataset here refers to the image data for training obtained again when applying to a new microscopic system or a new type of sub-visible particle other than the exceptions.
[0063] Step Four: Testing
[0064] The defocus microscopic data to be detected is sheared by a sliding window of 256*256. After performing the same background removal operation on the particles to be restored, it is input into the network to obtain the prediction result of the target defocus parameter and the restoration result of the blurred image, which is used for further analysis of the flow system such as particle statistics, size statistics, etc. The restoration results can be integrated into a large image through the method of average superposition. At the same time, the system's judgment of the image defocus parameter will also be explicitly saved and output in the form of an array record.
[0065] It can be gradually analyzed and restored according to a sliding window of 256*256 pixels in a larger field of view. The final image restoration result can be obtained by weighted averaging the image results restored from each sliding window, or by tiling the restored images of particles in different orientations onto the same large image to complete the final restoration of the image.
[0066] The following provides a more specific implementation description.
[0067] A method for defocus identification and repair of a microscopic system, which uses a two-stage cyclic microscopic image defocus prediction and repair network as shown in Figure 1 . The network is composed of a defocus parameter prediction network based on the compound of time domain and frequency domain channels, and an image repair network based on the generative adversarial network (GAN) technology, which are spliced into a two-stage network system, and a cyclic structure is formed by connecting the convolution and deconvolution operations based on the point spread function (PSF) blur kernel; the method includes the following steps:
[0068] Step 1: Preparation of the microscopic defocus image dataset and degradation modeling. Here, the modeling of the image degradation in the microscopic system is based on the prior knowledge of the 4th-order Zernike polynomial coefficients of the system:
[0069]
[0070] where s is the coordinate on the camera image plane, and a n (s) is the coefficient of the nth term Z n of the corresponding Zernike polynomial, represents the vector formed by the origin of the object plane and the coordinates of the traced point in the object plane, and the coefficients of the 4th-order polynomial can be approximated to stop at 15 terms. According to the above wavefront function composed of the Zernike polynomial, the PSF function can be formed as follows, where λ refers to the central wavelength of the light used in the system:
[0071]
[0072] By reorganizing the above 15-term polynomial coefficients according to their relationship with the coordinates, a new wavefront function with specific physical meanings for each coefficient can be constructed as follows:
[0073] Wf(ρ,θ) = e A(ρ,θ)
[0074] ρ 2 = x 2 + y 2
[0075] A(ρ,θ) = a defocus ρ 2 + a ast (xcos(θ) + ysin(θ)) 2 + a sphρ 4 +(a coma ρ 2 +a tilt )(x cos(θ)+y sin(θ))
[0076] Among them, (x, y) represents the pixel position on the target image coordinate system, ρ is the distance from this pixel to the center, and θ is the polar coordinate angle of the image center relative to the image plane center. The PSF convolution kernel is obtained by performing Fourier transform on the wavefront function Wf(ρ, θ) and processing its amplitude, etc. In the annotation, the annotation that needs to be performed on the blurred image is a series of coefficients a in the above formula (including a defocus , corresponding to the defocus coefficient; a ast , representing the astigmatism coefficient; a sph , representing the spherical aberration coefficient; a coma , representing the coma coefficient, a tilt , representing the tilt coefficient) and θ, as well as the parameter a bkg for predicting whether the image is a simple background image.
[0077] When preparing the training set, it is necessary to perform background removal on the image, cut out the blurred particles in the dataset through adjacent rectangles, flatten the background, and paste them onto a pure background image of 256 * 256 pixels with a fixed background. In addition, when the particles are dense and there are particles with a large difference in focal plane in the same area, the network system tends to identify the defocus parameters of the particles closer to the focal plane as the final output result. Therefore, it is necessary to separate and cut out the particles with too large a focal plane difference, split and paste the particles with significantly different clarity onto two backgrounds, and then perform test restoration to ensure the stability of the network's defocus parameter recognition.
[0078] Step 2, design and training of the network model, including the following sub-steps:
[0079] Step 2.1, design and training of the defocus parameter prediction network. When sampling and inputting the dataset here, first perform Fourier transform on the image grayscale value to obtain the amplitude spectrum and phase spectrum in the image frequency domain, and splice them with the original RGB channels to form a 5-channel input. Since the frequency domain image is linearly related to the parameters in the wavefront function and is more sensitive to the recognition of defocus parameters, the network can more stably recognize the relevant defocus parameters.
[0080] Furthermore, as Figure 2, in this part of the network, ResNeSt50 is adopted to replace the traditional ResNet50 as the image feature extractor. By using the multi-scale attention mechanism of the ResNeSt system, the feature information of the image at multiple scales can be emphasized more accurately. At the output head part of the network, a 3-layer MLP structure is adopted to output continuous parameter predictions more stably. The loss function of this step adopts weighted MSE. However, considering that the out-of-focus blur on both sides of the focal plane is approximately symmetric, and the out-of-focus blurred speckle images are closer to each other as they are farther away from the focal plane, and also considering that the calculation of positive and negative values is very likely to cause network confusion, the proportion of the absolute value of the defocus distance is increased, while the proportion of values such as θ that contribute less to the out-of-focus blur is reduced. Considering that there are no clear particles in the pure background image that can be used to judge defocus, it is very likely to cause ambiguity here, so it is listed separately and represented by a parameter value a bkg Distinguish. The final loss function can be written as:
[0081]
[0082] Among them, the superscript k represents the prediction result of a batch of image data in the training, and the superscript represents the corresponding true label data of this batch of image data. The meanings of data such as a bkg , a defocus are as described above, and a n includes a ast , a sph , a coma , a tilt , as well as the θ value, which is the parameter prediction result.
[0083] In the training of step 2.1, the initial learning rate is set to 0.01. The stage optimization method is adopted, and the learning rate is reduced to 0.2 of the original at the 50th and 100th epochs respectively. It is trained for 150 epochs, and the batch size is set to 16 during training. For the fine-tuning training of the actual shooting dataset, the initial learning rate is reduced to 0.001.
[0084] Step 2.2, training of the image restoration network. Here, the LowQuality dataset that has undergone blur prediction and deconvolution iteration is input into the GAN network for restoration. The generator adopts the residual dense connection module (RRDB) of the Real-ESRGAN model with pruning and added attention mechanism, that is, RRDAB, as Figure 3 , to fully extract and emphasize the contour and texture information of the particles to be analyzed. In the discriminator part, a 10-layer U-Net design is used to perform pixel-level authenticity determination of image details.
[0085] Furthermore, during network training, the loss functions involved in the network are optimized to varying degrees. For the perceptual loss, a VGG model pre-trained on streaming imaging particle data is used to replace the publicly available ImageNet pre-trained VGG19 model to determine feature extraction that is more consistent with particle semantics for the contour; for the GAN loss, to address the problem of gradient disappearance caused by BCEWithLogits in the general model during training and to accelerate network convergence, l2 loss is used as a replacement, where G(x) represents the generation result of the input low-quality image x, D(G(x)) represents the discriminator's judgment on the similarity between this generation result and the real in-focus image, and D(GT) represents the discriminator's judgment on the authenticity of the real in-focus image itself:
[0086]
[0087] For the pixel-level pixel loss, in addition to comparing the MSE between the generated image and the real in-focus image, the concept of structural similarity (SSIM) is introduced to compare information such as the brightness, contrast, and dynamic changes in pixel values of the two images, ensuring that the overall structure of the image remains unchanged. Among them, μ lq is the mean of the blurred image restoration result, μ gt is the mean of the real in-focus image for comparison, σ lq , σ gt are the standard deviations of the two groups of images themselves, and σ lq-gt is the covariance of the two groups of images:
[0088]
[0089] In addition, to address the common artifact problem in GAN image restoration, during the separate training of this part of the network, exponential moving average (EMA) training is introduced. In addition to the image generated by the model with the current training parameters, an image generated by weighted averaging the parameters of the previous training batches is used as a control, and LDL loss is introduced:
[0090] L pixel-LDL =||M refine L l1 ||, R=I GT -I Gen
[0091]
[0092] That is, the residuals, i.e., artifacts, between the above two images and the real in-focus image are compared. If the high-frequency artifact information generated by the current training model is relatively severe, a penalty term is incorporated. In the above formula, L l1represents the mean error between pre-computed images. k refers to the neighborhood size required when using a sliding window to calculate the LDL loss for a single pixel, and I GT and I Gen respectively represent the tensors formed by the real in-focus image and the generated image; the calculations of the remaining intermediate variables such as the residual R, M refine are given by the above formula. Thus, in step 2.2, more specifications for image contour and texture restoration are introduced for image inpainting. The network in this stage is trained for 32000 iterations, the batch size is set to 8, and the initial learning rate is 4*10 -4 , and stage-wise optimization is also adopted. The learning rate is reduced to 0.25 of the original value at the 8000th and 24000th iterations.
[0093] Step 2.3, the second round of training, splices the above two-stage models using a loop structure and continues to train for 4000 - 8000 rounds. At this time, the batch size is uniformly set to 8. In the loop, in addition to the above loss functions, loop losses such as Figure 4 are added. The loop losses are all composed of the pixel MSE loss and SSIM loss of the image pair, respectively considering the losses between the loop-blurred image L F formed by the PSF identified by the convolution of the network-generated image G R and the original blurred image L, the losses between the pseudo-blurred image L F formed by the real in-focus image GT (Groundtruth) through the identified PSF convolution and L, and the losses between the loop high-quality image G F formed by the same RL deconvolution and image restoration steps of L R and the Ground Truth.
[0094] Step three: Model fine-tuning. For the new measured data set, image preprocessing is also carried out according to the above steps; the basic parameters of the new microscopy imaging system are modified in the settings of the program wavefront parameter calculation, and accordingly, the model for the new system is trained on the new image training set with a relatively small learning rate such as 1*10 -4 .
[0095] Step four: Input the preprocessed new microscopy imaging data to be detected into the network, obtain the prediction and restoration results, and conduct further analysis, such as flow imaging particle statistics, size statistics, etc.
[0096] Example 1
[0097] In the steps of the present invention, the network architecture is uniformly written for parameters, optimization methods, and training steps based on the Python 3.8.10 version and the OpenMMlab format. Training and testing are carried out on a server based on the Ubuntu20.04 system, using Cuda11.3 and an RTX4090D graphics card.
[0098] For the commercial flow imaging system FlowCAM, under its two modes of counting (Con) and calibration (PS&MO), different particle conditions are photographed respectively to observe the improvement of image quality and analysis effect. For a 6μm polystyrene microsphere particle, the sharpening ability of the system is analyzed. After obtaining the microsphere particle image, preprocessing is carried out according to the specific implementation steps, and then it is input into the restoration network for image restoration. Before and after image restoration, the LoG operator is used to extract the change of the particle edge gradient under the two modes, as Figure 5 (as shown in (b - c)), there is an obvious improvement in the particle edge gradient. Among a total of 78,870 particles in the Con mode, according to the gradient standard of the PS&MO mode, more than 20,000 particles in the original data need to be ignored, while the number of such particles in the restored image drops to 220.
[0099] For the ThermoFisher Duke series standard microspheres of 10μm (10.0 ± 0.09μm) and 5μm (4.993 ± 0.04μm), the particle sizes in the original image and after model restoration are statistically analyzed. The average diameter of the 10μm microsphere image before restoration is 10.8409 ± 1.4122μm, and the variance is 0.0462μm 2 , and two relatively obvious peaks can be seen in the histogram; while after restoration, the average diameter is 10.5674 ± 1.3899μm, and the variance is 0.0819μm 2 , although the two statistical peaks still exist, the spacing between them decreases, and the peak corresponding to the 10μm standard diameter significantly increases. The average difference in diameter between the same particles before and after restoration reaches 0.2838μm. For the 5μm particles, the average particle sizes before and after restoration are 7.3652 ± 1.5037μm and 6.7895 ± 1.4539μm respectively. The average difference in particle size of the same microspheres before and after restoration reaches 0.5848μm. After restoration, the particle diameter is closer to the standard value of 5μm, which proves that the system detects more obvious defocus blur changes and performs restoration, which also conforms to the actual detection situation where due to using a 15μm standard bead for focusing, the 5μm microspheres may be far from the most suitable focal plane during detection, and are affected by more obvious edge diffraction interference and pixel size during imaging. Thus, it can be proved that the system effectively restores particle defocus in microfluidic imaging, making the statistical values of particle size, particle count, and other traits affected by defocus more accurate.
[0100] Above, the present invention focuses on the out-of-focus blur problem caused by the fact that the depth of the flow channel in the microfluidic system far exceeds the effective depth of field of the system, and constructs an out-of-focus parameter evaluation and image restoration post-processing network for out-of-focus particles. This network integrates an out-of-focus parameter prediction network and an image restoration network, and can be used for the recognition and restoration of out-of-focus blur in the microscopic imaging of various systems. The out-of-focus parameter prediction network predicts out-of-focus distortion parameters based on multi-label parameter reasoning and prior knowledge of the microscopic system. The image restoration network restores low-resolution images. The out-of-focus parameter prediction network of the microscopic system and the deep learning generative adversarial network are interconnected through convolution and deconvolution operations based on the predicted point spread function to form a circulatory system. The present invention helps to accurately perform functions such as system particle counting and particle size statistics, thereby further promoting the application of the panoramic depth-of-field microscopic system in biomedical detection.
Claims
1. A method for defocus identification and repair of a microscopic system, characterized in that, Including the following steps: 1) Preparation of the microscopic defocus image dataset and degradation modeling Prepare two microscopic defocus image datasets for training. One dataset is a simulated dataset, which is a blurred dataset with randomly defocused distances simulated from in-focus particle data based on a flow imager; the other dataset is a real-shot dataset, which is a real-shot blurred particle dataset captured in a microscopic imaging system. The microscopic defocus image datasets used for training are all pre-processed including background removal before training. 2) Construction and training of the microscopic image defocus prediction and restoration network The microscopic image defocus prediction and restoration network includes a defocus parameter prediction network and an image restoration network, and is connected by convolution and deconvolution operations based on the point spread function (PSF) blur kernel to form a loop structure. The defocus parameter prediction network is used to predict the defocus parameters of the blurred image based on multi-label parameter inference and prior knowledge of the microscopic system. The image restoration network restores the blurred image based on the generative adversarial network. The defocus parameter prediction network includes an image feature extractor and a multi-layer perceptron (MLP). After the input of the defocus parameter prediction network passes through the image feature extractor, it outputs continuous defocus parameter prediction values through the multi-layer perceptron (MLP) to simulate the Zernike polynomial coefficients. The defocus parameters include defocus distance, astigmatism, spherical aberration, coma, and tilt. The prediction result of the defocus parameter prediction network is transformed into a PSF convolution kernel tensor through wavefront function transformation, and deconvolution iterative operation is performed on the image, which is used as the input of the image restoration network. Use the simulated dataset to train the constructed microscopic image defocus prediction and restoration network. First, train the defocus parameter prediction network and the image restoration network independently, and then connect the defocus parameter prediction network and the image restoration network through convolution and deconvolution operations based on the PSF blur kernel to form a loop structure for joint loop training. 3) Model fine-tuning According to the model of the microscopic imaging system to be trained, update the parameters used to calculate the model wavefront function in the system, and use the real-shot dataset to fine-tune the model based on the existing model with a reduced learning rate. 4) Testing After pre-processing the defocus microscopic data to be detected by removing the background, input it into the trained microscopic image defocus prediction and restoration network to obtain the prediction result of the target defocus parameters and the restoration result of the blurred image.
2. The defocus identification and repair method according to claim 1, characterized in that In step 1), the simulated dataset is Based on Zemax simulation software, with a step size of 2 μm, simulate the parameters of the 4th-order Zernike polynomial within the range of ±50 μm on the z-axis of the microscopic system, fit the relationship function with the defocus distance, and then randomly sample to obtain the PSF convolution kernel after conversion; the real-shot dataset is a real-shot blurred particle dataset captured in the microscopic imaging system with a step size of 2 μm.
3. The defocus identification and repair method according to claim 1, characterized in that In step 1), the image pre-processing method is as follows: The particles are randomly rotated after being cropped from the original background according to the adjacent rectangle, the background brightness is uniformly adjusted to a fixed value, and then they are tiled in order on a unified background image with a size of 256 * 256 pixels and added with random Gaussian white noise with an intensity variance of 1, and the background brightness is unified.
4. The defocus identification and repair method according to claim 1, wherein In step 2), the defocus parameter prediction network uses ResNeSt50 as the backbone of the feature extractor, and the input end adopts a multi-channel input that concatenates the RGB input of the blurred image with its frequency-domain amplitude spectrum and phase spectrum. The loss function for training the defocus parameter prediction network is l2loss, that is, the loss function for calculating the weighted mean square error, which emphasizes the absolute value of the defocus distance, takes into account astigmatism, spherical aberration, coma, and tilt losses, and is premised on the judgment of whether the image is a simple background image.
5. The defocus identification and repair method according to claim 1, characterized in that In step 2), an optimized Real-ESRGAN model is adopted in the image restoration network, and a residual dense connection RRDB module with a self-attention mechanism is added to its generator; the loss function used when training the image restoration network is a function that combines pixel difference, perceptual loss, and GAN loss; among them, the pixel difference is the average pixel value difference, structural similarity difference, and artifact constraint in the previous and subsequent training between the generated image of the generative adversarial network and the real in-focus image GT; the perceptual loss is the mean error between the feature map content extracted by the pre-trained VGG network and the generated image of the generative adversarial network; the GAN loss is the MSE difference loss between the discriminator's judgment of the true and false image labels and the actual labels.
6. The defocus identification and repair method according to claim 1, wherein In step 2), the specific method of connecting convolution and deconvolution operations based on the point spread function PSF blur kernel to form a loop structure includes: the network generated image G of the image inpainting network F After deconvolution, it is convolved again with the point spread function PSF inferred by the defocus parameter prediction network to obtain the loop detection blurred image L R ; the real in-focus image GT also generates a pseudo-blurred image L after deconvolution and convolution with the point spread function PSF inferred by the defocus parameter prediction network F , the pseudo-blurred image L F After deconvolution again, the steps of generative adversarial inpainting are compared with the real in-focus image GT to complete the loop feedforward process of the network.
7. The defocus identification and repair method according to claim 6, characterized in that During joint cycle training, the cycle loss consists of the pixel MSE loss and the SSIM loss of the image pair, which are the losses between the image G F formed by the PSF identified by convolution and the loop-blurred image L R and the original blurred image L, the losses between the pseudo-blurred image L obtained by convolving the real in-focus image GT with the identified PSF F and the original blurred image L, and the loop high-quality image G F formed by the same RL deconvolution and image restoration steps R and the real in-focus image GT.
8. The defocus identification and repair method according to claim 1, wherein In step 3), the microscopic system parameters are directly modified as training parameters, and based on the data pairs of some in-focus and defocus images and the defocus distance labels in this system, a new dataset is simulated or collected, and the model is fine-tuned; among them, the microscopic system parameters include the imaging center wavelength, camera pixel size pixel size, physical aperture, magnification, and PSF convolution kernel sampling size.
9. The defocus identification and repair method according to claim 1, characterized in that, In step 4), based on the field of view, step-by-step analysis and restoration are carried out according to a sliding window of 256*256 pixel size; the final image restoration result is weighted and averaged according to the image results restored by each sliding window, or restored and tiled into the same large image according to the particles in different orientations to complete the final restoration of the image.
Citation Information
Cited By
Automatic focusing method based on micro flow cell, electronic equipment and storage medium
CN121037694A
Micro-flow cell-based autofocusing method, electronic device, and storage medium
CN121037694B
Fluorescence signal recovery method and system including defocused background filtering
CN121053029A
A method and system for fluorescence signal recovery including defocus background filtering
CN121053029B
Efficient automatic focusing method and system
CN121603779A