Efficient universal pathological image deblurring method and device based on diffusion model

By combining the pathological image defuzzing method of DTransformer and diffusion model, the problems of high computational cost and insufficient generalization ability in the prior art are solved, and efficient and generalized pathological image defuzzing is achieved, and the generated image quality is significantly improved to adapt to pathological images of different tissues and cell types.

CN120235784APending Publication Date: 2025-07-01WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510606107.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing pathological image defuzzing methods have high computational cost, insufficient generalization ability, poor image generation quality, and lack of modeling of tissue structure and blur directionality, resulting in increased diagnostic errors.

Method used

Using an efficient and general method combining DTransformer and diffusion model, the feature extraction network, cross-transpose attention module CTA and neighborhood aggregation feedforward neural network NAFN are used to enhance feature interaction and spatial expression capabilities, and the diffusion model is used to restore clear image features through iterative denoising process.

Benefits of technology

Efficient and universal pathological image defuzzing is achieved, which significantly improves the accuracy and quality of image generation, reduces artifacts, enhances the generalization ability of the model, and adapts to pathological images of different tissues and cell types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235784A_ABST
    Figure CN120235784A_ABST
Patent Text Reader

Abstract

The invention provides an efficient general pathological image deblurring method and device based on a diffusion model. The method comprises the steps of obtaining a to-be-processed out-of-focus blurred pathological image; inputting the image into a pre-training deblurring model, and outputting a clearly focused pathological image; wherein the pre-trained deblurring model comprises a feature extraction network, a DTransform model and a diffusion model; the feature extraction network is used for compressing an original image to a low-dimensional potential space; the DTransform model comprises a cross transpose attention module (CTA) and a neighborhood aggregation feedforward neural network (NAFN), and is used for enhancing feature interaction; the diffusion model is used for restoring clear image features through an iterative denoising process. The problems of high calculation cost, insufficient generalization ability, poor image generation quality and the like of the traditional pathological image deblurring method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - field of deep learning and biomedicine, and particularly relates to a method for de - blurring whole - slide pathological images, belonging to the application of deep learning in the medical field. Background Art

[0002] Medical images play an important role in clinical medicine. Computed tomography, X - ray radiography, magnetic resonance imaging, and optical imaging are common and indispensable medical imaging modalities, which are used for diagnosis, guiding treatment decisions, and monitoring treatment responses. As an advanced medical diagnosis mode, it plays an increasingly important role in patient care and is now used for non - invasive cancer detection, brain tumor diagnosis, and surgical specimen analysis, etc.

[0003] However, the basic prerequisite for imaging is to capture high - quality and fully - focused images at high speed. Defocus blur is the main source of low - quality imaging. Blurred defocused images have a significant negative impact on clinical diagnosis, which may lead to misdiagnosis and misleading treatment. At the same time, there are many reasons for defocus blur: for example, the scattering of incident light on tissues causes blur; biological tissues contain pigments that absorb light of specific wavelengths, reducing the signal intensity and thus causing defocus blur; interference from external instruments and electronic components, etc. Existing image de - blurring technologies mainly focus on the de - blurring of natural images, and there is less research on the de - blurring of pathological images. The cells and tissue structures in pathological images are very complex, and the model needs to be able to learn and accurately restore these microscopic structures. Therefore, it is very necessary to explore a general and efficient method for de - blurring pathological images.

[0004] In recent years, with the development of deep learning technology, the blurring method based on pathological images has changed from traditional methods based on filtering, interpolation, and models to a neural network-based pathological image deblurring method. Although traditional methods can improve the clarity of images to a certain extent, they often lead to the loss of image details or the introduction of artificial traces. The deep learning-based method can learn the complex features and structural information of images by constructing an end-to-end neural network model, thus achieving more accurate and efficient image deblurring. However, the current deep learning-based pathological image deblurring methods still have some problems, including insufficient generalization ability of the model, high computational cost, and lack of training data. Previous methods require a large number of paired low-quality / high-quality pathological images as a dataset for training, which greatly increases the difficulty of obtaining the dataset. Therefore, an unpaired image restoration task has emerged, such as CycleGAN, which is a generative network for unpaired image-to-image. However, when this method is applied to biomedical imaging, it may result in unrealistic and perceptually poor-quality reconstructions. Importantly, poor-quality reconstructions will have harmful downstream effects on medical imaging, increasing the risk of undiagnosed or misdiagnosed images. Therefore, an ideal pathological image deblurring method should only require a small amount of dataset and not produce unrealistic artifacts.

[0005] However, most of the current deep learning-based pathological image deblurring methods follow natural image processing strategies. Common methods such as DeblurGAN, EDVR, MPRNet, etc. have achieved good performance in natural image scenarios, but have the following deficiencies in pathological images: First, the structure is complex, the training is time-consuming, and the resource requirements are high; second, there is a lack of modeling of tissue structure and blurring directionality, and the processing results lack realism or consistency; third, the generalization ability is poor, and the model can often only adapt to images of specific staining or organ types. Summary of the Invention

[0006] In view of the above problems, the present invention proposes an efficient and general method combining DTransformer and diffusion model, which has the advantages of simple structure, strong modeling ability, and stable reconstruction quality, and can effectively improve the deficiencies of the above traditional methods.

[0007] The above technical problems of the present invention are mainly solved by the following technical solutions:

[0008] An efficient and general pathological image deblurring method based on a diffusion model, comprising the following steps:

[0009] Obtain a defocused and blurred pathological image to be processed;

[0010] Input the image into a pre-trained deblurring model and output a focused and clear pathological image;

[0011] Among them, the pre-trained deblurring model includes a feature extraction network, a DTransformer model, and a diffusion model; the feature extraction network is used to compress the original image into a low-dimensional latent space; the DTransformer model contains a cross transposed attention module CTA and a neighborhood aggregation feed-forward neural network NAFN, which is used to enhance feature interaction; the diffusion model is used to restore clear image features through an iterative denoising process.

[0012] Moreover, the processing process of the cross transposed attention module CTA is as follows:

[0013] Input image features and latent representations are used to generate a query matrix, a key matrix, and a value matrix;

[0014] Calculate the cross-modal attention map and weighted fuse it with the value matrix to obtain the output features.

[0015] Moreover, the neighborhood aggregation feed-forward neural network NAFN introduces a convolutional structure and a gating mechanism to enhance the spatial expression ability.

[0016] Moreover, the feature extraction network includes a residual module, a Patch mixing module, and a channel mixing module. The Patch mixing module models the spatial dimension of image patches through a fully connected layer, and the channel mixing module models the interdependence between channels through a fully connected layer.

[0017] Moreover, the training of the pre-trained deblurring model includes the following stages:

[0018] In the first stage, jointly train the feature extraction network and the DTransformer model;

[0019] In the second stage, jointly train the feature extraction network, the diffusion model, and the DTransformer.

[0020] Moreover, the diffusion model is composed of T Denoising U-Nets, where T is the diffusion step; the Denoising U-Net is a U-Net network structure for denoising, and each Denoising U-Net is composed of 4 cross-attention modules.

[0021] Moreover, the reverse denoising process of the diffusion model estimates clear image features through a noise predictor.

[0022] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned efficient and general pathological image deblurring method based on the diffusion model.

[0023] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the efficient and general pathological image deblurring method based on the diffusion model as described above is implemented.

[0024] On the other hand, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the efficient and general pathological image deblurring method based on the diffusion model as described above is implemented.

[0025] The technical solution proposed by the present invention has the following advantages:

[0026] On the one hand, by improving the traditional Transformer, the cross transposed attention module (CTA) and the neighborhood aggregation feed-forward neural network (NAFN) are proposed, enabling the model to learn from both the input image and the image features simultaneously, and enhancing the model's ability to learn image neighborhood features. On the other hand, through the regression-based image feature extraction model, the defect of low distortion accuracy of the generation model is improved, and the relationship between patches and channels is fully considered. Generally speaking, this technical solution solves the problems of high computational cost, insufficient generalization ability, and poor image generation quality in the previous deblurring methods through the combination of the diffusion model and the improved Transformer.

[0027] The implementation of the solution of the present invention is simple and convenient, with strong practicability. It solves the problems of low practicability and inconvenience in actual application in the related technologies, can improve the user experience, and has important market value. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the model architecture in the embodiment of the present invention.

[0029] Figure 2 It is a schematic diagram of the diffusion process of the diffusion model in the embodiment of the present invention.

[0030] Figure 3 It is the pathological image used in the embodiment of the present invention, where part (a) is the blurred image, part (b) is the corresponding real focused image, and part (c) is the clear image after deblurring. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following will further illustrate the concept, specific structure, and technical effects generated by the present invention in combination with the drawings and embodiments, so as to fully understand the purpose, features, and effects of the present invention.

[0032] Example 1:

[0033] The embodiment of the present invention proposes an efficient and general pathological image deblurring method based on the diffusion model, including the following processes:

[0034] Obtain a to-be-processed pathological image that needs to be deblurred; wherein, the to-be-processed pathological image has a defocus blur representation;

[0035] Input the to-be-processed pathological image into a pre-trained deblurring model to obtain a target pathological image that matches the to-be-processed pathological image and is output by the pre-trained deblurring model;

[0036] Wherein, the target pathological image has a clear focus representation; the pre-trained deblurring model is trained using sample pathological images containing defocus blur representations of different degrees; the defocus blur representations of different degrees refer to the defocus blur caused by different defocus distances; the pre-trained deblurring model includes a deblurring Transformer model (DTransformer model), a feature extraction network (FE), and a diffusion model.

[0037] See Figure 1 , in the first training stage, the feature extraction network and the DTransformer model participate in the training, and the DTransformer model learns the ability to restore the blurred pathological image according to the clear pathological image features. In the second training stage, the DTransformer model, the feature extraction network (FE), and the diffusion model jointly participate in the training to train the ability of the diffusion model to restore the blurred pathological image features. By inputting the blurred image feature D, the pure noise Z T is restored T times to obtain an output feature, which is sent to the DTransformer to obtain a predicted clear pathological image I HQ to further improve the ability of the DTransformer. In the inference and verification stage, the blurred pathological image is input into the feature extraction network (FE) and the diffusion model to obtain the restored image feature, and then input into the DTransformer model to deblur the blurred image and output the predicted clear pathological image. Specifically, the output of the feature extraction network is used as the input of the diffusion model and the DTransformer respectively. Then, the output of the diffusion model is input into the DTransformer and input together with the output of the feature extraction network to obtain the final predicted image result.

[0038] The DTransformer model is first proposed in the present invention and is composed of DTransformer Blocks to form a UNet architecture.

[0039] Encoding stage: The input feature is calculated step by step through several layers (preferably set to 3 layers in the embodiment) of DTransformer Blocks, and each stage is followed by a downsampling layer. A total of three downsampling operations are completed to gradually compress the spatial resolution and extract multi-level feature abstractions, and finally the feature is mapped to a high-dimensional semantic feature space.

[0040] Bottleneck layer processing: The encoded high-level features are subjected to global context modeling through the underlying DTransformer Block to achieve cross-scale feature fusion;

[0041] Decoding stage: The output of the bottleneck layer is gradually reconstructed through the corresponding number of layers (3 levels are set in the embodiment) of DTransformer Block. An upsampling layer is connected before each level, and a total of three upsampling operations are completed to gradually restore the spatial resolution and refine the feature representation, and finally a high-fidelity clear image I is output. HQ 。

[0042] The DTransformer Block is composed of a Cross Transposed Attention (CTA) module and a Neighbourhood Aggregation Feed-Forward Network (NAFN).

[0043] The Cross Transposed Attention module CTA is composed of the fusion of a cross-attention mechanism and a dedicated attention mechanism. The Neighbourhood Aggregation Feed-Forward Network (NAFN) includes a fully connected layer, a convolutional sub-unit, a GELU activation function layer, and a gating logic.

[0044] The preferred implementation solutions further proposed in the embodiment are as follows:

[0045] The Cross Transposed Attention module CTA calculates two different sources at the input end: features and the blurred image input First, perform a dot product accumulation operation through parallel fully connected layers. Then calculate the cross-attention to obtain the query Q, key K, and value V, where the feature Z is the real clear image I GT and the blurred pathological image I LQ After the Concat splicing operation, the features obtained by inputting into the feature extraction network FE, then perform a dot product on Q and K to generate a transposed attention map A of size C×C, and then perform a dot product operation with V and connect it to the fully connected layer. Here, C is the dimension of the feature, and the preferred value is recommended to be 256.

[0046] In the embodiment, the inputs of the CTA module are respectively the image feature and the latent image representation where R is the real number field and N is the length of the feature, and the preferred value is recommended to be 64. The matrices for constructing the query (Q), key (K), and value (V) are respectively

[0047] Q = XW Q 、K = ZW K 、V = ZW V

[0048] The cross-modal attention map (Transposed-Attention Map) is calculated and denoted as A

[0049]

[0050] where softmax() is the normalized exponential function, and d k is a learnable parameter matrix, and W Q , W K , W V are respectively the combination of the normalization layer Norm and the fully connected layer Linear.

[0051] Then, the weighted feature fusion of the attention map is output as X out

[0052] X out = AVW O + X in

[0053] where W O is the fully connected layer Linear. To reduce the computational complexity, CTA converts the traditional spatial dimension attention into channel dimension attention, thus reducing the computational amount from O(H 2 ) to O(C2), significantly improving the inference efficiency.

[0054] In the neighborhood aggregation feed-forward neural network (NAFN), at the input end, a convolutional layer with a 1×1 convolutional kernel is used, and then it is connected to two transposed convolutional layers with 3×3 convolutional kernels, where one is connected to the GELU activation function layer to generate gating weights.

[0055] The NAFN module is based on the traditional feed-forward neural network, introducing a convolutional structure and a gating mechanism to enhance the spatial expression ability. Its input is X in which is the output of the cross transposed attention module CTA. First, a dot product and accumulation operation are performed through parallel fully connected layers, then a linear transformation is obtained through a 1×1 convolution to get Y, and then through two 3×3 transposed convolutions. One branch is connected to the GELU activation function to obtain weights through the gating structure, and the final result is obtained through dot product and input into a 1×1 convolution to get the output feature X out′ . This process can be described as: first calculate the intermediate result U

[0056] U = W2·GELU(W1*Y)

[0057] Output is

[0058]

[0059] where W1 and W2 are 3×3 transposed convolutions, and W3 is a 1×1 convolution.

[0060] The feature extraction network is proposed by the present invention and includes a number of Residual Blocks, a patch Mix module, and a Channel Mix module.

[0061] The Residual Block is stacked by a convolution with a 3×3 convolutional kernel, a GELU activation function layer, and a convolution with a 3×3 convolutional kernel; when specifically implemented, it is preferably recommended to cascade 4 Residual Blocks.

[0062] The patch Mix module is stacked by a fully connected layer 1, a GELU activation function layer, and a fully connected layer 2. Among them, the input dimension of the fully connected layer 1 is the number of image patches, the output size is the number of patches × expansion degree, the input dimension of the fully connected layer 1 is the number of patches × expansion degree, and the output size is the number of image patches;

[0063] The Channel Mix module is stacked by a fully connected layer 3, a GELU activation function layer, and a fully connected layer 4. Among them, the input dimension of the fully connected layer 3 is the number of channels, the output size is the number of channels × expansion degree, the input dimension of the fully connected layer 4 is the number of channels × expansion degree, and the output size is the number of image channels.

[0064] When specifically implemented, prior calculations can also be performed by a PixelUnshuffle operation on image pixels, a convolution with a 3×3 convolutional kernel, and a GELU activation function layer before the Residual Block. A convolution with a 3×3 convolutional kernel and an Avg Pool layer are connected between the Residual Block and the patch Mix module. A fully connected layer 5 and a GELU activation function layer are connected after the Channel Mix module.

[0065] In the embodiment, the feature extraction network includes three parts: a Residual Block, a patch Mix module, and a Channel Mix module, which are used to compress the image from the original space to the latent space.

[0066] In the Residual Block, a 3×3 convolution, a GELU activation, and a 3×3 convolution are sequentially set, and a residual connection is made between the input and the output to construct a shallow feature encoding.

[0067] The patch Mix module realizes spatial block-level modeling through two fully connected layers and a non-linear activation, and the specific form is

[0068] PatchMix(X) = FC2(GELU(FC1(X)))

[0069] where X is the input, the input of the fully connected layer FC1 is the number of patches N, the output dimension r < N, and the output dimension of the fully connected layer FC2 is restored to N.

[0070] The channel mixing module adopts

[0071] ChannelMix(X) = FC4(GELU(FC3(X)))

[0072] to model the redundancy or dependency relationships between different channels, where FC3 and FC4 are fully connected layers.

[0073] Through the cascading of these three modules, the image is input from three channels and mapped to a latent representation where N << H×W, greatly reducing the computational burden of the subsequent diffusion model. Here, H and W are the length and width of the input image, and C' is the channel size.

[0074] The diffusion model consists of T Denoising U-Nets, where T is the diffusion step; the Denoising U-Net is a U-Net network structure for denoising proposed in the present invention, and each Denoising U-Net consists of 4 cross-attention modules; the input of the cross-attention module is the current step t and the image feature D, and the cross-attention modules are divided into groups of two. The first group performs convolution operations with a 3×3 convolution kernel on the output of each cross-attention module, and the second group performs deconvolution operations with a 3×3 convolution kernel on the output of each cross-attention module.

[0075] In the embodiment, the diffusion model consists of a T-step iterative denoising process, and the input of each step includes the image feature vector and the current time step t, and the output is the gradually restored clear image feature. In the forward diffusion process, Gaussian noise is gradually added to the clear image, and the intermediate generation result x t is

[0076]

[0077] where the true noise weight simulates the Markov process from the real sample to the noise, s is the subscript of the cumulative process, and α s is the weight corresponding to the subscript s, is the standard normal distribution. In the reverse denoising stage, the model needs to learn a noise predictor ε θ (x t , t) ≈ ε, so as to estimate the clear image is

[0078]

[0079] where, x t is the intermediate result at time step t. is the weight, ε θ (x t , t) is the predicted noise at the current time step and parameter θ.

[0080] This inverse diffusion process is reversible, ensuring that the generated images have reasonable structures, realistic contents, and significantly reduced artifacts.

[0081] The above defocusing model has high efficiency and generality; the above high efficiency means that the resources and time required by the above defocusing network during training and inference are less than those of traditional diffusion models of the same scale; the above generality means that when the above defocusing network processes pathological images of different tissues and cell types and pathological images with different degrees of defocus blur, it can ensure the quality of the final output target pathological images.

[0082] The high efficiency of the above defocusing model is due to the fact that the feature extraction network in the model performs dimensionality reduction on the original input image. The dimension of the image space is H×W×3, while the dimension of the compressed latent space is N×C′, where H and W are the length and width of the input image, and N and C′ represent the number of markers and the channel size for extracting features, and the number of markers N is a constant much smaller than H×W. Since the diffusion model part in this network model requires 1000 iterations with a default value during the training stage and the inference stage, it more significantly reduces the training and inference time consumed by the reduction of the image dimension, improving the efficiency of the model.

[0083] The generality of the above defocusing model is due to the fact that this network model does not rely on any unique features of the input image to ensure the reliability of the final result, but guarantees the generality and effectiveness for any type of input through strict mathematical and logical reasoning; in addition, the generality of the diffusion model is natural because their basic principles are based on probability theory and stochastic processes, that is, as long as the samples conform to the defined distribution, the diffusion model can be used for provable processing. At the same time, the process of the diffusion model is realized through a Markov chain, and this process is reversible, that is, from any state, it can return to the initial state through the reverse process. Therefore, by adjusting the parameters of the Markov chain, it can adapt to image inputs with different features, thus ensuring the generality.

[0084] Example 2:

[0085] Based on Example 1, a method for defocusing pathological images is further provided, including the following steps:

[0086] Step 1, create a training dataset;

[0087] Among them, the training data set includes a defocused blurred sample pathological image set and a focused clear sample pathological image set; the defocused blurred sample pathological image set includes multiple defocused blurred sample pathological images; the focused clear sample pathological image set includes multiple focused clear sample pathological images.

[0088] The creation of the training data set may include the following methods:

[0089] Step 1: Scan samples in different batches according to different defocus distances to obtain multiple initial sample pathological images;

[0090] Step 2: Cluster the multiple initial sample pathological images based on whether the defocus distance is equal to 0 to obtain the defocused blurred pathological image set and the focused clear pathological image set.

[0091] Step 2, Training Phase 1: Train the feature extraction network and the DTransformer.

[0092] The present invention further proposes the implementation method of Training Phase 1 as follows:

[0093] 1) Based on the focused clear sample pathological image set, input the multiple focused clear sample pathological images into the feature extraction network to be trained. The image feature Z of the focused clear sample pathological images is generated by the feature extraction network to be trained, and the image feature Z is input into the DTransformer to be trained; based on the corresponding defocused blurred sample pathological image set, input the multiple defocused blurred sample pathological images into the DTransformer to be trained, and the clear prediction image corresponding to each defocused blurred sample pathological image is output by the DTransformer to be trained;

[0094] 2) Use a preset loss function to obtain the function value of the preset loss function based on the clear sample pathological images and the clear prediction images;

[0095] 3) Based on the function value of the preset loss function, train the feature extraction network to be trained and the DTransformer to be trained until the preset training conditions are met.

[0096] Based on this, the embodiment further proposes that the specific implementation scheme of Step 2 of the preferred suggestion includes the following steps:

[0097] Step 21: Based on the set of focused clear sample pathological images, input the multiple focused clear sample pathological images into the feature extraction network to be trained. The feature extraction network to be trained generates the image feature Z of the focused clear sample pathological images, and input the image feature Z into the DTransformer to be trained. Based on the corresponding out-of-focus blurred sample pathological image set, input the multiple out-of-focus blurred sample pathological images into the DTransformer to be trained, and the DTransformer to be trained outputs the clear prediction image corresponding to each out-of-focus blurred sample pathological image.

[0098] Step 22: Use a preset loss function to obtain the function value of the preset loss function based on the clear sample pathological image and the clear prediction image.

[0099] Among them, the preset loss function is an image loss function; the image loss function obtains the function value of the preset loss function based on the clear sample pathological image and the clear prediction image. The specific loss function L trans is:

[0100]

[0101] Among them, I GT is the real clear sample pathological image, is the clear prediction image, and || ||1 is the absolute value loss function.

[0102] Step 23: Based on the function value of the preset loss function, train the feature extraction network to be trained and the DTransformer to be trained until the preset training conditions are met.

[0103] Among them, step 23 may include the following steps:

[0104] Step 231: Backpropagate the function value of the image loss function to the feature extraction network to be trained and the DTransformer.

[0105] Step 232: The feature extraction network to be trained and the DTransformer adjust the network parameters of each network layer according to the function value of the image loss function.

[0106] Step 233: Iteratively execute step 231 and step 232 until the preset training completion condition is met; among them, the preset training completion condition is that the number of iterative training times is greater than the preset iterative training times threshold.

[0107] Step 3: Training stage 2: Conduct joint training of the feature extraction network, the diffusion model and the DTransformer.

[0108] The present invention proposes to jointly train the to-be-trained diffusion model and the DTransformer after being trained in training stage 1 based on the function values of the first set of preset loss functions and the function values of the second set of preset loss functions until a preset training completion condition is met, including:

[0109] 1) Backpropagate the function value of the image feature loss function to the to-be-trained diffusion model;

[0110] 2) The to-be-trained diffusion model adjusts the network parameters of each network layer according to the function value of the image feature loss function;

[0111] 3) Backpropagate the function value of the image loss function to the DTransformer after being trained in training stage 1;

[0112] 4) The DTransformer after being trained in training stage 1 adjusts the network parameters of each network layer according to the function value of the image loss function;

[0113] 5) Iteratively execute steps 1)-4) until a preset training completion condition is met; wherein, the preset training completion condition is that the number of iterative training times is greater than a preset iterative training times threshold.

[0114] Based on this, the embodiment further proposes a specific implementation scheme for step 3 of the preferred suggestion, including the following steps:

[0115] Step 31: Input the multiple out-of-focus blurred sample pathological images into the feature extraction network trained in training stage 1 based on the out-of-focus blurred sample pathological image set, output the image features of the out-of-focus blurred sample pathological images, and input the image features into the to-be-trained diffusion model;

[0116] Step 32: The Denoising U-Net in the above to-be-trained diffusion model passes in the above image features, and at the same time inputs the current step size t (initially T) into the Denoising U-Net in the above to-be-trained diffusion model, and the Denoising U-Net outputs the image features corresponding to the step size t - 1;

[0117] Specifically in implementation, T can be specified by the user in advance, and the default value of the preferred suggestion is 1000.

[0118] Step 33: Repeat the above step 32 for T times to obtain the predicted image features of the clear focused clear sample pathological images corresponding to each out-of-focus blurred sample pathological image;

[0119] Among them, steps 32-33 involve the forward and backward processes in the diffusion model.

[0120] Forward process: Given a set of data \(x_0\sim q(x)\) sampled from the true data distribution, i.e., clear pathological image features, in \(T\) steps (note that \(T\) is a variable parameter during training), Gaussian noise is gradually added to the sample step by step. Finally, a series of samples \(x_1, x_2, \cdots, x\) after noise addition are obtained. T represent the intermediate results generated at each time step. Among them, the size of the number of steps \(T\) is restricted by the weight \(\beta\). t Constraint:

[0121]

[0122] where \(q()\) represents the probability distribution and \(N()\) represents the normal distribution, i.e., \(q(x)\) is the probability distribution of the sample \(x\), and \(q(x|x)\) is the probability distribution of the sample \(x\) under the condition \(x\), and \(I\) is the unit variance. t |x t-1 ) is the probability distribution of the sample \(x\) t under the condition \(x\) t-1 .

[0123] It can be derived from the total probability theorem that:

[0124]

[0125] where \(q(x|x_0)\) is the joint probability distribution of the samples \(x_1\) to \(x\). 1:T |x0) is the joint probability distribution of the samples \(x_1\) to \(x\). T

[0126] During this period, the original data \(x_0\) gradually loses its unique and distinct features during the \(t\) - step iteration of forward diffusion. Finally, when \(T\) approaches infinity, the final output \(x\) T is equivalent to an isotropic Gaussian - distributed noise, as Figure 2 shown.

[0127] Define the weight parameter \(\alpha\) t = 1 - \(\beta\) t , Given Gaussian noise: \(\epsilon\) t-1 , \(\epsilon\) t-2 , \(\cdots\sim N(0, I)\), then according to the parameter re - normalization technique, the final result derived is denoted as:

[0128]

[0129] where \(\epsilon\) is Gaussian noise that satisfies the standard normal distribution.

[0130] Reverse process: The diffusion process is to add noise to the data, and the reverse process is a denoising process. During the reverse diffusion process, the Gaussian noise \(x\) T \(\sim N(0, I)\) is used as the input, and from the probability distribution \(q(x|x)\) t-1|x t ) sample, infer and reconstruct the real sample. It should be noted here that if the parameter β t is small enough, the sampling result of the probability distribution q(x t-1 |x t ) is still a Gaussian distribution, and this value is also difficult to evaluate. Therefore, it is difficult to infer the original real distribution step by step through the formula-solving method. Here, the idea of the diffusion model is also very straightforward: the entire dataset already exists. Since it cannot be directly solved, we can try to train a model to predict the conditional probability of these noises. Since we want to make predictions, the label is to record the real noises generated in each step of the forward propagation as the label. The process of forward diffusion, in addition to inference, also includes a process similar to the "construction of the dataset" used in this mathematical model. When the model performs reverse diffusion, it can predict the Gaussian noises generated in the forward diffusion and infer step by step to restore the initial sample data.

[0131] By the total probability theorem, the posterior estimate distribution of removing the Gaussian noise added to the intermediate result x t-1 at the previous time step can be used to obtain the conditional probability p t-1 of x t under the condition x θ (x t-1 |x t ):

[0132] p θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , t), ∑ θ (x t , t))

[0133] where θ represents the given parameter, μ θ (x t , t) is the predicted mean of the sample x t , and ∑ θ (x t , t) represents the predicted variance of the sample x t .

[0134] Although the probability distribution q t-1 of x t under the condition x θ (x t-1 |x t ) cannot be directly calculated, but adding the posterior distribution q(x t-1 |x t, x0) can be obtained through computational processing. It should be noted that the conditional probability of the inverse diffusion process can be calculated from the forward diffusion process. Based on the property of the Markov chain, given the initial sampling, its probability distribution is:

[0135]

[0136] where I is the unit variance of the normal distribution.

[0137] The posterior distribution q(x t-1 |x t , x0) can be obtained, and its variance and mean are as follows:

[0138]

[0139] where α t = 1 - β t , β t is the variance of the distribution of the sample x t at time step t, and the parameter meaning corresponding to t - 1 is the same.

[0140] Here, it can be seen that the variance is a constant value, and the change of the mean is regulated by the sample x t at time step t.

[0141] In summary, what step 33 needs to do is to use the image feature D as the conditional vector of the conditional diffusion model to estimate the degradation prior from a Gaussian noise through the reverse process by gradually denoising

[0142] Step 34: Using the first set of preset loss functions, based on the predicted image features, the image features obtained from the clear focused sample pathological images corresponding to the out-of-focus blurred sample pathological images through the above-mentioned feature extraction network, obtain the function value of the first set of preset loss functions.

[0143] The first set of preset loss functions is the image feature loss function; the image feature loss function is based on the predicted image features, the image features obtained from the clear focused sample pathological images corresponding to the out-of-focus blurred sample pathological images through the above-mentioned feature extraction network, and obtain the function value of the image feature loss function :

[0144]

[0145] where C' is the number of channels of the feature, Z(i) is the true feature, is the predicted feature.

[0146] Step 35: Input the predicted image features and the out-of-focus blurred sample pathological images into the DTransformer after being trained in the first training stage, and the trained DTransformer outputs clear predicted images corresponding to each out-of-focus blurred sample pathological image;

[0147] Step 36: Use the second set of preset loss functions to obtain the function values of the second set of preset loss functions based on the clear sample pathological images and the clear predicted images corresponding to the out-of-focus blurred sample pathological images;

[0148] The second set of preset loss functions is an image loss function; the image loss function is based on the clear sample pathological images and the clear predicted images corresponding to the out-of-focus blurred sample pathological images to obtain the function value of the image loss function L trans :

[0149]

[0150] where, I GT is the true clear sample pathological image, is the clear predicted image, and || ||1 is the absolute value loss function.

[0151] Step 37: Based on the function values of the first set of preset loss functions and the function values of the second set of preset loss functions, obtain the overall loss function of the second training stage. Train the diffusion model to be trained and the DTransformer after being trained in the first training stage until the preset training completion condition is met, and obtain the pre-trained deblurring model from the diffusion model to be trained and the DTransformer after being trained in the first training stage.

[0152] Step 37 may include the following steps:

[0153] Step 371: Feed back the function value of the image loss function to the feature extraction network to be trained and the DTransformer;

[0154] Step 372: The feature extraction network to be trained and the DTransformer adjust the network parameters of each network layer according to the function value of the image loss function;

[0155] Step 373: Iteratively execute Step 371 and Step 372 until the preset training completion condition is met; where, the preset training completion condition is that the number of iterative training times is greater than the preset iterative training times threshold.

[0156] where, the overall loss function is:

[0157]

[0158] As described above, by using the efficient and general pathological image deblurring method provided in the embodiments of the present disclosure, on the one hand, by improving the traditional Transformer, a cross transposed attention module (CTA) and a neighborhood aggregation feed-forward neural network (NAFN) are proposed, enabling the model to learn from both the input image and image features simultaneously and enhancing the model's ability to learn image neighborhood features; on the other hand, through a regression-based image feature extraction model, the defect of low distortion accuracy of the generation model is improved, and the relationship between patches and channels is fully considered. Generally speaking, this method solves the problems of high computational cost, insufficient generalization ability, and poor image generation quality in previous deblurring methods by combining a diffusion model with an improved Transformer. See Figure 3 , the clear pathological images generated by this method effectively restore the blurred and missing detailed textures, and effectively suppress the appearance of artifacts and non-existent textures. The gap with real clear pathological images is extremely small, which is convenient for diagnosis.

[0159] For the convenience of understanding the technical effects of the present invention, the following table shows the experimental comparison between the present invention and the prior art:

[0160] Comparison method Peak signal-to-noise ratio Structural similarity index DBGAN-v2 32.49 0.9212 ASNet 31.86 0.8963 DRBNet 33.68 0.9344 MedDeblur 34.44 0.9433 LoFormer 33.99 0.9408 Restormer 35.12 0.9498 DiffIR 36.27 0.9621 This method 36.89 0.9678

[0161] It can be seen from the above experimental results that this method shows excellent performance in both peak signal-to-noise ratio and structural similarity index evaluation metrics. Specifically, the peak signal-to-noise ratio of this method is 36.89, significantly higher than other comparison methods. The peak signal-to-noise ratio of DiffIR is 36.27, and that of Restormer is 35.12. The peak signal-to-noise ratio values of other methods do not exceed 35, which indicates that while this method restores the details of pathological images and reduces noise, it retains more image information, making the deblurred images clearer and of higher quality. In terms of the structural similarity index metric, the structural similarity index of this method is 0.9678, leading all comparison methods. The structural similarity index of DiffIR is 0.9621, and that of Restormer is 0.9498, indicating that this method not only surpasses other methods in terms of image quality but also performs well in terms of structure preservation and detail fidelity. By comparing with other methods, this method has achieved the best results in two core metrics (peak signal-to-noise ratio and structural similarity index), which means that this method can more effectively remove the blur in pathological images and restore more delicate image details, especially in the field of pathological images where the requirements for details are higher. Therefore, the efficient and general pathological image deblurring method based on the diffusion model, with its high peak signal-to-noise ratio and structural similarity index performance, proves its application potential in pathological image deblurring and can provide more reliable image quality support for the subsequent analysis and diagnosis of pathological images.

[0162] In specific implementation, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including the operation of the corresponding computer program, should also be within the protection scope of the present invention.

[0163] Next, the efficient and general pathological image deblurring electronic device based on the diffusion model provided by the present invention will be described. The efficient and general pathological image deblurring electronic device described below can be correspondingly referred to the efficient and general pathological image deblurring method described above.

[0164] The electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communications interface, and the memory complete communication with each other through the communication bus. The processor can call the logical instructions in the memory to execute the efficient and general pathological image deblurring method based on the diffusion model, mainly including the software processing part in the above steps.

[0165] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the essence of the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0166] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the software processing part in the efficient and general pathological image deblurring method provided by the above-mentioned various methods.

[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the software processing part of the efficient general pathological image deblurring method based on the diffusion model provided by the above-mentioned various methods.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An efficient and universal pathological image deblurring method based on a diffusion model, characterized in that: The following steps are involved: Acquire an out-of-focus and blurred pathological image to be processed; Inputting the image into a pre-trained deblurring model to output a clearly focused pathological image; Among them, the pre-trained deblurring model includes a feature extraction network, a DTransformer model and a diffusion model; the feature extraction network is used to compress the original image into a low-dimensional latent space; the DTransformer model contains a cross-transposition attention module CTA and a neighborhood aggregation feedforward neural network NAFN, which are used to enhance feature interaction; the diffusion model is used to restore clear image features through an iterative denoising process.

2. The method according to claim 1, characterized in that: The processing process of the cross-transposition attention module CTA is to input image features and potential representations to generate a query matrix, a key matrix and a value matrix; The cross-modal attention map is calculated and weightedly fused with the value matrix to obtain the output features.

3. The method according to claim 1, characterized in that: The neighborhood aggregation feedforward neural network NAFN introduces a convolutional structure and a gating mechanism to enhance spatial expression capabilities.

4. The method according to claim 1, characterized in that: The feature extraction network includes a residual module, a patch mixing module and a channel mixing module. The patch mixing module models the spatial dimension of the image block through a fully connected layer, and the channel mixing module models the dependency relationship between channels through a fully connected layer.

5. The method according to claim 1, characterized in that: The training of the pre-trained deblurring model includes the following stages: in the first stage, jointly training the feature extraction network and the DTransformer model; In the second stage, the feature extraction network, diffusion model and DTransformer are jointly trained.

6. The method according to claim 1, characterized in that: The diffusion model is composed of T Denoising U-Nets, where T is the diffusion step length; the Denoising U-Net is a U-Net network structure used for denoising, and each Denoising U-Net is composed of 4 cross-attention modules.

7. The method according to claim 1, characterized in that: The inverse denoising process of the diffusion model estimates the clear image features through the noise predictor.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the efficient and universal pathological image deblurring method based on the diffusion model as described in any one of claims 1 to 7 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the efficient and universal pathological image deblurring method based on a diffusion model as claimed in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the efficient and universal pathological image deblurring method based on a diffusion model as claimed in any one of claims 1 to 7 is implemented.