Medical Image Generation Based on Label Evolution and Class-Mask Self-Attention Diffusion Model

By introducing the label evolution module and the class mask attention module in the medical image generation model, combined with the diffusion model, the shortcomings of the existing models in terms of generation diversity and accuracy are solved, and more efficient medical image generation is achieved.

CN118537478BActive Publication Date: 2025-06-10TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410364380.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-06-10
Estimated Expiration
2044-03-28

AI Technical Summary

Technical Problem

The existing medical image generation models have shortcomings in generating diversity and accuracy, especially when the Unet network processes medical images, the types of images generated are not clear enough and fail to fully capture the feature differences between different categories.

Method used

The tag evolution module and the class mask attention module are introduced, combined with the diffusion model, and guide the generation of diverse cell images by gradually updating point labels and fusion cell type information.

Benefits of technology

Effectively generating diverse medical images improves the accuracy and diversity of model generation and solves the shortcomings of Unet network in medical image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118537478B_ABST
    Figure CN118537478B_ABST
Patent Text Reader

Abstract

Medical image generation based on a diffusion model with label evolution and class mask self-attention belongs to the fields of computer vision and medical image generation. Based on the diffusion model as the basic framework, we perform data preprocessing on the obtained Lizard dataset and propose targeted improvements to the problems existing in the field of medical image generation. We replace the self-attention module in the diffusion model with the designed class mask self-attention module, which can capture the information of cell types. In addition, in order to generate diverse medical images, we introduce a label evolution module, which is incorporated into Unet in the form of conditional normalization by gradually evolving point labels into their complete Mask labels, enabling the network to reasonably generate diverse medical images. The present invention is applicable to all medical image datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical image generation based on a diffusion model with label evolution and class mask self-attention, belonging to the fields of computer vision and medical image generation. Background Art

[0002] Medical image generation is a research field that focuses on the generation, processing, and analysis of medical image data. With the continuous progress of medical image generation technology, the scale of medical image data has gradually increased, providing extensive development opportunities for the field of medical image generation. The goal of medical image generation is to generate virtual medical image data to assist doctors in making more accurate diagnoses and treatments. Early medical image generation methods were mainly based on manually designed rules and traditional graphics techniques, including physical model-based rendering and simple texture mapping. These methods were limited by computing resources and the model's expressive power. With the rise of deep learning, neural networks have shown great potential in the field of image generation. Early convolutional neural networks (CNNs) were used for tasks such as image classification, laying the foundation for image generation. The introduction of variational autoencoders enabled image generation to learn the distribution in the latent space, thus generating more diverse images. VAEs combine probabilistic graphical models and neural networks. Generative adversarial networks were proposed by Ian Goodfellow et al., triggering a revolution in the field of image generation. GANs can generate high-quality, realistic images through the confrontation between the generator and the discriminator in the game process.

[0003] Although generative adversarial networks (GANs) have achieved remarkable success in image generation and other tasks, they also have some drawbacks: The training process of GANs can be unstable. The game between the generator and the discriminator may lead to model oscillation or collapse, and sometimes the generator cannot generate high-quality samples. GANs may sometimes fall into mode collapse, with limited diversity in the generated samples, resulting in overly similar or uncreative generated images.

[0004] To address these deficiencies, we introduced a label evolution module and a class mask mechanism into the original diffusion model. First is the label evolution module. The main body of this module is a DNANet network structure, which is used to extract the information of the Mask labels of cells to generate corresponding mapping graphs. Then, through label evolution, the original point Labels of cells are updated and used to replace the initial point Labels, so as to gradually update the point Labels. The new point Labels learned at each stage are used as conditions to guide the diffusion model to generate diverse medical images. Secondly, since the guidance of the diffusion model by the individual point label evolution has infinitely many solutions, we added a class mask attention module to constrain the cell types corresponding to each point label. Through the above two modules, it is possible to ensure that the diffusion model generates diverse cells of multiple types. Summary of the Invention

[0005] The object of the present invention is to provide a medical image generation network based on a diffusion model with label evolution and class mask attention, so as to generate diverse medical images and solve the problems existing in the field of medical image generation.

[0006] To achieve the above object, the solution of the present invention is: based on the medical image generation network of the diffusion model as the basic framework, a more advanced network model is designed, which can make the generated pictures more diverse and accurate on the premise of improving the model performance. The current specific steps are as follows:

[0007] (1) Obtain a total of 384 lizard datasets commonly used in medical image generation, and divide them into a training set, a test set, and a validation set according to a ratio of 7:2:1 and randomly divide them five times respectively;

[0008] (2) Perform data preprocessing. For each picture in the lizard dataset, perform a cropping operation and split it into 10 small pictures;

[0009] (3) Select a suitable point label evolution network model, which can update the point Labels in the lizard dataset, and pre-train a model that can achieve ideal results.

[0010] (4) Design a medical image generation network based on a diffusion model with label evolution and class mask attention. The updated labels output by the point label evolution module each time will be input into the Decoder part of the Unet, so that the Unet network can generate diverse pictures according to the gradually updated label information. At the same time, to ensure the categories of the generated cells, we propose a class mask self-attention module to fuse the information of cell classes. The combination of the two modules is used to guide the generation of diverse cell images;

[0011] (5) Conduct multiple experiments to explore the optimal network parameters and evaluate the model using FID and IS evaluation metrics;

[0012] (6) Design ablation experiments to explore the effectiveness of each proposed module and the specific problems it solves;

[0013] (7) Prove the feasibility and superiority of the method, that is, the method can achieve the generation of medical images.

[0014] The beneficial effects of the present invention are as follows: This method can effectively generate diverse medical images. Based on the diffusion model, common datasets in medical images are selected and their formats and sizes are adjusted. Further, some challenging problems are solved: The Unet network in the diffusion model may have some deficiencies when processing medical images. One of them is that the boundaries between the types of images generated by the generation network are not clearly defined. This means that the generated images may not be visually clear enough to distinguish the subtle differences between different categories. Another problem is the insufficient extraction of type information in medical images by the Unet network. In other words, the Unet network may not fully capture the feature differences between different categories in medical images, resulting in the generated images lacking corresponding type information, thus affecting the accuracy and diversity of model generation. Brief Description of the Drawings

[0015] Figure 1 is the overall algorithm flowchart of the present invention.

[0016] Figure 2 is the specific implementation method diagram of the label evolution module of the present invention.

[0017] Figure 3 is the network diagram of the encoded class mask attention mechanism module of the present invention. Detailed Embodiment

[0018] Obtain a dataset of lizard in JPG format, and divide it into a training set, a test set, and a validation set according to 7:2:1 and randomly divide them five times respectively. Then, perform data preprocessing on each picture in the lizard dataset. This includes cropping each picture, splitting it into 10 small pictures, converting each image into an npy format image, and then searching for an advanced generation network model. After putting the preprocessed dataset into the network for training, it is found that there are problems such as blurred class division and poor quality in the generated images compared with the actual images.

[0019] To address the above problems, we constructed a medical image generation algorithm based on a diffusion model with label evolution. In our neural network, we first proposed adding a label evolution module to the diffusion model. The specific implementation method is as followsFigure 2 As shown, where we will send the Mask Label corresponding to the Point Label into a feature map extraction network (here we use the Unet network to extract features) to obtain the feature map corresponding to the Mask Label. According to the centroid in the Point Label, local neighborhood candidate pixels are extracted through an adaptive threshold. As shown in the formula

[0020]

[0021] where represents candidate pixels, ⊙ represents element-wise multiplication, and T adapt represents whether the adaptive threshold is related to the current prediction . The positive pixels in the label are arranged according to

[0022]

[0023] where h and w are the height and width of the input image, and r is set to 0.15%. T b is the minimum threshold, here set to 0.5, and k is the control threshold growth rate. As increases, the threshold also increases, which can reduce the error accumulation of low-contrast targets and strong background clutter.

[0024] The updated label obtained each time will replace the original Point Label for subsequent label updates. Secondly, it is input into the diffusion model. The Decoder in Unet is used to guide the generation of images.

[0025] For the class mask attention module as Figure 3 shown. We first input the mask information of the class into the shared multi-layer perceptron module (shared_conv). This shared multi-layer perceptron consists of a convolutional layer and a ReLU activation function. The role of the convolutional layer is to map the input class mask conditional information from the cell type dimension (label_nc) to the hidden layer dimension (nhidden). Here, label_nc represents the dimension of the cell type, and nhidden represents the dimension of the hidden layer. Next, the output hidden representation will be passed to the Gconv and Bconv modules. These modules map the hidden representation to the dimension of the output space (norm_nc). Specifically, Gconv and Bconv respectively generate the gamma and beta parameters, which will be used for subsequent conditional normalization operations. Therefore, through this process, we can effectively utilize the input class mask information to generate the gamma and beta parameters, thus providing the necessary parameters for subsequent conditional normalization operations. Then, the output X is obtained through the formula out :

[0026] X out = X in *(1 + gamma) + beta (3)

[0027] X in represents the extracted feature map.

[0028] Next, we separately input the hidden representations obtained from the shared multi-layer perceptron module into the linear layer to obtain the query (Q), key (K), and value (V) vectors for subsequent self-attention mechanism calculations. Finally, we compare the images generated by the diffusion model with the actual images. Among them, the network uses the Mean Squared Error (MSE) as the loss function for network training.

[0029] To fairly evaluate the generation effect of the model, we will use two metrics: FID (Frechet Inception Distance) and IS (Inception Score). FID is used to evaluate the similarity between the generated images and the real images, and it calculates the Frechet distance between the feature distributions of the generated images and the real images. While IS is used to evaluate the diversity and quality of the generated images, and it measures the diversity and authenticity by calculating the entropy of the class distribution of the generated images and the KL divergence of the conditional distribution.

[0030] FID = ||μ r - μ g || 2 + Tr(∑ r + ∑ g - 2(∑ r ∑ g ) 1 / 2 ) (4)

[0031] IS g = exp(E x~p D KL (p(y|x)||p(y))) (5)

[0032] The smaller the FID, the more similar the generated images are to the real images, and the larger the IS, the higher the diversity and authenticity of the generated images. Therefore, for the evaluation of the model, we hope that the FID is as small as possible and the IS is as large as possible to ensure that the generated images have both high similarity and rich diversity.

[0033] Based on the existing medical image generation models, we conducted experiments and found some problems. In response to these problems, we proposed corresponding improvement solutions and verified them through experiments. The results show that the improvement modules we proposed can more effectively enhance the performance of the network in generating medical images. It is worth noting that these improvement modules can not only be directly applied to the current medical image generation network, but also can be easily embedded into other medical image generation networks to achieve a plug-and-play effect. These improvements not only help doctors in disease treatment, but also provide useful support for research in the medical field.

[0034] It should be noted that the above description is only an embodiment of the present invention, which is only to explain the present invention and does not limit the scope of the present invention. Modifications that are merely obvious within the technical concept of the present invention are also within the protection scope of the present invention.

Claims

1. A medical image generation algorithm based on a diffusion model of label evolution and class mask self-attention, characterized in that: The steps include: (1) Obtain a total of 384 Lizard datasets commonly used in medical image generation, and divide them into training set, test set and validation set according to the ratio of 7:2:1 and randomly divide them five times; (2) Perform data preprocessing and crop each image in the lizard dataset to split it into 10 small images; (3) Select a point label evolution network model, update and diffuse the point labels in the lizard dataset, and pre-train a model that can achieve ideal results; (4) Design a medical image generation network based on a diffusion model of label evolution and class-complementary self-attention. The updated label output by each point label evolution module is input into the Decoder part of the Unet, which allows the Unet network to generate diverse images based on the gradually updated label information. At the same time, the class-masked self-attention module is used to fuse the information of cell classes. The combination of the two modules is used to guide the generation of diverse cell images. (5) Conduct multiple experiments to explore the optimal network parameters and use FID and IS evaluation indicators to evaluate the model; (6) Design ablation experiments to explore the effectiveness of the point label evolution module and the class mask self-attention module and the specific problems they solve; According to step (4), the Mask Label corresponding to the Point Label is sent to the Unet feature map extraction network to extract features, and the feature map corresponding to the Mask Label is obtained. According to the centroid in the Point Label, the local neighborhood candidate pixels are extracted through an adaptive threshold. Each updated label obtained will replace the initial Point Label for subsequent label updates, and then input into the Decoder in the Unet in the diffusion model to guide the generation of images. The class mask information is input into the shared multi-layer perceptron module shared_conv, which consists of a convolutional layer and a ReLU activation function; the function of the convolutional layer is to map the input class mask conditional information from the cell type dimension 1abel_nc to the hidden layer dimension nhidden, where label_nc represents the dimension of the cell type and nhidden represents the dimension of the hidden layer; next, the output hidden representation will be passed to the Gconv and Bconv modules to map the hidden representation to the dimension norm_nc of the output space; specifically, the Gconv and Bconv modules generate gamma and beta parameters, respectively, for subsequent conditional normalization operations.

2. The medical image generation algorithm based on the diffusion model of label evolution and class mask self-attention as claimed in claim 1, characterized in that: According to the Lizard dataset commonly used in the medical image field in step (1), five cross-validations were performed and the average value was calculated.

Citation Information

Patent Citations

  • Small sample image classification method based on feature adaptation

    CN114898136A

  • Breast medical image segmentation method based on improved U-net network

    CN116580202A