Medical image generation method and device based on migration model
Through the migration model-based method, the text-to-image generation model based on the diffusion model is constructed and optimized, which solves the problems of data finiteness and distribution inconsistency in medical image generation, and achieves high-quality medical image generation.
Patent Information
- Application Number
- CN202510470596.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-27
AI Technical Summary
In the field of medical imaging, traditional data augmentation methods cannot meet the complex medical image generation needs, mainly due to the limitations of medical data and inconsistencies in data distribution.
Using a medical image generation method based on migration model, the initial text-to-image generation model based on diffusion model is constructed by acquiring the medical image report data set for preprocessing, and the model is migrated and adjusted and fine-tuned, and the generative model is optimized to adapt to pathological text data.
It has achieved high-quality generation from pathological text to medical images, overcome the problems of limited medical data and inconsistency in data distribution, and improved the effectiveness and application value of medical image generation.
Smart Images

Figure CN120048448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a medical image generation method, device, computer equipment and storage medium based on a migration model. Background Art
[0002] In recent years, generative AI technology has made significant progress in many fields, especially in the field of image generation. Since OpenAI proposed the DALL-E model in 2021, the development of generative AI technology has been changing with each passing day. In 2022, the StableDiffusion model based on the Latent Diffusion Model (LDM) came out, marking an important breakthrough in text-to-image generation technology. Compared with traditional generative adversarial networks (GANs), the Stable Diffusion model can not only generate higher quality and more diverse images, but also provides more fine-grained control during the reasoning process, making it a dominant player in the field of image synthesis.
[0003] However, in the field of medical imaging, although generative models have shown broad application prospects, they also face a series of unique challenges. First, the limited medical data is a major problem. Since the acquisition of medical images usually requires expensive equipment and professional operations, and due to the requirements of patient privacy protection, the sharing and use of medical data are strictly restricted. Secondly, the inconsistency of data distribution between different hospitals is also a major obstacle in medical research and clinical applications. These problems make it difficult to obtain medical images for certain specific diseases and the amount of training samples is limited, which in turn affects the quality and application effect of the generative model.
[0004] In order to solve these problems, data augmentation technology has become an important method to improve the quality of medical image generation. Traditional data augmentation methods include transformation techniques such as rotation, flipping, and scaling. These methods help improve model performance to a certain extent, but they are still insufficient in the face of complex medical image generation requirements. Based on the above background, developing a new method to address these technical challenges in medical image generation has become an urgent task. Summary of the invention
[0005] The main purpose of the present invention is to provide a medical image generation method, device, computer equipment and storage medium based on a migration model to solve the problem that traditional data enhancement methods cannot meet the needs when facing complex medical image generation in the field of medical imaging.
[0006] To achieve the above-mentioned objectives, the present invention provides a medical image generation method based on a migration model, comprising: obtaining a medical image report data set, and preprocessing the medical image data set to obtain an initial medical image and initial pathology text data corresponding to the initial medical image; constructing an initial text-to-image generation model based on a diffusion model; using the initial medical image and the initial pathology text data corresponding to the initial medical image, performing migration adjustment on the initial text-to-image generation model based on the diffusion model to obtain an optimized medical text-to-medical image generation model based on the diffusion model, wherein the optimized medical text-to-medical image generation model based on the diffusion model is used to generate corresponding medical images according to the pathology text data; obtaining multiple target pathology text data, and using the optimized medical text-to-medical image generation model based on the diffusion model to analyze the target pathology text data to generate multiple target medical images.
[0007] Furthermore, the initial medical image and the initial pathological text data corresponding to the initial medical image are used to fine-tune the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, including: using the initial pathological text data corresponding to the initial medical image, adjusting the CLIP module of the initial diffusion model-based text-to-image generation model, the adjusted CLIP module is used to organize and output the target pathological text data; using the initial medical image and the initial pathological text data corresponding to the initial medical image, adjusting the VAE module of the initial diffusion model-based text-to-image generation model, the adjusted VAE module is used to output a basic medical image for describing the target pathological text data based on the target pathological text data output by CLIP; using the initial medical image, adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, the adjusted U-Net module is used to perform noise prediction processing and iterative denoising processing on the basic medical image output by the VAE module, and then output the target medical image; obtaining the optimized diffusion model-based medical text-to-medical image generation model.
[0008] Furthermore, after fine-tuning the initial diffusion model-based text-to-image generation model using the initial medical image and the initial pathological text data corresponding to the initial medical image, the medical image generation method also includes: performing data enhancement processing on the initial medical image to obtain an X-ray enhanced image, wherein the data enhancement processing includes: slight rotation, flipping and cropping; and performing secondary training on the fine-tuned diffusion model-based text-to-image generation model using the X-ray enhanced image to obtain an optimized diffusion model-based medical text-to-medical image generation model.
[0009] Furthermore, before adjusting the initial U-Net module of the diffusion model-based text-to-image generation model, the medical image generation method also includes: defining the sampling process of the fine-tuned U-Net module of the diffusion model-based text-to-image generation model as a non-Markov chain, allowing the U-Net module to sample the basic medical image based on the non-Markov chain, and completing iterative denoising based on the sampling results.
[0010] Furthermore, before adjusting the initial U-Net module of the diffusion model-based text-to-image generation model, the medical image generation method also includes: introducing negative prompts into the fine-tuned U-Net module of the diffusion model-based text-to-image generation model to enhance the image generation constraints of the U-Net module when performing noise prediction processing and iterative denoising processing on the basic medical image.
[0011] The present invention also provides an image classification method based on medical images, comprising: acquiring multiple target medical images, wherein the target medical images are generated by the above-mentioned medical image generation method, and the multiple target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; training a medical image training classification model based on the multiple target medical images to obtain a target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical images, and add pathological result labels to the target medical images according to the analysis results, wherein the pathological result labels include: negative labels and positive labels.
[0012] The present invention also proposes a medical image generation device based on a migration model, comprising: a processing unit, used to obtain a medical image report data set, and pre-process the medical image data set to obtain an initial medical image and initial pathological text data corresponding to the initial medical image; a construction unit, used to construct an initial diffusion model-based text-to-image generation model; a migration unit, used to use the initial medical image and the initial pathological text data corresponding to the initial medical image to perform migration adjustment on the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, wherein the optimized diffusion model-based medical text-to-medical image generation model is used to generate corresponding medical images according to the pathological text data; an analysis unit, used to obtain multiple target pathological text data, and use the optimized diffusion model-based medical text-to-medical image generation model to analyze the target pathological text data to generate multiple target medical images.
[0013] The present invention also provides an image classification device based on medical images, comprising: an acquisition unit, used to acquire multiple target medical images, wherein the target medical images are generated by the medical image generation method described in any one of claims 1 to 6 above, and the multiple target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; a training unit, used to train a medical image training classification model based on the multiple target medical images to obtain a target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical images, and add pathological result labels to the target medical images according to the analysis results, wherein the pathological result labels include: negative labels and positive labels.
[0014] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0016] The present invention provides a method, device, computer equipment and storage medium for generating medical images based on a migration model. By acquiring and preprocessing a medical image report data set, an initial medical image and corresponding initial pathology text data are obtained. This step ensures the quality and consistency of the data and provides a reliable basis for subsequent model training. An initial text-to-image generation model based on a diffusion model is constructed, and the preliminary conversion capability from text to image is achieved using advanced generation technology. The generation model is fine-tuned using the initial medical image and the corresponding initial pathology text data. The optimized model can more accurately generate corresponding medical images based on the pathology text, thereby improving the adaptability and generation quality of the model. Multiple target pathology text data are obtained, and the optimized model is used for analysis to generate multiple target medical images. This step realizes the effectiveness of the model in practical applications.
[0017] The entire technical solution combines transfer learning and diffusion models to achieve high-quality generation from pathology text to medical images, overcome the problems of limited medical data and inconsistent data distribution, and improve the effect and application value of medical image generation. This method not only improves the generation ability of the model, but also provides important technical support for clinical research and medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of steps of a medical image generation method based on a migration model in one embodiment of the present invention; Figure 2 A comparison of medical images generated by various image generation methods in one embodiment of the present invention Figure 1 ; Figure 3 A comparison of medical images generated by various image generation methods in one embodiment of the present invention Figure 2 ; Figure 4 Schematic diagram of the classification results of data sets generated by different methods in one embodiment of the present invention Figure 1 ; Figure 5 Schematic diagram of the classification results of data sets generated by different methods in one embodiment of the present invention Figure 2 ; Figure 6 is a schematic diagram of steps of an image classification method based on medical images in one embodiment of the present invention; Figure 7 is a structural block diagram of a medical image generation device based on a migration model in one embodiment of the present invention; Figure 8 is a structural block diagram of an image classification device based on medical images in one embodiment of the present invention; Fig. 9It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0019] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0021] Reference Figure 1 , this embodiment provides a medical image generation method based on a migration model, comprising: S1, obtaining a medical image report data set, and preprocessing the medical image data set to obtain an initial medical image and initial pathology text data corresponding to the initial medical image.
[0022] In an optional example, the process of acquiring and preprocessing a medical image report dataset may extract and organize relevant data from multiple medical datasets that have been published and widely used in research.
[0023] For example, we extract data from Stanford University’s CheXpert dataset, which contains 224,316 chest X-ray images and related radiology reports of 65,240 patients who underwent radiology examinations at Stanford Medical Center between October 2002 and July 2017. These data are labeled with the presence of 14 common chest X-ray examinations, and we convert these labels into concise natural language descriptions.
[0024] Also, we obtained the MIMIC Chest X-ray Database (MIMIC-CXR) from Beth Israel Deaconess Medical Center. This database includes 227,835 imaging studies of 65,379 patients in the emergency department between 2011 and 2016. Each study can include one or more images, usually containing frontal and lateral views, totaling 377,110 images. These images are accompanied by semistructured free-text radiology reports written by practicing radiologists, and we applied natural language processing techniques to extract content describing disease manifestations from these reports.
[0025] In addition, the lung adenocarcinoma STAS dataset is sourced from Shanghai Zhongshan Hospital and Shanghai Cancer Center of Fudan University. The dataset includes CT scans and corresponding formatted text. There are 203 cases (106 positive and 97 negative) from Zhongshan Hospital of Fudan University, and 103 cases (39 positive and 64 negative) from Shanghai Cancer Center of Fudan University. We extract the slice with the largest lesion area from the CT dataset and crop it around the center of the lesion to obtain an image of size 64x64. At the same time, we process the formatted labels corresponding to each CT scan into a concise text form and divide the dataset into a training set and a test set in a ratio of 7:3.
[0026] Through these steps, the preprocessed initial medical images and their corresponding initial pathological text data were obtained, laying the foundation for the subsequent construction of a text-to-image generation model based on a diffusion model.
[0027] S2, build an initial diffusion model-based text-to-image generation model.
[0028] It should be noted that the text-to-image generation model based on the diffusion model can be a Stable Diffusion model. The Stable Diffusion model is a text-to-image model that allows realistic images to be generated from text prompts. The model uses the diffusion model to generate realistic images through a denoising process. The diffusion model generates realistic images by removing the noise added to the real image.
[0029] Stable Diffusion consists of three main modules: encoder (VAE), UNet, and text encoder (CLIP). The VAE module speeds up the calculation and improves the image quality by encoding and decoding the image into a smaller latent space. The UNet module plays a key role in noise prediction and iterative denoising, predicting the noise pattern by iteratively removing noise from the image. The CLIP module is used to convert the user's input prompt text into text embedding.
[0030] S3, using the initial medical image and the initial pathological text data corresponding to the initial medical image, fine-tuning the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, wherein the optimized diffusion model-based medical text-to-medical image generation model is used to generate corresponding medical images according to the pathological text data.
[0031] The core idea of the fine-tuning process is to transfer knowledge learned from one domain to another through transfer learning. Transfer learning identifies similarities between existing and new knowledge to ensure that the performance of the downstream task model remains consistent, even when processing images from different sources. By fine-tuning the U-Net component of the Stable Diffusion model, high-fidelity chest X-ray (CXR) images can be generated. The generative model has the ability to learn radiological concepts and can be used to insert realistic anomalies.
[0032] At this point, the constructed and optimized medical text to medical image generation model based on the diffusion model can generate corresponding medical images according to the pathological text data, providing strong support for medical image generation and analysis.
[0033] S4, acquiring a plurality of target pathology text data, and using the optimized diffusion model-based medical text to medical image generation model to analyze the target pathology text data to generate a plurality of target medical images.
[0034] In summary, this embodiment provides a medical image generation method based on a migration model. By acquiring a medical image report data set and performing preprocessing, an initial medical image and corresponding initial pathology text data are obtained. This step ensures the quality and consistency of the data and provides a reliable foundation for subsequent model training. An initial text-to-image generation model based on a diffusion model is constructed, and the preliminary conversion capability from text to image is achieved using advanced generation technology. The generation model is fine-tuned using the initial medical image and the corresponding initial pathology text data. The optimized model can more accurately generate corresponding medical images based on the pathology text, thereby improving the adaptability and generation quality of the model. Multiple target pathology text data are obtained, and the optimized model is used for analysis to generate multiple target medical images. This step realizes the effectiveness of the model in practical applications.
[0035] The entire technical solution combines transfer learning and diffusion models to achieve high-quality generation from pathology text to medical images, overcome the problems of limited medical data and inconsistent data distribution, and improve the effect and application value of medical image generation. This method not only improves the generation ability of the model, but also provides important technical support for clinical research and medical diagnosis.
[0036] It should be noted that large text-to-image models are capable of generating high-quality and diverse images from given text prompts. However, these models lack the ability to imitate the appearance of subjects in a given reference set and synthesize their new representations in different contexts. To address this challenge, the present application embodiment proposes a solution in which a pre-trained text-to-image model is fine-tuned using a few images of a specific subject as input. This fine-tuning process enables the model to learn to associate unique identifiers with specific subjects, thereby facilitating the generation of realistic and novel subject images in a variety of scenarios using these identifiers.
[0037] The main purpose is to fine-tune the U-Net component while keeping other components frozen to improve the baseline stablediffusion model, aiming to generate better images specific to the medical field. During the fine-tuning process, a large amount of task-specific medical image data is first prepared, which covers various situations that may occur in the target task and is pre-processed, such as cleaning, word segmentation, encoding, etc. Then, a pre-trained model suitable for the task is selected, such as Stable Diffusion. The fine-tuning strategy involves freezing certain layers of the model, especially the layers close to the input layer, which usually contain more general feature extractors, while training the layers close to the output layer to learn task-specific features.
[0038] Through this fine-tuning strategy, we are able to leverage the knowledge learned by the pre-trained model on a wide range of data, reduce computing resources and time, and improve the accuracy, efficiency, and robustness of the model on specific tasks. In addition, fine-tuning can also improve the generalization ability of the model, allowing it to perform well on new samples. During the fine-tuning process, we used different techniques, such as LoRA, Adapter Tuning, Prefix Tuning, etc. These techniques achieve the purpose of fine-tuning by introducing additional trainable parameters or adjusting specific parts of the model, while keeping most of the pre-trained weights unchanged, reducing the required computing resources and storage space. Finally, the technical effects achieved by the medical image generation method based on the migration model provided in the embodiments of the present application are as follows: Figure 2 As shown, Figure 2From left to right are the original x-ray image, the image generated by the initial stable diffusion, and the image generated by our improved stable diffusion. Among them, the pathological analysis of the x-ray image is: mild opacity in the bilateral lower lobes, enlarged cardiac contour, no pleural effusion or pneumothorax. It can be observed that the original x-ray image on the left shows an undesirable black border, while the image generated by the initial stable diffusion in the middle is of low quality. The image generated by fine-tuning the stable diffusion not only successfully eliminates the irrelevant black border, but also accurately captures the expected content specified in the given prompt.
[0039] In one example, the initial medical image and the initial pathology text data corresponding to the initial medical image are used to fine-tune the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, including: using the initial pathology text data corresponding to the initial medical image, adjusting the CLIP module of the initial diffusion model-based text-to-image generation model, the adjusted CLIP module is used to organize and output the target pathology text data; using the initial medical image and the initial pathology text data corresponding to the initial medical image, adjusting the VAE module of the initial diffusion model-based text-to-image generation model, the adjusted VAE module is used to output a basic medical image for describing the target pathology text data based on the target pathology text data output by CLIP; using the initial medical image, adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, the adjusted U-Net module is used to perform noise prediction processing and iterative denoising processing on the basic medical image output by the VAE module, and then output the target medical image; obtaining the optimized diffusion model-based medical text-to-medical image generation model.
[0040] It should be noted that due to the lack of training data, only 106 cases of STAs-positive lung adenocarcinoma and 97 cases of STAs-negative lung adenocarcinoma were included. In addition, during CT scans, patients may show slight angle changes relative to the CT scanner. In order to solve this problem, in an optional example, after fine-tuning the initial diffusion model-based text-to-image generation model using the initial medical image and the initial pathological text data corresponding to the initial medical image, the medical image generation method also includes: performing data enhancement processing on the initial medical image to obtain an X-ray enhanced image, wherein the data enhancement processing includes: slight rotation, flipping and cropping; using the X-ray enhanced image to perform secondary training on the fine-tuned diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text to medical image generation model.
[0041] In other words, in order to cope with the slight angle changes and limited data volume during CT scanning, the initial medical images were augmented by slight rotation, flipping and cropping, thereby expanding the training dataset. Then, these enhanced initial medical images were used to retrain the fine-tuned diffusion model-based text-to-image generation model to further optimize the model's generation capabilities.
[0042] In addition, in the stable diffusion denoising process, this application adopts two sampling methods: Denoising diffusion implicit model (DDIMS) and k-sample.
[0043] First, before adjusting the initial U-Net module of the diffusion model-based text-to-image generation model, the medical image generation method also includes: defining the sampling process of the fine-tuned U-Net module of the diffusion model-based text-to-image generation model as a non-Markov chain, allowing the U-Net module to sample the basic medical image based on the non-Markov chain, and completing iterative denoising based on the sampling results.
[0044] That is, the DDIM model is an improvement on the denoising diffusion probability model (DDPM) model. DDPM is a class of latent variable models inspired by non-equilibrium thermodynamics. It generates samples by simulating multiple steps of the Markov chain, so that it can achieve high-quality image generation without adversarial training, thereby solving the common mode collapse problem in generative model training. However, the sampling speed of DDPM is slow, while DDIM defines the image sampling process as a non-Markov chain, making it more efficient than DDPM in generating high-quality samples.
[0045] Secondly, before adjusting the initial U-Net module of the diffusion model-based text-to-image generation model, the medical image generation method also includes: introducing negative prompts into the U-Net module of the diffusion model-based text-to-image generation model that has been fine-tuned to enhance the image generation constraints of the U-Net module when performing noise prediction processing and iterative denoising processing on basic medical images.
[0046] Specifically, introducing negative prompts in the U-Net module of the diffusion model-based text-to-image generation model can effectively enhance the image generation constraints during noise prediction and iterative denoising of basic medical images. The specific content of negative prompts can include the following aspects: 1. Avoid unnecessary details: For example, “do not include blurred edges” or “avoid excessive background noise” to ensure that the resulting image is clearer.
[0047] 2. Exclude specific features: such as “no artifacts” or “no unnatural colors” to prevent interference features commonly used in medical image analysis from appearing in the generated images.
[0048] 3. Control image quality: Use hints such as "Avoid low contrast" or "No low-resolution details" to ensure that the generated images meet the quality standards for medical applications.
[0049] 4. Restrictions to specific objects: for example, “do not include non-relevant anatomical structures” or “avoid unnecessary lesion labeling” to ensure the medical relevance of the generated images.
[0050] 5. Prevent overfitting: Hints such as “no excessive detail” or “avoid excessive smoothing” prevent the model from relying too much on specific features in the training data when generating images.
[0051] Through these negative cues, the U-Net module can better control the characteristics of the generated image, thereby improving the accuracy and reliability of medical image processing. This method has important application value in the process of medical image generation and denoising, and can help doctors make more accurate diagnoses and treatment plans.
[0052] That is, to ensure that unwanted content is excluded, the given prompt is enhanced by incorporating negative keywords for each generated image. This approach is able to provide additional image generation constraints during the generation process, ensuring the quality and accuracy of the generated images.
[0053] In summary, by adopting DDIM and K-Sample sampling methods, combined with the stable diffusion method, 500 cases of STAS-positive and STAS-negative lung adenocarcinoma were generated. This technical solution that combines multiple sampling methods and negative prompt enhancement not only improves the efficiency and quality of the generation model, but also ensures the stability and reliability of the generated images. In other words, before adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, by defining the sampling process as a non-Markov chain and introducing negative prompts, the medical image generation method can more efficiently perform noise prediction and iterative denoising, thereby generating high-quality target medical images.
[0054] like Figure 3 As shown in Figure 2, other techniques such as DCGAN and RAT-GAN have certain limitations when generating images based on the provided prompts. The images generated by DCGAN have rough textures and low quality, which can be easily identified as artificial images. Although RAT-GAN has improved in image quality, the quality of the generated images is inconsistent. Some images look very realistic, while others can be clearly identified as generated images. In addition, these models are insufficient in generating specific details such as hybrid ground glass nodules (mGGN) and cavity features.
[0055] The technology used in this application combines transfer learning and diffusion models, especially using DDIM and K-sample sampling methods, to achieve high-quality generation from pathology text to target medical images. This technology not only improves the quality and stability of generated images, but also overcomes the problems of limited medical data and inconsistent data distribution. By fine-tuning the StableDiffusion model and introducing negative prompts and data augmentation processing, the generation ability of the model is further optimized.
[0056] The generated images are very similar to real lung adenocarcinoma images in terms of color contrast and texture, and can generate images with complex details such as needles, lobes, and vacuoles according to the prompts. This method not only improves the robustness and generation quality of the model, but also provides important technical support for clinical research and medical diagnosis. Under limited data conditions, high-quality medical images can still be generated, ensuring the stability and reliability of the generated images.
[0057] It should be noted that: Figure 3 The target pathology text data provided were “STAS-negative group: mixed ground glass nodules (mGGN), no spicules, no lobulation, no cavitation, no vacuoles, and clear borders” and “STAS-positive group: mixed ground glass nodules (mGGN), spicules, lobulation, cavitation, no vacuoles, and clear borders”.
[0058] Finally, this application tests the image generation effect of the optimized diffusion model-based medical text to medical image generation model, as follows: First, the STAS dataset is binary classified into STAS positive or STAS negative using the ResNet101 network. In the experimental design, ORI represents a training dataset with only original image data but no generated image data. ORI+"X" means that the training dataset includes both original image data and images generated by DCGAN, RATGAN, or TDASD ("X" represents one of the models). DCGAN, RATGAN, TDASD (DDIM), and TDASD (Ksample) mean that the training dataset does not contain any original real data and consists only of generated data for training purposes.
[0059] like Figure 4 As shown in the figure, compared with the traditional data augmentation method, the data augmentation method using synthetic images generated by GAN and TDASD slightly improved the positive or negative classification performance of the STAS dataset. In subsequent experiments, only the generated lung adenocarcinoma nodule images were used as the training set, while the test set of real lung adenocarcinoma nodule images was kept unchanged. It was observed that when the images generated by the DCGAN and RAT-GAN methods were used for training alone, the accuracy and area under the curve (AUC) were significantly reduced compared with the previous methods that used a combination of real data and generated data or only used real data for training. However, when training with images generated by the TDASD method, it was observed that both the classification accuracy and AUC were close to the performance when training with real data.
[0060] Experimental results show that the data generation method based on Stable Diffusion is able to learn non-pseudo image representations in the training dataset. When DCGAN and RATGAN map high-dimensional data to a set of labels for image representation learning, they may capture spurious features. These spurious features can often accurately predict labels, but they usually lack practical utility for predictions on different datasets or performing other downstream tasks.
[0061] In summary, by comparing the performance of different generation models, this application verifies the superiority of the text-to-image generation model based on Stable Diffusion in medical image generation, and further proves the effectiveness of this method in improving classification accuracy and image quality.
[0062] Second, the dataset of Shanghai Zhongshan Hospital was used as the training set, and the dataset of Shanghai Cancer Center of Fudan University was used as the test set. The model was retrained on DCGAN, RAT-GAN, and Stable Diffusion (SD) to generate 500 STAS-positive and 500 STAS-negative lung nodule images. In the training dataset, there were 106 STAS-positive and 97 STAS-negative. In the test dataset, there were 39 STAS-positive and 64 STAS-negative. Figure 5 As shown, it can be observed that the model trained only on generated data shows classification results close to the accuracy obtained when trained on real data, and even exceeds the real data in terms of AUC metric.
[0063] Experimental results show that by using data generated by Stable Diffusion for training, the model can learn non-pseudo image representations and has good domain adaptation capabilities. When DCGAN and RAT-GAN map high-dimensional data to a set of labels for image representation learning, they may capture false features, while the images generated by Stable Diffusion are closer to real data in terms of image quality and detail fidelity, thus achieving better migration performance between different domains.
[0064] In summary, by comparing the performance of different generation models, this application verifies the superiority of the text-to-image generation model based on Stable Diffusion in medical image generation, and further proves the effectiveness of this method in improving classification accuracy, image quality and domain adaptability.
[0065] In summary: Deep learning provides an important solution in medical image analysis, but its performance is heavily dependent on large amounts of data, which is limited in the field of medical images due to ethical and privacy issues in data collection. To address this challenge, this application uses the Stable Diffusion model and achieves excellent performance in specific domain tasks through large-scale domain model training and customization.
[0066] This application uses the publicly available CheXpert and MIMIC-CXR datasets to fine-tune the Stable Diffusion (SD) model, enabling it to generate corresponding lung X-ray images based on the description. However, due to limited training data, the initial results were of poor quality. To this end, this application designed the TDASD method, which combines traditional data augmentation techniques with Stable Diffusion to significantly improve the quality of the generated nodule images.
[0067] Experimental results show that compared with traditional data augmentation techniques, the TDASD method does not significantly improve accuracy, but shows a slight advantage in AUC metrics. Compared with DCGAN and RAT-GAN methods, the images generated by the TDASD method perform well in both in-domain and out-of-domain classification tasks, indicating that the TDASD model can simplify complex raw data, filter out invalid or redundant information, and generate medical images containing more meaningful and valuable information.
[0068] In addition, this application uses two sampling methods, DDIM and K-sample, to enhance the image generation ability of the model by introducing negative cues. Although the TDASD method shows superior performance in generating high-quality images, it still faces challenges in accurately depicting complex contour descriptions, such as images of needles, lobes, cavities, or vacuoles.
[0069] In summary, this application demonstrates the potential of the Stable Diffusion model in generating medical images, and effectively improves the image generation quality and classification model performance through the TDASD method. This method provides an effective solution to the limited sample size and domain transfer problems in the field of medical images, and has the potential to be extended to other disease areas and different types of images.
[0070] refer to Figure 6 The present invention also provides an image classification method based on medical images, comprising: S5, acquiring a plurality of target medical images, wherein the target medical images are generated by the above-mentioned medical image generation method, and the plurality of target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; S6, training a medical image training classification model based on the multiple target medical images to obtain a target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical images, and add pathological result labels to the target medical images according to the analysis results, wherein the pathological result labels include: negative labels and positive labels.
[0071] refer to Figure 7 , an embodiment of the present invention provides a medical image generation device based on a migration model, comprising: Processing unit 1 is used to obtain a medical image report data set, and pre-process the medical image data set to obtain an initial medical image and initial pathology text data corresponding to the initial medical image; Construction unit 2, used to construct an initial diffusion model-based text-to-image generation model; A migration unit 3, used to adopt the initial medical image and the initial pathological text data corresponding to the initial medical image, to perform migration adjustment on the initial diffusion model-based text-to-image generation model, to obtain an optimized diffusion model-based medical text-to-medical image generation model, wherein the optimized diffusion model-based medical text-to-medical image generation model is used to generate a corresponding medical image according to the pathological text data; The analysis unit 4 is used to obtain a plurality of target pathology text data, and use the optimized diffusion model-based medical text to medical image generation model to analyze the target pathology text data to generate a plurality of target medical images.
[0072] refer to Figure 8 The embodiment of the present invention provides an image classification device based on medical images, comprising: An acquisition unit 5 is used to acquire a plurality of target medical images, wherein the target medical images are generated by the above-mentioned medical image generation method, and the plurality of target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; The training unit 6 is used to train the medical image training classification model based on the multiple target medical images to obtain the target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical image, and add a pathological result label to the target medical image according to the analysis result, wherein the pathological result label includes: a negative label and a positive label.
[0073] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be repeated here.
[0074] Reference Fig. 9 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0075] Those skilled in the art will understand that Fig. 9The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0076] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0077] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0078] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0079] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A medical image generation method based on a migration model, characterized in that: include: Acquire a medical image report data set, and preprocess the medical image data set to obtain an initial medical image and initial pathology text data corresponding to the initial medical image; Build an initial diffusion-based text-to-image generation model; Using the initial medical image and the initial pathological text data corresponding to the initial medical image, the initial diffusion model-based text-to-image generation model is transferred and adjusted to obtain an optimized diffusion model-based medical text-to-medical image generation model, wherein the optimized diffusion model-based medical text-to-medical image generation model is used to generate a corresponding medical image according to the pathological text data; A plurality of target pathology text data are acquired, and the target pathology text data are analyzed using the optimized diffusion model-based medical text-to-medical image generation model to generate a plurality of target medical images.
2. The medical image generation method according to claim 1, characterized in that: Using the initial medical image and the initial pathological text data corresponding to the initial medical image, fine-tuning the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, including: Using the initial pathological text data corresponding to the initial medical image, adjusting the CLIP module of the initial diffusion model-based text-to-image generation model, the adjusted CLIP module is used to organize and output the target pathological text data; Using the initial medical image and the initial pathological text data corresponding to the initial medical image, adjusting the VAE module of the initial diffusion model-based text-to-image generation model, the adjusted VAE module is used to output a basic medical image for describing the target pathological text data based on the target pathological text data output by CLIP; Using the initial medical image, adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, the adjusted U-Net module is used to perform noise prediction processing and iterative denoising processing on the basic medical image output by the VAE module, and then outputting a target medical image; The optimized diffusion model-based medical text to medical image generation model is obtained.
3. The medical image generation method according to claim 1, characterized in that: After fine-tuning the initial diffusion model-based text-to-image generation model using the initial medical image and the initial pathological text data corresponding to the initial medical image, the medical image generation method further includes: Performing data enhancement processing on the initial medical image to obtain an X-ray enhanced image, wherein the data enhancement processing includes: slight rotation, flipping and cropping; The X-ray enhanced image is used to perform secondary training on the fine-tuned diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model.
4. The medical image generation method according to claim 1, characterized in that: Before adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, the medical image generation method further includes: The sampling process of the U-Net module of the text-to-image generation model based on the diffusion model that has been fine-tuned is defined as a non-Markov chain, and the U-Net module is used to sample the basic medical image based on the non-Markov chain, and iterative denoising is performed according to the sampling results.
5. The medical image generation method according to claim 1, characterized in that: Before adjusting the U-Net module of the initial diffusion model-based text-to-image generation model, the medical image generation method further includes: Negative cues are introduced into the fine-tuned U-Net module of the diffusion model-based text-to-image generation model to enhance the image generation constraints of the U-Net module when performing noise prediction and iterative denoising on basic medical images.
6. An image classification method based on medical images, characterized in that: include: Acquire a plurality of target medical images, wherein the target medical images are generated by the medical image generation method described in any one of claims 1 to 7, and the plurality of target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; The medical image training classification model is trained based on the multiple target medical images to obtain a target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical images, and add pathological result labels to the target medical images according to the analysis results, wherein the pathological result labels include: negative labels and positive labels.
7. A medical image generation device based on a migration model, characterized in that: include: A processing unit, used to obtain a medical image report data set, and pre-process the medical image data set to obtain an initial medical image and initial pathology text data corresponding to the initial medical image; A construction unit, used to construct an initial diffusion model-based text-to-image generation model; A migration unit, used to adopt the initial medical image and the initial pathological text data corresponding to the initial medical image, to perform migration adjustment on the initial diffusion model-based text-to-image generation model to obtain an optimized diffusion model-based medical text-to-medical image generation model, wherein the optimized diffusion model-based medical text-to-medical image generation model is used to generate a corresponding medical image according to the pathological text data; The analysis unit is used to obtain a plurality of target pathology text data, and use the optimized diffusion model-based medical text to medical image generation model to analyze the target pathology text data to generate a plurality of target medical images.
8. An image classification device based on medical images, characterized in that: include: An acquisition unit, configured to acquire a plurality of target medical images, wherein the target medical images are generated by the medical image generation method according to any one of claims 1 to 6, and the plurality of target medical images include: target medical images with negative pathological results and target medical images with positive pathological results; A training unit is used to train a medical image training classification model based on the multiple target medical images to obtain a target medical image training classification model, wherein the medical image training classification model is used to identify and analyze the input target medical image, and add a pathological result label to the target medical image according to the analysis result, wherein the pathological result label includes: a negative label and a positive label.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 6 are implemented, or when the processor executes the computer program, the steps of the method described in claim 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented; or, when the processor executes the computer program, the steps of the method according to claim 7 are implemented.
Citation Information
Patent Citations
Medical image classification method based on transfer learning
CN113592027A
Image generation model training method and device, image generation method and device, equipment and medium
CN116912187A
CT image generation method and device based on large model
CN117115291A
Animation image style migration method and system based on Stable Diffusion
CN117495662A
Medical image analysis method and storage medium
CN117853435A
Cited By
Three-dimensional medical image generation method based on cross attention mechanism
CN121170146A
Chest X-ray film data generation method and system, electronic equipment and storage medium
CN122391420A
Chest x-ray data generation method and system, electronic device and storage medium
CN122391420B