Medical image generation method and device based on improved large model

By conducting migration training of the pre-trained large model with medical image data sets, the large model is improved, which solves the problem that the large model cannot generate medical images and achieves the effect of generating high-quality medical images.

CN120163902APending Publication Date: 2025-06-17LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228813.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Current big models cannot generate medical images.

Method used

By obtaining the marked medical image dataset, the pre-trained large model is migrated and trained to obtain the improved large model, and then the medical image corresponding to the input information is generated based on the improved large model.

Benefits of technology

The large model can generate high-quality medical images, solving the problem that the large model cannot generate medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163902A_ABST
    Figure CN120163902A_ABST
Patent Text Reader

Abstract

The invention provides a medical image generation method and device based on an improved large model. The method comprises the following steps: acquiring a labeled medical image data set; obtaining a pre-training large model, wherein the pre-training large model is obtained by training a general image data set; migrating the pre-trained large model to a medical image data set for secondary training to obtain an improved large model; and generating a medical image corresponding to the input information according to the improved large model. According to the method and the device, the medical image data set is acquired, and the large model is subjected to migration training, so that the large model can generate the corresponding medical image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image recognition technology. Specifically, it relates to a method and device for generating medical images based on an improved large model. Background Art

[0002] Large image generation models are an important research direction in the field of deep learning in recent years, and they perform well in generating high-quality and high-resolution images.

[0003] Current large models cannot generate medical images due to the small number of training samples of medical images. Summary of the Invention

[0004] The problem solved by this application is that current large models cannot generate medical images.

[0005] To solve the above problem, the first aspect of this application provides a method for generating medical images based on an improved large model, including:

[0006] Obtain a labeled medical image dataset;

[0007] Obtain a pre-trained large model, which is trained through a general image dataset;

[0008] Transfer the pre-trained large model to the medical image dataset for secondary training to obtain an improved large model;

[0009] Generate a medical image corresponding to the input information according to the improved large model.

[0010] The second aspect of this application provides a device for generating medical images based on an improved large model, which includes:

[0011] An image acquisition module, which is used to obtain a labeled medical image dataset;

[0012] A model pre-training module, which is used to obtain a pre-trained large model, which is trained through a general image dataset;

[0013] A transfer training module, which is used to transfer the pre-trained large model to the medical image dataset for secondary training to obtain an improved large model;

[0014] An image generation module, which is used to generate a medical image corresponding to the input information according to the improved large model.

[0015] The third aspect of this application provides an electronic device, which includes: a memory and a processor;

[0016] The memory is used to store programs;

[0017] The processor, coupled to the memory, is configured to execute the program for:

[0018] Obtain an annotated medical image dataset;

[0019] Obtain a pre-trained large model, which is trained by a general image dataset;

[0020] Transfer the pre-trained large model to the medical image dataset for secondary training to obtain an improved large model;

[0021] Generate a medical image corresponding to the input information according to the improved large model.

[0022] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned medical image generation method based on an improved large model.

[0023] In the present application, by obtaining a medical image dataset and performing transfer training on the large model, the large model can generate corresponding medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of a medical image generation method based on an improved large model according to an embodiment of the present application;

[0025] Figure 2 Flowchart of a medical image generation method based on an improved large model according to an embodiment of the present application;

[0026] Figure 3 Architecture diagram of an adaptive adjustment module of a medical image generation method based on an improved large model according to an embodiment of the present application;

[0027] Figure 4 Structural block diagram of a medical image generation device based on an improved large model according to an embodiment of the present application;

[0028] Figure 5 Structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] To make the above objects, features, and advantages of the present application more apparent and understandable, the following detailed description of the specific embodiments of the present application will be given with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0030] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application shall have the ordinary meanings understood by those skilled in the art to which this application pertains.

[0031] In view of the above problems, this application provides a new medical image generation solution based on an improved large model, which can perform transfer training on the large model through medical images to eliminate the problem that the current large model cannot generate medical images.

[0032] An embodiment of this application provides a medical image generation method based on an improved large model. The specific solution of this method is Figures 1 - 3 as shown. This method can be executed by a medical image generation device based on an improved large model, and this medical image generation device based on an improved large model can be integrated in electronic devices such as computers, servers, computers, server clusters, data centers, etc. Combining Figure 1 、 Figure 2 as shown, it is a flowchart of a medical image generation method based on an improved large model according to an embodiment of this application; wherein, the medical image generation method based on an improved large model includes:

[0033] S101, obtain a labeled medical image data set;

[0034] Obtain a large amount of medical image data, including X-ray films, CT scans, MRIs, etc., which can be obtained from public medical data sets (such as NIH Chest X-ray, MIMIC-CXR, etc.) or in cooperation with medical institutions. Ensure the diversity and representativeness of the data set, covering different disease types, imaging devices, and patient groups.

[0035] Label the medical images to ensure that each image has a corresponding label or description. The labels can be disease types, anatomical structures, lesion areas, etc.

[0036] Use data augmentation techniques (such as rotation, scaling, flipping, noise addition, etc.) to increase the diversity of the data and improve the generalization ability of the model.

[0037] S102, obtain a pre-trained large model, where the pre-trained large model is trained through a general image data set;

[0038] Select a basic large model suitable for image generation and pre-train the model on a large-scale general image data set (such as ImageNet) to enable it to have basic image generation capabilities.

[0039] S103, transfer the pre-trained large model to the medical image data set for secondary training to obtain an improved large model;

[0040] Transfer the pre-trained model to the medical image dataset for secondary training. In this way, the model can learn the unique features of medical images.

[0041] Design a loss function suitable for medical image generation, such as combining pixel-level loss (L1 / L2 loss) and perceptual loss, to ensure that the generated images are of high quality both visually and medically.

[0042] S104, Generate a medical image corresponding to the input information according to the improved large model.

[0043] After obtaining the corresponding improved large model through training, a predetermined medical image can be obtained by inputting the corresponding information: for example, after inputting "I want to get an X-ray image of the knee joint", an X-ray image of the knee joint can be output based on the improved large model; inputting "Generate a brain MRI image showing a 2-cm diameter tumor in the left frontal lobe with mild edema around it", the corresponding output image can be obtained.

[0044] In this application, by obtaining a medical image dataset and performing transfer training on the large model, the large model can generate corresponding medical images.

[0045] In this application, by performing transfer training on the large model, the existing pre-trained large model can be utilized, greatly reducing the data volume and training workload of the entire improved large model.

[0046] In one implementation, the pre-trained large model is a ChatGPT model, a DeepSeek model, or a StyleGAN model.

[0047] In this application, ChatGPT is a text generation model based on Transformer, which is good at understanding and generating natural language but does not directly support image generation.

[0048] In this application, DeepSeek is a multi-modal large model with text generation ability, but it does not directly support image generation or can only generate simple images.

[0049] In one implementation, the S101, obtaining the labeled medical image dataset, includes:

[0050] Obtain multiple different types of medical images;

[0051] Perform data augmentation and type augmentation on the medical images to obtain enhanced medical images;

[0052] Generate corresponding text descriptions based on the medical images and construct the corresponding relationship between the medical images and the text descriptions;

[0053] The text descriptions are used as training samples and the corresponding medical images are used as annotations to construct the medical image dataset.

[0054] The above steps are aimed at building a high-quality medical image dataset for training the corresponding model.

[0055] In this application, a variety of medical image data are collected, covering different modalities, anatomical parts and disease types, which may include:

[0056] Get data from publicly available medical image datasets, such as:

[0057] BraTS: Brain tumor MRI images. CheXpert: Chest X-ray images. NIHChestX-ray: Chest X-ray images. ISIC: Dermoscopy images.

[0058] Obtain private datasets from hospitals or research institutions (ensuring data compliance and privacy protection).

[0059] Ensure diversity in your dataset, including images from different devices, resolutions, and patient populations.

[0060] In this application, a detailed text description is generated for each medical image, and a correspondence between the image and the text is constructed.

[0061] In this application, the implementation method of generating corresponding text description based on medical images is as follows:

[0062] Manual annotation: Medical experts write descriptions for each image, including anatomical location, lesion characteristics, diagnosis results, etc.

[0063] Automatic generation: Use pre-trained multimodal models (such as CLIP or BLIP) to generate preliminary descriptions. Perform manual revisions on the generated descriptions to ensure accuracy and professionalism.

[0064] Describe the following: Image modality (e.g., MRI, CT, X-ray). Anatomical location (e.g., brain, lung, liver). Lesion characteristics (e.g., tumor size, location, shape). Diagnosis (e.g., benign, malignant, inflammatory).

[0065] Constructing the correspondence between medical images and text descriptions is to establish a mapping relationship between images and texts. The mapping relationship between images and texts is established by storing each medical image and its corresponding text description as a key-value pair (such as JSON or CSV format).

[0066] In this application, the goal of constructing a medical image dataset is to build a graphic and text dataset that can be used to train a multimodal model.

[0067] Dataset Structure: Image Folder: Stores all medical images; Annotation File: Stores the correspondence between images and descriptions (such as JSON or CSV files).

[0068] Dataset Partition: The dataset is partitioned into a training set, a validation set, and a test set (such as 70% training set, 15% validation set, 15% test set).

[0069] In one implementation, the data augmentation and type augmentation of medical images to obtain augmented medical images includes:

[0070] Performing rotation, flipping, scaling, or cropping operations on medical images to generate corresponding medical images;

[0071] Constructing a conversion model between different types of medical images;

[0072] Based on the conversion model, performing format conversion on medical images to generate full-type medical images; The full-type medical images at least include X-ray images, MRI images, CT images, and three-dimensional images.

[0073] Rotation operation: By rotating a medical image by a certain angle (such as 90°, 180°, 270°), images from different perspectives can be simulated, increasing the diversity of the dataset.

[0074] Flipping operation: Horizontally or vertically flipping an image can generate a mirror image.

[0075] Scaling operation: By adjusting the resolution of an image, the imaging effects under different devices or different distances can be simulated. The scaling operation can generate images with different resolutions, enhancing the model's ability to generate multi-scale images.

[0076] Cropping operation: Randomly cropping a sub-region from an image can simulate local lesions or images of different regions.

[0077] It should be noted that in this application, the flipping operation and the rotation operation have different effects on different medical images: For symmetric medical images, different perspectives of medical images can be simulated through the flipping operation or the rotation operation; For asymmetric medical images, corresponding words such as flipping and mirroring will also appear in the text descriptions of the generated medical images after the flipping operation or the rotation operation.

[0078] By constructing a conversion model, converting one type of medical image to another type (such as converting an X-ray image to an MRI image), thereby expanding the type diversity of the dataset.

[0079] In this application, Cycle GAN is an unsupervised image-to-image conversion model that can achieve the conversion between different modality images without paired data. For example, converting X-ray images into MRI images.

[0080] In this application, Pix2Pix is a supervised image-to-image conversion model suitable for conversion tasks of paired data. For example, converting CT images into MRI images.

[0081] In this application, for a supervised model (such as Pix2Pix), paired training data (such as CT-MRI image pairs) need to be prepared. For an unsupervised model (such as Cycle GAN), only images of two modalities need to be prepared.

[0082] Through format conversion, various types of medical images (such as X-ray images, MRI images, CT images, and three-dimensional images) are generated, enabling the dataset to cover a more comprehensive range of medical image types.

[0083] In this application, X-ray image generation: generating X-ray projection images from CT or MRI images.

[0084] In this application, MRI image generation: generating MRI images from CT images.

[0085] In this application, CT image generation: generating CT images from MRI images.

[0086] In this application, three-dimensional image generation: reconstructing three-dimensional images from two-dimensional images or converting from one three-dimensional modality to another.

[0087] In this application, the full type of medical images refers to medical images with all the types defined in this application; for example, if it is specified that the full type of medical images includes at least X-ray images, MRI images, CT images, and three-dimensional images, then a medical image that has X-ray images, MRI images, CT images, and corresponding three-dimensional images can be considered as full-type medical images.

[0088] In this application, if a medical image only has X-ray images, the X-ray images can be respectively converted into MRI images, CT images, and three-dimensional images through a conversion model, thereby making the medical image a full-type medical image.

[0089] In one embodiment, step S103 of migrating the pre-trained large model to a medical image dataset for secondary training to obtain an improved large model includes:

[0090] Obtaining pairs of text descriptions and medical images in the medical image dataset;

[0091] Input the text description into the pre-trained large model to obtain the large model image;

[0092] Input the large model image into the adaptive adjustment module to obtain the predicted medical image;

[0093] Calculate the overall loss based on the large model image, the annotated medical image, and the predicted medical image;

[0094] Iterate the pre-trained large model and the adaptive adjustment module based on the overall loss until the loss converges.

[0095] In this application, read the images and corresponding text descriptions (such as JSON or CSV files) in the dataset, load the images as tensors (such as using PyTorch or TensorFlow), and convert the text descriptions into an input format acceptable to the model (such as tokenized text).

[0096] In this application, use a pre-trained large model (such as ChatGPT or DeepSeek) to generate a preliminary image representation: adjust the output layer of the large model, which can include an image generation module; in this way, input the text description into the pre-trained large model to generate a text representation or feature vector of the image; pass these text representations or feature vectors to the image generation module (such as StyleGAN or diffusion model) to generate a preliminary large model image.

[0097] In this application, optimize the preliminarily generated large model image into a high-quality medical image through the adaptive adjustment module.

[0098] In this application, forward propagation: input the text description into the pre-trained large model to generate the large model image; input the large model image into the adaptive adjustment module to generate the predicted medical image. Calculate the loss: calculate the overall loss based on the large model image, the annotated medical image, and the predicted medical image. Backward propagation: calculate the gradients and update the parameters of the pre-trained large model and the adaptive adjustment module.

[0099] In one implementation, the calculating the overall loss based on the large model image, the annotated medical image, and the predicted medical image includes:

[0100] Split the text description to obtain split keywords;

[0101] Calculate the first loss based on the split keywords, the large model image, and the annotated medical image;

[0102] Calculate the second loss based on the split keywords, the predicted medical image, and the annotated medical image;

[0103] Calculate the overall loss based on the first loss and the second loss.

[0104] In this application, key information (such as anatomical parts, lesion characteristics, etc.) is extracted from the text description to guide loss calculation.

[0105] In this application, after extracting keyword information from the text description, the keyword information is converted into an input format acceptable to the model.

[0106] In this application, a tokenization tool (such as NLTK, spaCy) is used to tokenize the text description, and noun phrases or medical terms are extracted as keywords.

[0107] In this application, the text description can be split, for example, when constructing a medical image dataset, so as to maintain the consistency of keywords throughout the training process.

[0108] In this application, the first loss is to evaluate the consistency between the large model image and the annotated medical image in terms of keyword-related features.

[0109] In this application, the specific calculation process of the first loss can be as follows: Use a pre-trained feature extraction model (such as VGG, ResNet) to extract image features from the large model image and the annotated medical image; According to the split keywords, calculate the part of the image features related to the keywords. For example, if the keyword is "tumor", then extract the features related to tumors; Calculate the difference between the large model image features and the annotated medical image features (such as L1 loss, L2 loss).

[0110] In this application, the second loss is to evaluate the consistency between the predicted medical image and the annotated medical image in terms of keyword-related features.

[0111] In this application, the specific calculation process of the second loss can be as follows: Feature extraction: Use the same feature extraction model to extract features from the predicted medical image and the annotated medical image; Keyword attention mechanism: According to the split keywords, calculate the part of the image features related to the keywords; Loss calculation: Calculate the difference between the predicted medical image features and the annotated medical image features (such as L1 loss, L2 loss).

[0112] In this application, the first loss and the second loss are combined to obtain the final overall loss, which is used to guide model optimization.

[0113] In this application, by splitting the keywords, the consistency between the large model image and the predicted medical image and the annotated medical image is evaluated separately.

[0114] In this application, the first loss and the second loss respectively focus on the quality of the initially generated image and the optimized image.

[0115] In one implementation, combine Figure 3As shown, the process of inputting the large model image into the adaptive adjustment module to obtain the predicted medical image includes:

[0116] Divide the large model image into blocks to obtain independent blocks;

[0117] For each independent block, obtain the first neighborhood block and the second neighborhood block with different spacings;

[0118] Generate the first feature block based on the independent block and the first neighborhood block;

[0119] Generate the second feature block based on the independent block and the second neighborhood block;

[0120] Perform feature compression on the first feature block and the second feature block to obtain a compressed block;

[0121] Traverse all independent blocks and generate the predicted medical image based on the obtained compressed blocks.

[0122] In this application, dividing the large model image into blocks means dividing the large model image into corresponding image blocks through a checkerboard; among them, the image block can be at the pixel level (that is, each pixel is an image block), or at other levels, and the specific division depends on the actual processing situation.

[0123] In this application, the image is divided into blocks of the same size using a sliding window or a fixed step size.

[0124] It should be noted here that if the large model image is a two-dimensional image, it is directly divided into a checkerboard, and each grid is an image block; if the large model image is a three-dimensional image, a plane is selected for checkerboard division, and each grid is a strip-shaped grid with a lot of depth (the depth is the depth of the three-dimensional image), and this strip-shaped grid is an image block.

[0125] Preferably, in this application, each image block is 100 - 1000 pixels, so as to perform more feature calculations between local regions on the basis of ensuring the generation accuracy and reducing the calculation amount.

[0126] In this application, an image block is selected as the independent block, and the adjacent image blocks above, below, left, and right of this independent block are the first neighborhood blocks; the image blocks separated by one grid above, below, left, and right of this independent block are the second neighborhood blocks. The spacings between the first neighborhood block and the second neighborhood block and the independent block are different.

[0127] In this application, neighborhood information is extracted for each independent block to capture local structures.

[0128] In this application, the generation of the first feature block is to generate a local feature representation using an independent block and its first neighborhood blocks. Specifically, it can be: performing convolutional layer and attention layer processing on the independent block and the first neighborhood blocks to obtain the first feature block.

[0129] In this application, the specific structure and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.

[0130] It should be noted that in this application, there are four first neighborhood blocks and multiple first feature blocks.

[0131] In this application, the process of performing convolutional layer and attention layer processing on the independent block and the first neighborhood blocks to obtain the first feature block is as follows: splicing the independent block and the four neighborhood blocks together to form a multi-channel input, using the convolutional layer to extract features from the spliced blocks; using the self-attention mechanism or the channel attention mechanism to enhance important features, calculating the attention weights, and weighting the output of the convolutional layer to enhance important features; splitting the output of the attention layer into multiple feature blocks, and each feature block corresponds to the processing results of the independent block and at least one neighborhood block.

[0132] In this application, the generation of the second feature block is to generate a more extensive local feature representation using the independent block and its second neighborhood blocks. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.

[0133] In this application, the generated feature blocks are compressed into a more compact representation to reduce the computational amount and retain key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.

[0134] In this way, through compression, multiple first feature blocks and second feature blocks are compressed into a compressed block, which corresponds to the independent block in size and position and is used to replace the independent block. All image blocks are replaced by compressed blocks to obtain the predicted medical image.

[0135] In this application, by means of traversal, each image block of the large model image is traversed to obtain the corresponding compressed block.

[0136] In this application, for the image blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are not complete. At this time, the first neighborhood blocks and second neighborhood blocks in the relative positions are copied for complementation. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.

[0137] In this application, through complementation, the processing accuracy of the edge image blocks is greatly improved.

[0138] In this application, an adaptive adjustment module is used to capture the similarity relationships between local regions, thereby enhancing feature representation. For images with rich textures or complex structures, high-quality medical images can be generated.

[0139] In this application, due to its own defects, even after secondary training, the pre-trained large model cannot generate high-quality images. By adding an adaptive adjustment module, the pre-trained large model can overcome its own defects and achieve the generation of high-quality medical images.

[0140] An embodiment of this application provides a medical image generation device based on an improved large model, which is used to execute a medical image generation method based on the above content of this application. The following provides a detailed description of the medical image generation device based on an improved large model.

[0141] As Figure 4 shown, the medical image generation device based on an improved large model includes:

[0142] An image acquisition module 101, which is used to acquire an annotated medical image data set;

[0143] A model pre-training module 102, which is used to acquire a pre-trained large model, and the pre-trained large model is trained through a general image data set;

[0144] A transfer training module 103, which is used to transfer the pre-trained large model to the medical image data set for secondary training to obtain an improved large model;

[0145] An image generation module 104, which is used to generate a medical image corresponding to the input information according to the improved large model.

[0146] In one implementation, the pre-trained large model is a ChatGPT model, a DeepSeek model, or a StyleGAN model.

[0147] In one implementation, the image acquisition module 101 is further used for:

[0148] Acquire multiple different types of medical images; perform data enhancement and type enhancement on the medical images to obtain enhanced medical images; generate corresponding text descriptions based on the medical images and construct a correspondence between the medical images and the text descriptions; use the text descriptions as training samples and the corresponding medical images as annotations to construct the medical image data set.

[0149] In one implementation, the image acquisition module 101 is further used for:

[0150] Perform operations such as rotation, flipping, scaling, or cropping on medical images to generate corresponding medical images; construct a conversion model between different types of medical images; perform format conversion on medical images based on the conversion model to generate full-type medical images; the full-type medical images at least include X-ray images, MRI images, CT images, and three-dimensional images.

[0151] In one implementation, the transfer training module 103 is further configured to:

[0152] Obtain the data pairs of text descriptions and medical images in the medical image dataset; input the text description into the pre-trained large model to obtain the large model image; input the large model image into the adaptive adjustment module to obtain the predicted medical image; calculate the overall loss based on the large model image, the labeled medical image, and the predicted medical image; perform iteration on the pre-trained large model and the adaptive adjustment module based on the overall loss until the loss converges.

[0153] In one implementation, the transfer training module 103 is further configured to:

[0154] Split the text description to obtain split keywords; calculate the first loss based on the split keywords, the large model image, and the labeled medical image; calculate the second loss based on the split keywords, the predicted medical image, and the labeled medical image; calculate the overall loss based on the first loss and the second loss.

[0155] In one implementation, the transfer training module 103 is further configured to:

[0156] Divide the large model image into blocks to obtain independent blocks; for each independent block, obtain the first neighborhood block and the second neighborhood block with different spacings; generate the first feature block based on the independent block and the first neighborhood block; generate the second feature block based on the independent block and the second neighborhood block; perform feature compression on the first feature block and the second feature block to obtain the compressed block; traverse all independent blocks and generate the predicted medical image based on the obtained compressed blocks.

[0157] An apparatus for generating medical images based on an improved large model provided by the above embodiments of the present application and a method for generating medical images based on an improved large model provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0158] The above describes the internal functions and structures of an apparatus for generating medical images based on an improved large model. As Figure 5 shown, in practice, the apparatus for generating medical images based on an improved large model can be implemented as an electronic device, including: a memory 301 and a processor 303.

[0159] A memory 301 that can be configured to store programs.

[0160] In addition, the memory 301 can also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0161] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks or optical discs.

[0162] A processor 303, coupled to the memory 301, for executing the programs in the memory 301 for:

[0163] Obtaining an annotated medical image dataset;

[0164] Obtaining a pre-trained large model, where the pre-trained large model is trained by a general image dataset;

[0165] Transferring the pre-trained large model to the medical image dataset for secondary training to obtain an improved large model;

[0166] Generating a medical image corresponding to the input information according to the improved large model.

[0167] In one embodiment, the pre-trained large model is a ChatGPT model, a DeepSeek model or a StyleGAN model.

[0168] In one embodiment, the image acquisition module 101 is further configured to:

[0169] Obtain multiple different types of medical images; perform data enhancement and type enhancement on the medical images to obtain enhanced medical images; generate corresponding text descriptions based on the medical images and construct a correspondence between the medical images and the text descriptions; use the text descriptions as training samples and the corresponding medical images as annotations to construct the medical image dataset.

[0170] In one embodiment, the image acquisition module 101 is further configured to:

[0171] Perform operations such as rotation, flipping, scaling, or cropping on medical images to generate corresponding medical images; construct a conversion model between different types of medical images; perform format conversion on medical images based on the conversion model to generate full-type medical images; the full-type medical images at least include X-ray images, MRI images, CT images, and three-dimensional images.

[0172] In one implementation manner, the transfer training module 103 is further configured to:

[0173] Obtain the data pairs of text descriptions and medical images in the medical image dataset; input the text description into the pre-trained large model to obtain the large model image; input the large model image into the adaptive adjustment module to obtain the predicted medical image; calculate the overall loss based on the large model image, the labeled medical image, and the predicted medical image; perform iteration on the pre-trained large model and the adaptive adjustment module based on the overall loss until the loss converges.

[0174] In one implementation manner, the transfer training module 103 is further configured to:

[0175] Split the text description to obtain split keywords; calculate the first loss based on the split keywords, the large model image, and the labeled medical image; calculate the second loss based on the split keywords, the predicted medical image, and the labeled medical image; calculate the overall loss based on the first loss and the second loss.

[0176] In one implementation manner, the transfer training module 103 is further configured to:

[0177] Divide the large model image into blocks to obtain independent blocks; for each independent block, obtain the first neighborhood block and the second neighborhood block with different spacings; generate the first feature block based on the independent block and the first neighborhood block; generate the second feature block based on the independent block and the second neighborhood block; perform feature compression on the first feature block and the second feature block to obtain the compressed block; traverse all independent blocks and generate the predicted medical image based on the obtained compressed blocks.

[0178] In this application, Figure 5 only some components are schematically shown, which does not mean that the electronic device only includes Figure 5 the components shown.

[0179] The electronic device provided in this embodiment has the same inventive concept as a medical image generation method based on an improved large model provided in the embodiments of this application, and has the same beneficial effects as the method adopted, run, or implemented by the application program stored therein.

[0180] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0181] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0184] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0185] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (Flash RAM). The memory is an example of a computer-readable medium.

[0186] The present application also provides a computer-readable storage medium corresponding to a medical image generation method based on an improved large model provided by the foregoing embodiments. A computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it will execute a medical image generation method based on an improved large model provided by any of the foregoing embodiments.

[0187] Computer-readable media include both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0188] The computer-readable storage medium provided by the above embodiments of the present application and a medical image generation method based on an improved large model provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored thereon.

[0189] It should be noted that a large number of specific details are set forth in the specification provided herein. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0190] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0191] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for generating medical images based on an improved large model, characterized in that: include: Obtain annotated medical image datasets; Obtain a pre-trained large model, where the pre-trained large model is trained by a general image dataset; Migrating the pre-trained large model to a medical image dataset for secondary training to obtain an improved large model; Based on the improved large model, a medical image corresponding to the input information is generated.

2. The method for generating medical images based on an improved large model according to claim 1, characterized in that: The pre-trained large model is a ChatGPT model, a Deepseek model or a StyleGAN model.

3. The method for generating medical images based on an improved large model according to claim 1, characterized in that: The step of obtaining a labeled medical image dataset includes: Acquire multiple different types of medical images; Perform data enhancement and type enhancement on medical images to obtain enhanced medical images; Generate corresponding text descriptions based on medical images and construct the corresponding relationship between medical images and text descriptions; The text descriptions are used as training samples and the corresponding medical images are used as annotations to construct the medical image dataset.

4. The method for generating medical images based on an improved large model according to claim 3, characterized in that: The step of performing data enhancement and type enhancement on the medical image to obtain an enhanced medical image includes: Rotate, flip, scale or crop the medical image to generate the corresponding medical image; Constructing conversion models between different types of medical images; The medical images are format converted based on the conversion model to generate all types of medical images; the all types of medical images include at least X-ray images, MRI images, CT images and three-dimensional images.

5. The method for generating medical images based on an improved large model according to any one of claims 1 to 4, characterized in that: The step of migrating the pre-trained large model to a medical image dataset for secondary training to obtain an improved large model includes: Obtain data pairs of text descriptions and medical images in a medical image dataset; Input the text description into the pre-trained large model to obtain the large model image; The large model image is input into the adaptive adjustment module to obtain a predicted medical image; Calculate the overall loss based on the large model image, the annotated medical image, and the predicted medical image; The pre-trained large model and the adaptive adjustment module are iterated based on the overall loss until the loss converges.

6. The method for generating medical images based on an improved large model according to claim 5, characterized in that: The processor executes a medical image generation method based on an improved large model as described in any one of claims 1-7.