Medical image generation method, system and device based on artificial intelligence, medium and program product

By constructing a text image pre-training model and a medical image generation model, and employing a contrastive learning strategy and imaging parameter control, the problem of generating target sequences with specific imaging parameters in existing models is solved. This achieves cross-modal data conversion and data augmentation, thereby improving the quality and efficiency of generated images.

CN121170079APending Publication Date: 2025-12-19SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511259874.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing medical image generation models struggle to generate target sequences for specific imaging parameters, resulting in insufficient consistency and practicality of the generated images in downstream clinical applications.

Method used

We construct and train a text-image pre-training model and a medical image generation model. We adopt a contrastive learning strategy to align paired text and images in the feature space, widen the feature distance between unpaired samples, and use text metadata as instructions for image translation. We also control imaging parameters by scaling factors and bias terms of source and target modalities.

Benefits of technology

It enables cross-modal data conversion between MRI, PET, and CT, multi-parameter magnetic resonance data generation, data enhancement, and data harmonicization, improving the quality and efficiency of generated images and simplifying model computation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170079A_ABST
    Figure CN121170079A_ABST
Patent Text Reader

Abstract

The invention provides a medical image generation method, system and device based on artificial intelligence, a medium and a program product, and relates to the field of image processing. The method comprises the steps that a special text image pre-training model is constructed and trained, the text image pre-training model adopts a comparative learning strategy, paired texts and images are aligned in a feature space, and the feature distance between non-paired samples is increased; and a medical image generation model is constructed and trained, the medical image generation model uses the text image pre-training model to extract text metadata of the source image and the target image, and the text metadata is used as an instruction in the image translation process to translate any available source image into the target image specified by the text. According to the method, the contrast change between the source image and the target image can be modeled more accurately, the operation time of the model can be greatly shortened, the process is simplified, and the relationship between the text and the image is established more explicitly.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical image processing, in particular to a medical image generation method, system, device, medium and program product based on artificial intelligence. BACKGROUND

[0002] Radiological medical images such as magnetic resonance imaging (MRI), positron emission tomography (PET), and computed tomography (CT) provide comprehensive anatomical, physiological, and pathological information for clinicians and researchers as important support tools for disease diagnosis and treatment and scientific research. However, due to the high cost of acquisition, long imaging time, and potential radiation risk to patients, it is difficult to obtain high-quality, multi-modal radiological image data in clinical practice. In addition, the scarcity of high-quality and de-identified medical image data resources has become a key bottleneck for the continuous development of artificial intelligence (AI) in the field of medical imaging. The rapid development of generative AI technology provides a revolutionary solution for medical image data acquisition and application. By generating multi-modal, multi-task, and multi-disease image data, generative AI is expected to break through the whole-chain application system from image generation, image analysis to intelligent diagnosis and treatment, significantly improving clinical diagnosis and treatment efficiency and scientific research capability.

[0003] However, existing medical image generation models focus on specific tasks, such as MRI to PET generation, MRI to CT generation, and MRI cross-sequence generation network structure and training strategy optimization. Such models rely on a single dataset, resulting in insufficient generalization ability for new tasks and real clinical scenarios, and failing to fully exploit the general representation features between multi-modal images. For the MRI cross-sequence generation task, existing methods usually focus on generating one sequence (such as T2w) from another sequence (such as T1w), but ignore the fact that even for the same sequence, the image appearance can change significantly due to differences in imaging parameters such as scanner manufacturer, magnetic field strength, and pulse sequence type.

[0004] Therefore, existing models can only achieve general cross-sequence generation, and are difficult to generate target sequences for specific imaging parameters, thereby limiting the consistency and practicality of generated images in downstream clinical applications. SUMMARY

[0005] In view of the above-mentioned shortcomings of the prior art, the present application aims to provide a medical image generation method, system, device, medium and program product based on artificial intelligence, which solves the technical problem that the prior art is difficult to generate target sequences for specific imaging parameters.

[0006] To achieve the above object and other related objects, the first aspect of the present application provides a medical image generation method based on artificial intelligence, comprising: constructing and training a special text-image pre-training model, the text-image pre-training model adopts a contrast learning strategy, aligns paired text and image in a feature space, and pulls apart the feature distance between non-paired samples; constructing and training a medical image generation model, the medical image generation model extracts text metadata of a source image and a target image using the text-image pre-training model, uses the text metadata as an instruction in an image translation process, and translates any available source image into a text-specified target image; and using the trained medical image generation model for inference to convert the source image into the specified target image.

[0007] In some embodiments of the first aspect of the present application, the constructing and training of the special text-image pre-training model that adopts a contrast learning strategy to align paired text and image in a feature space and pull apart the feature distance between non-paired samples further comprises: constructing a text encoder for encoding relevant text corresponding to a three-dimensional medical image into a feature vector; constructing an image encoder for encoding a three-dimensional medical image into a feature vector; minimizing the feature distance between paired samples and maximizing the distance between non-paired samples to achieve consistent learning in a cross-modal feature space; wherein the paired samples refer to a three-dimensional medical image and text corresponding thereto, and the non-paired samples refer to a three-dimensional medical image and text not corresponding thereto.

[0008] In some embodiments of the first aspect of the present application, the constructing and training of the medical image generation model that extracts text metadata of a source image and a target image using the text-image pre-training model, uses the text metadata as an instruction in an image translation process, and translates any available source image into a text-specified target image further comprises: a text encoding stage that adopts the text-image pre-training model and freezes the parameters therein, is used for extracting text features of source and target modalities, and respectively obtains scaling factors and bias terms of the source and target modalities; an image encoding stage that adopts a ResNet architecture, is used for extracting image features of the source modality; a multi-modal fusion stage that regulates imaging parameters of the source and target modalities through the scaling factors and bias terms of the source and target modalities, so that the image features of the source modality are converted into image features with contrast features of the target modality; and an image decoding stage that generates an image of the target modality.

[0009] In some embodiments of the first aspect of the present application, the text encoding stage that adopts the text-image pre-training model and freezes the parameters therein, is used for extracting text features of source and target modalities, and respectively obtains scaling factors and bias terms of the source and target modalities further comprises: the scaling factor of the source modality is sThe bias term is β s The scaling factor for the target mode is α. t The bias term is β t .

[0010] In some embodiments of the first aspect of the present invention, in the multimodal fusion stage, the imaging parameters of the source modality and the target modality are adjusted by scaling factors and bias terms of the source modality and the target modality, so that the image features of the source modality are converted into image features with the contrast features of the target modality, and the process further includes:

[0011] Equation 1 is used to transform the image features of the source modality into image features with the contrast features of the target modality.

[0012]

[0013] Among them, f t For image features with target modal contrast characteristics, f s For the image features of the source mode, α s β is the scaling factor for the source mode. s α is the bias term of the source mode. t β is the scaling factor for the target mode. t This is the bias term for the target mode.

[0014] In some embodiments of the first aspect of the present invention, the medical image generation model is further comprising: employing a unified multi-task learning paradigm during training and supervising the quality of the generated images through an L1 loss function.

[0015] To achieve the above and other related objectives, a second aspect of the present invention provides an artificial intelligence-based medical image generation system, comprising: a text-image pre-training module for aligning paired text and images in a feature space and widening the feature distance between unpaired samples; and a general medical image generation module for extracting text metadata of source and target images, and using the text metadata as instructions during image translation to translate any available source image into a target image specified by the text.

[0016] To achieve the above and other related objectives, a third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the artificial intelligence-based medical image generation method provided in the first aspect of the present invention.

[0017] To achieve the above and other related objectives, a fourth aspect of the present invention provides a computer program product comprising computer program code, wherein when the computer program code is run on a computer, the computer implements the artificial intelligence-based medical image generation method provided in the first aspect of the present invention.

[0018] To achieve the above and other related objectives, a fifth aspect of the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the artificial intelligence-based medical image generation method provided in the first aspect of the present invention.

[0019] As described above, the artificial intelligence-based medical image generation method, system, device, medium, and program product provided by the present invention have the following beneficial effects:

[0020] This invention realizes cross-modal data conversion, multi-parameter magnetic resonance data generation, data augmentation, data harmonization, and data amplification tasks among MRI, PET, and CT during a single model inference process. Specifically, this invention pre-trains a dedicated text-image pre-training model, which is used as the text encoder in the medical image generation model, simultaneously obtaining the scaling factors and bias terms of the source and target images. In the multimodal fusion stage, using the scaling factors and bias terms as two parameters enables more accurate modeling of the contrast changes between the source and target images, while significantly reducing the model's computation time, simplifying the process, and more explicitly establishing the relationship between text and images. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the artificial intelligence-based medical image generation method of the present invention.

[0022] Figure 2 This is a flowchart illustrating the sub-steps of step S1 in this invention.

[0023] Figure 3 This is a flowchart illustrating the sub-step of step S2 in this invention.

[0024] Figure 4 This is a flowchart illustrating the sub-step of step S3 in this invention.

[0025] Figure 5 This is the pre-training framework for the text image encoder based on the contrastive learning strategy in this invention.

[0026] Figure 6 This is the framework for a medical image generation model based on imaging parameters in this invention.

[0027] Figure 7The visualization results obtained using the medical image generation model provided by this invention are shown below. (a), (b), and (c) illustrate the unified multi-task image translation results for three cases. (a) Synthesizes various MRI sequences, PET images with different tracers, CT images, and a 7T T1 MRI from a 3T T1 MRI scan. Black boxes with white diagonal stripes indicate modes not acquired from the scanner. (b) Multimodal MRI, CT, and AV45 PET images generated from FDG PET. (c) Synthesizes multimodal MRI, AV45 PET, and FDG PET images from CT.

[0028] Figure 8 This is a schematic block diagram of the artificial intelligence-based medical image generation system of the present invention.

[0029] Figure 9 This is a schematic block diagram of the computer device provided by the present invention. Detailed Implementation

[0030] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0031] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:

[0032] <1> ResNet is a type of deep network that can address the degradation problem of deep networks through residual learning.

[0033] <2> Cross-attention: Cross-attention establishes an attention relationship between two different sequences. The core of cross-attention is that it allows a query from one sequence to focus on the keys and values ​​of another sequence, thereby achieving information fusion.

[0034] <3> L1 loss function: minimizes the sum of the absolute differences between the target value and the estimated value.

[0035] To address related technical problems, this invention introduces an artificial intelligence-based medical image generation method capable of cross-modal conversion between three-dimensional medical images, such as conversion between MRI, PET, and CT images. It should be understood that the method provided in this embodiment, besides being applied to three-dimensional medical images, can also be applied, after improvement and adjustment, to other types of medical images, such as ultrasound images, pathological slide images, and endoscopic images. Furthermore, in addition to its application in the medical field, it can also be used in industrial inspection, video surveillance, military defense, aerospace, autonomous driving, and other fields; this invention does not limit its application to these areas.

[0036] Figure 1 This document illustrates a flowchart of an artificial intelligence-based medical image generation method according to an embodiment of the present invention. The method in this embodiment mainly includes the following steps:

[0037] Step S1: Construct and train a dedicated text image pre-training model. The text image pre-training model adopts a contrastive learning strategy to align paired text and images in the feature space and widen the feature distance between unpaired samples.

[0038] In the embodiments of the present invention, the specific implementation process of step S1 can be further divided into: Figure 2 and Figure 5 The following steps are shown:

[0039] Step S1a: Construct a text encoder to encode the relevant text corresponding to the 3D medical image into feature vectors. The text image pre-trained model can be based on the CLIP model, while extending the structure of the text encoder in the CLIP model by increasing the length of the original token (sequence fragment) to enhance the text encoder's ability to process long texts. Additionally, the relevant text can include: Age, Gender, Scanner, Modality, Voxel size, etc.

[0040] Step S1b: Construct an image encoder to encode 3D medical images into feature vectors. The text image pre-trained model can be based on the CLIP model, while the image encoder in the CLIP model is structurally extended to support 3D medical image input. The 3D medical images can include six structural MRI sequences (T1w, T2w, FLAIR, PD, T2, and contrast-enhanced T1w), PET scans with three different radiotracers (FDG, AV45, and TAU), and CT images.

[0041] Step S1c: Minimize the feature distance between paired samples and maximize the distance between unpaired samples to achieve consistent learning across the modal feature space. Paired samples refer to 3D medical images and their corresponding text, while unpaired samples refer to 3D medical images and text that do not correspond to them. For example, patient 1's CT image and patient 1's CT data are paired samples, while patient 1's CT image and patient 2's CT data are unpaired samples, and so on.

[0042] Specifically, during the training process, the text-image pre-trained model uses the InfoNCE loss function as supervision. By bringing paired text and images closer together and pushing unpaired text and images further apart, the text encoder and image encoder are trained together, enabling the encoder to effectively extract corresponding semantic and content information. The InfoNCE (NoiseContrastive Estimation) loss function is a commonly used loss function in contrastive learning, aiming to maximize the distance between paired samples and minimize the distance between unpaired samples.

[0043] Step S2: Construct and train a medical image generation model. The medical image generation model uses the text image pre-training model to extract text metadata of the source image and the target image. During the image translation process, the text metadata is used as instructions to translate any available source image into a target image specified by the text.

[0044] In the embodiments of the present invention, the specific implementation process of step S2 can be further divided into: Figure 3 and Figure 6 The following steps are shown:

[0045] Step S2a: In the text encoding stage, a pre-trained text image model is used, and its parameters are frozen. This model is used to extract text features from both the source and target modalities, and to obtain the scaling factors and bias terms for both modalities. These are used to capture and simulate changes in image contrast and intensity features during cross-modal transformation. The scaling factor for the source modality is α. s The bias term is β s The scaling factor for the target mode is α. t The bias term is β t .

[0046] Step S2b: Image encoding stage, using a ResNet architecture to extract image features from the source modality. This stage typically includes an image encoder, which can employ a ResNet architecture, such as ResNet-18, ResNet-34, or ResNet-51. Specifically, the image encoding stage uses an improved ResNet architecture that omits downsampling operations to preserve the original resolution of the image's anatomical structure.

[0047] Step S2c: In the multimodal fusion stage, the imaging parameters of the source and target modalities are adjusted by using scaling factors and bias terms of the source and target modalities, transforming the image features of the source modal into image features with the contrast characteristics of the target modal. Specifically, the image features of the source modal are transformed into image features with the contrast characteristics of the target modal using Equation 1:

[0048]

[0049] Among them, f t For image features with target modal contrast characteristics, f s For the image features of the source mode, α s β is the scaling factor for the source mode. s α is the bias term of the source mode. t β is the scaling factor for the target mode. t This is the bias term for the target modality. Specifically, during image generation, the scaling factor α of the source modality is first used. s and bias term β s The source mode f is obtained through mathematical operations. s Contrast features are removed to obtain general image features, and then the scaling factor α of the target modality is applied. t and bias term β t As an instruction, it guides the generation of a target modality f with specified contrast features from general image features. t .

[0050] Compared to single-parameter modulation methods (which encode imaging parameters as learnable features), the two-parameter method (which employs scaling factors and bias terms as provided in this invention) more accurately models the contrast variations between the source and target images. Furthermore, compared to commonly used cross-attention mechanisms, the two-parameter modulation method significantly reduces model computation time, simplifies the process, and more explicitly establishes the relationship between text and images.

[0051] Step S2d: Image decoding stage, generating the image of the target modality. This stage typically includes an image decoder, which can be... During backpropagation, a unified multi-task learning paradigm is employed, and the quality of the generated image is supervised using the L1 loss function.

[0052] In summary, in step S2, the text image pre-trained model from step S1 is used as a text encoder to extract text features of the source modality and the target modality, and obtain the scaling factor and bias term of the source modality and the target modality respectively. After the source modality image is encoded to obtain the source features, the imaging parameters of the source modality and the target modality are adjusted, such as the scaling factor and the bias term, so that the image features of the source modality are converted into image features with the contrast features of the target modality. Finally, the image is decoded by the image decoder to generate the target modality image.

[0053] In the embodiments of the present invention, the specific implementation process of step S3 can be further divided into: Figure 4 The following steps are shown:

[0054] Step S3: Use the trained medical image generation model to perform inference to convert the source image into the specified target image.

[0055] Step S3a: Input the source image and the corresponding text, as well as the target text to be generated.

[0056] Step S3b: The medical image generation model converts the source image into a target image based on the input target text, and generates the target text corresponding to the target image.

[0057] Verification Example

[0058] Dataset Construction: A large-scale brain imaging dataset was constructed, drawing from 11 publicly available datasets and 5 internal clinical datasets, containing 67,974 3D scans from 16,104 subjects. In terms of diseases, the dataset covers a variety of common neurological disorders, such as Alzheimer's disease, white matter hyperintensity lesions, stroke, and brain tumors.

[0059] In terms of image data, the data covers six structural MRI sequences (T1w, T2w, FLAIR, PD, T2, and contrast-enhanced T1w) throughout the entire life cycle from infancy to old age; PET scans with three different radiotracers (FDG, AV45, and TAU); and CT data. Regarding text data, the content covers multiple equipment manufacturers, including Siemens, GE, Philips, and UIH. The magnetic field strength of the MRI ranges from 1.5T to 7T.

[0060] Preprocessing: Since T1w images have clearer structural information than other modal images, the ANTsPy tool was first used to register all modal images of each subject to their corresponding T1w images. Then, the BrainParc tool was used to perform skull stripping on all images.

[0061] Pre-training Phase 1: The entire dataset is divided into different categories based on imaging parameters, and the training batch size is designed accordingly. In each forward propagation, each batch contains an image-text pairing sample from each category.

[0062] The second stage of pre-training involves constructing a dedicated text image pre-training model and a medical image generation model. The text image pre-training model is used to pair samples and serves as the text encoder in the medical image generation model. The medical image generation model is used to generate target images and their corresponding text based on source images and the text in the source images.

[0063] Test data processing: A raw T1-weighted MRI scan image in DICOM format was collected and processed. The image was then converted to Nifti format using the dicom2nifti software package. Simultaneously, the textual imaging parameter information and the age and gender information of the subjects were read from the DICOM header file. Subjects could be from all regions of the world and of all ages.

[0064] Model input data preprocessing: For text data, the read information is written using a specific template to construct text prompts that the model can use. A specific text example is as follows: 1) MRI, {Age:{34};Gender:{M};Scanner:{3.0Siemens};Modality:{T1w};Voxel size:{(0.8,0.8,0.8)};Imagingparameter TR(ms),TE(ms),TI(ms),and FA(degree):{(2500.0,2.2,1000.0,8.0};2) CT, {Age:{50};Gender:{M};Scanner:{Siemens};MFM:{BioGraph 40};Modality:{CT};Voxelsize:{(1.0,1.0,1.5)};TV(kVp),TC(mA):{(120.0,120.0)}};PET,{Age:{22};Gender:{F};Scanner:{Philips};Modality:{tau PET};MFM:{BioGraph mMR};Voxel size:{(1.0,1.0,1.0)};Isotope:{F18},Dose(mC):9.6}. To uniformly implement multiple tasks in a single inference process, in addition to constructing the input T1-weighted MRI text, different text prompts need to be constructed for each task. For example, the cross-modal generation task requires constructing the corresponding generated PET or CT text; the data augmentation task requires constructing the corresponding high-field-strength T1-weighted MRI text; and the data normalization task requires constructing T1-weighted MRI text with specified scanner and imaging parameters. For image data, skull stripping is required.

[0065] Model Input: In this embodiment, the input to the medical image generation model is a T1-weighted MRI image, its corresponding text, and target-generated text prompts. Multiple target-generated texts can be placed in a single .txt file for easy sequential reading during a single inference process. Note that the medical image generation model is an end-to-end model; after input data is fed into the model, it automatically performs computations on the GPU without further processing. After automatic conversion by the model, multiple user-defined image data can be output.

[0066] Figure 7The visualization results obtained using the medical image generation model are shown. (a), (b), and (c) present the unified multi-task image translation results for three cases: (a) Synthesis of various MRI sequences, PET images with different tracers, CT images, and 7T T1 MRI from a 3T T1 MRI scan. Black boxes with white diagonal stripes indicate modalities not acquired from the scanner. (b) Multimodal MRI, CT, and AV45 PET images generated from FDGPET. (c) Synthesis of multimodal MRI, AV45PET, and FDG PET from CT.

[0067] Thus, the AI-based medical image generation and method provided by this invention realizes cross-modal data conversion, multi-parameter magnetic resonance data generation, data enhancement, data harmonization, and data amplification tasks between MRI, PET, and CT during a single model inference process.

[0068] Since the model training process, inference process, and the entire implementation process have been described in detail in the above embodiments, they will not be repeated here.

[0069] Figure 8 This is a schematic block diagram of an artificial intelligence-based medical image generation system provided in an embodiment of the present invention. Figure 5 As shown, the system includes: a text image pre-training module 51 and a general medical image generation module 52.

[0070] The text-image pre-training module 51 is used to align paired text and images in the feature space and increase the feature distance between unpaired samples. The general medical image generation module extracts text metadata from source and target images, using this metadata as instructions during image translation to translate any available source image into a text-specified target image.

[0071] According to the method provided in the embodiments of the present invention, the present invention also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the above-described method.

[0072] According to the method provided in the embodiments of the present invention, the present invention also provides a computer program product, the computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform... Figures 1 to 4 The method for generating medical images based on artificial intelligence is shown.

[0073] Figure 9This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device includes: at least one processor 601, a memory 602, at least one network interface 603, and a user interface 605. The various components in the device are coupled together via a bus system 604. It is understood that the bus system 604 is used to implement communication between these components. In addition to a data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 9 The general will label all buses as bus systems.

[0074] In this embodiment of the invention, the memory 602 is used to store various types of data to support the operation of the electronic terminal 600. Examples of this data include: any executable program for operation on the electronic terminal 600, such as the operating system 6021 and application program 6022; the operating system 6021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 6022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The implementation of the AI-based medical image generation provided in this embodiment of the invention can be included in the application program 6022.

[0075] The methods disclosed in the above embodiments of the present invention can be applied to, or implemented by, processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed through integrated logic circuits in the hardware of processor 601 or through software instructions.

[0076] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0077] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A medical image generation method based on artificial intelligence, characterized in that, include: A dedicated text-image pre-training model is constructed and trained. The text-image pre-training model adopts a contrastive learning strategy to align paired text and images in the feature space and widen the feature distance between unpaired samples. A medical image generation model is constructed and trained. The medical image generation model uses the text image pre-training model to extract text metadata of source images and target images. During the image translation process, the text metadata is used as instructions to translate any available source image into a target image specified by the text. The trained medical image generation model is used for inference to convert the source image into a specified target image.

2. The artificial intelligence-based medical image generation method according to claim 1, characterized in that, The construction and training of a dedicated text-image pre-training model, which employs a contrastive learning strategy to align paired text and images in the feature space and increase the feature distance between unpaired samples, further includes: A text encoder is constructed to encode the relevant text corresponding to the 3D medical image into a feature vector; Construct an image encoder to encode 3D medical images into feature vectors; Consistent learning across modal feature spaces is achieved by minimizing the feature distance between paired samples and maximizing the distance between unpaired samples; where paired samples refer to 3D medical images and their corresponding text, and unpaired samples refer to 3D medical images and their uncorresponding text.

3. The artificial intelligence-based medical image generation method according to claim 1, characterized in that, The construction and training of the medical image generation model, wherein the medical image generation model utilizes the text image pre-trained model to extract text metadata of the source image and the target image, and uses the text metadata as instructions during the image translation process to translate any available source image into a target image specified by the text, further includes: In the text encoding stage, a text image pre-trained model is used and its parameters are frozen to extract text features of the source modality and the target modality, and to obtain the scaling factor and bias term of the source modality and the target modality, respectively. In the image encoding stage, the ResNet architecture is used to extract image features from the source modality; In the multimodal fusion stage, the imaging parameters of the source and target modes are adjusted by scaling factors and bias terms of the source and target modes, so that the image features of the source mode are converted into image features with the contrast features of the target mode. In the image decoding stage, an image of the target modality is generated.

4. The artificial intelligence-based medical image generation method according to claim 3, characterized in that, In the text encoding stage, a pre-trained text image model is used, and its parameters are frozen. This model is used to extract text features from the source and target modalities, and to obtain the scaling factors and bias terms for the source and target modalities, respectively. The scaling factor for the source modality is α. s The bias term is β s The scaling factor for the target mode is α. t The bias term is β t .

5. The artificial intelligence-based medical image generation method according to claim 4, characterized in that, In the multimodal fusion stage, the imaging parameters of the source and target modalities are adjusted by scaling factors and bias terms of the source and target modalities to transform the image features of the source modalities into image features with the contrast features of the target modalities. This also includes: Equation 1 is used to transform the image features of the source modality into image features with the contrast features of the target modality. Among them, f t For image features with target modal contrast characteristics, f s For the image features of the source mode, α s β is the scaling factor for the source mode. s α is the bias term of the source mode. t β is the scaling factor for the target mode. t This is the bias term for the target mode.

6. The artificial intelligence-based medical image generation method according to claim 3, characterized in that, Also includes: The medical image generation model adopts a unified multi-task learning paradigm during training and uses the L1 loss function to supervise the quality of the generated images.

7. A medical image generation system based on artificial intelligence, characterized in that, include: A text-image pre-training module is used to align paired text and images in the feature space and increase the feature distance between unpaired samples; The general medical image generation module is used to extract text metadata from source and target images. During image translation, the text metadata is used as instructions to translate any available source image into a target image specified by the text.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the artificial intelligence-based medical image generation method according to any one of claims 1 to 6.

9. A computer program product, characterized in that, The computer program product includes computer program code, which, when run on a computer, causes the computer to implement the artificial intelligence-based medical image generation method as described in any one of claims 1 to 6.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the artificial intelligence-based medical image generation method according to any one of claims 1 to 6.