Lesion temporal evolution method, medium and device based on multimodal medical images

Through the lesion image generation model based on multi-attention generation adversarial network, the problem of cross-organ synergistic evolution mechanism and mutual change relationship between multi-modal medical images is solved, and the lesion time evolution law that cannot be predicted from easily obtained images is realized, which improves medical efficiency and treatment effect.

CN118822957BActive Publication Date: 2025-05-09CHINA COAL (TIANJIN) UNDERGROUND ENG INTELLIGENCE RES INST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410801868.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-05-09
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to explore the mechanism of synergistic evolution across organs and its mutual change relationships and lesion evolution laws of multimodal medical images, especially in how to predict poorly acquired images from easily acquired images.

Method used

The lesion image generation model based on the multi-attention generation adversarial network is adopted. By obtaining multimodal medical image data of different time nodes of animals under the preset environment, pre-processing and conditional input processing are performed, the lesion images of each time node of the animal are generated, and corresponding to the human time line is used to obtain the human lesion time evolution law.

Benefits of technology

The lesion time evolution law is extracted from multimodal medical images, helping doctors better understand the growth trend of the lesion, improve medical efficiency and treatment effects, and obtain difficult-to-obtain modal information through more efficient or safer imaging modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118822957B_ABST
    Figure CN118822957B_ABST
Patent Text Reader

Abstract

The present invention provides a method, medium and device for the time evolution of lesions based on multimodal medical images, including: an acquisition module acquires the associated modality A and modality B medical image data of animals at different time nodes under a preset environment; preprocesses the image data to obtain a training set and a test set after preprocessing; processes the modality medical image to obtain two sets of lesion abnormality images, calculates the lesion area and normalizes it; inputs the modality data into a lesion image generation model based on a multi-attention generative adversarial network, adjusts the parameters and saves them, and tests the quality of the generated images with a test set; visualizes the results of the time evolution of the lesion and performs regression modeling analysis. The present invention can obtain modality information through a more efficient or safer imaging modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method, medium and device for temporal evolution of lesions based on multimodal medical images. Background Art

[0002] The occurrence and development of a disease is often not limited to the lesions of a single organ, but is often accompanied by the appearance of other complications. For example, the progression of diabetes can cause the appearance of diabetic retinopathy, diabetic foot, diabetic nephropathy and other lesions.

[0003] Long-term exposure to special environments will cause corresponding changes in the human brain and fundus. For example, long-term exposure to dusty environments may cause dust to enter the brain through the blood, leading to cerebrovascular inflammation, which in turn leads to high signals in the white matter. In addition, retinal inflammation may be induced, leading to retinal hemorrhage, increased exudates, edema, etc. People who are exposed to high altitude environments for a long time may experience symptoms such as hypoxia and altitude sickness. In brain imaging, abnormal changes such as white matter lesions, brain atrophy, and cerebrovascular lesions may occur. In fundus images, abnormal changes such as retinal vascular lesions and macular degeneration may occur.

[0004] At present, most studies focus on predicting and classifying diseases through medical images, without exploring the co-evolution mechanism across organs and their mutual change relationship in multimodal medical images and the law of lesion evolution. In addition, some modal image data are difficult to obtain, and how to predict difficult-to-obtain images from easily accessible images is also a problem to be solved. For example, the function and composition of the blood-retinal barrier are similar to those of the blood-brain barrier, and there are certain similarities between the anatomical characteristics of the small blood vessels in the brain and the composition function of the barrier and the pathology of vascular lesions between the microvessels in the fundus. Therefore, there is a corresponding relationship between the changes in the brain and the fundus under special environments. This invention can predict the evolution of the fundus and the brain in this environment, and can predict the brain magnetic resonance imaging at a certain stage through the fundus images at that stage.

[0005] Patent document CN109844808A discloses an encoding and / or classification method based on training on multiple types of medical image data to identify regions of interest in medical images. Various aspects and / or embodiments seek to provide a method for training an encoder and / or classifier based on multimodal data input so as to classify regions of interest in medical images based on a single modality data input source. However, the invention does not explore the co-evolution mechanism across organs and their mutual change relationship in multimodal medical images and the law of lesion evolution. Summary of the invention

[0006] In view of the defects in the prior art, the object of the present invention is to provide a method, medium and device for temporal evolution of lesions based on multimodal medical images.

[0007] A lesion time evolution method based on multimodal medical images provided by the present invention comprises:

[0008] Step S1: The animal multimodal image acquisition module acquires the associated modality A and modality B medical image data of the animal at different time points under a preset environment;

[0009] Step S2: The preprocessing module preprocesses the image data to obtain a training set and a test set after preprocessing, so that the spatial and grayscale proximity of images of different periods and different modalities meets the preset standards;

[0010] Step S3: The conditional input module processes the modality A and modality B medical images at time node t to obtain two sets of lesion abnormality images, calculates the lesion areas respectively and normalizes them to obtain conditional input;

[0011] Step S4: The lesion time evolution module inputs the modal data into the lesion image generation model based on the multi-attention generative adversarial network, adjusts the parameters and saves them, and the test set tests the quality of the generated image;

[0012] Step S5: Result visualization and modeling analysis, visualizing the results of the time evolution of the lesion and performing regression modeling analysis;

[0013] Step S6: Mapping of animal model and human model: Match the time lines of the lesion features of a certain modality of human image with the lesion features of animal image, explore the correspondence between human and animal based on examples, and obtain the evolution law of multimodal lesion images of human under preset environment.

[0014] Preferably, in step S2:

[0015] The preprocessing steps include image alignment, denoising and pairing, image standardization, and partitioning into training, validation and test sets in proportion;

[0016] All images are standardized using the following formula:

[0017]

[0018] Among them, μ is the mean of the image, x represents the image matrix, σ and N represent the standard deviation and the number of image voxels respectively.

[0019] Preferably, in step S3:

[0020] The modality A and modality B medical images at time node n are A n and B n , A n Subtract A0 from B nSubtract it from B0 to get two sets of lesion abnormality maps. The lesion areas of the two sets of lesion abnormality maps are calculated and normalized to get the conditional input and As the input of the n+1 time node modal B lesion time evolution module, As the input of the time evolution module of the modal A lesion at the n+1 time node;

[0021] Conditional input module, the specific process is as follows:

[0022] A n Subtract it from A to get the abnormal lesion map;

[0023] Use OpenCV to convert the lesion abnormality image into a grayscale image;

[0024] Perform a binarization operation on the grayscale image to convert the target lesion in the image into a binary image;

[0025] Perform contour detection on the binarized image and use the findContours function of OpenCV to obtain the contour information of the lesion;

[0026] According to the contour information, the contourArea function of OpenCV is used to calculate the area of ​​the lesion. The lesion area is calculated and normalized to obtain the conditional input Serves as conditional input to the modality B lesion temporal evolution module.

[0027] Preferably, in step S4:

[0028] For modality A, input A0 and A1 into the lesion image generation model based on the multi-attention generative adversarial network, adjust the parameters and save them, and use the test set to test the quality of the image generated at t1; input A1 and A2 into the network, adjust the parameters and save them, and test the results, and so on, until A n-1 With A n Input the network, adjust the parameters, save them, and test the results; the processing of mode B is the same as that of mode A;

[0029] The lesion image generation model based on multi-attention generative adversarial network, the specific process is as follows:

[0030] For mode A, the resulting conditional vector and A0 as input images, and A1 as target image, and input them into the network; the network includes a generator G based on a multi-attention mechanism, a discriminator D based on Darknet, and a feature loss calculation module;

[0031] By synchronously training the generator and discriminator in the generative adversarial network, the target loss function is reduced until the error between the output image and the target image is less than a preset value, and the model from A0 to A1 is obtained and saved; similarly, A is obtained n-1 To A n and B n-1 To B n Model.

[0032] Preferably, the generator G includes: a lesion attention module and a feature attention module, specifically as follows:

[0033] The lesion attention module generates an attention map of the lesion, detects the lesion area, and keeps other parts unchanged; the lesion attention module inputs A0 and A1, and uses the attention localization loss function AL to minimize the difference between the generated lesion mask and the true mask:

[0034]

[0035] Among them, A is the original input image, A ′ is the generated image, A M is the true lesion mask, balanced by parameters λ1 and λ2, is the maximum likelihood estimate;

[0036] AL is back-propagated to the lesion attention module until the difference between the generated nodule mask and the true nodule mask is minimized; the lesion attention module is F s (·), the attention map uses m = F s (A0) is generated, where m∈[1,0]; the regions with non-zero values ​​are lesion-specific regions, and the rest are lesion-independent regions;

[0037] The feature attention module transforms image A0 into A ′ (G(A0,c)→A ′ ), G(·,c) is the generator output function, the feature attention module modifies the lesion area, and the rest of the image remains unchanged. In order to train the feature attention module, the reconstruction loss function RL is used:

[0038]

[0039] ‖IG(A ′ ,c)‖1 and ‖IG(A,c)‖1 represent the consistency loss and reconstruction loss respectively. For maximum likelihood estimation, the consistency loss ensures that when the image A generated using the conditional vector c ′ The translation back is the same as the original image A, while the reconstruction loss ensures that the input image remains unchanged when it is modified.

[0040] Preferably, the lesion attention module and the feature attention module are based on a U-shaped structure based on the attention mechanism, as follows:

[0041] The U-net architecture consists of an encoder and a decoder: the encoder has n stages for deep feature extraction, and the size of n is determined according to the input medical image; each stage consists of a downsampling layer; the previous layer in the encoder The image features are first converted into two feature spaces f and g to calculate the attention. It is a C×N dimensional vector space, where C is the number of channels and N is the number of feature positions from the previous hidden layer;

[0042] f(x)=W f x,g(x)=W g x, where W f is the weight of f, W g is the weight of g. According to the U-Net network structure, the decoder is symmetrical with the encoder, which also contains n stages. Each stage of the decoder consists of two multi-head self-attention modules and an upsampling layer. The upsampling layer includes bilinear interpolation and convolution layers. Skip connections are used for feature fusion between the encoder and decoder.

[0043] The self-attention module consists of a block multi-head self-attention plus normalization layer and a feed-forward network plus normalization layer, as follows:

[0044] F′=MSA(Norm(F in ))+F in

[0045] F o =FFN(Norm(F′))+F′

[0046] Among them, F in represents the input feature map of the self-attention module, Norm(·) represents the normalization layer, F′ and F o They represent the output features of block MSA and FFN respectively, MSA is a multi-head self-attention module, and FFN is a feed-forward network;

[0047] The specific process of block multi-head self-attention is as follows:

[0048] The input feature map X in Divided into L×L non-overlapping blocks, the i-th block is where i∈{1,2,…,N}, is flattened and transposed to X i , for X i Apply MSA, first X i Linear mapping to query:Q i , key value: Ki And value item value:V i .

[0049] Q i =X i W Q

[0050] K i =X i W K

[0051] V i =X i W V

[0052] Among them, W Q , W K and W V are learnable parameters, representing the projection matrices of query, key value, and value item; Q i , K i and V i Divide into k groups along the channel dimension: For the jth group of self-attention, it is expressed as:

[0053]

[0054] in, and are the query, key, and value items of the jth group, respectively, and d k For each group dimension, the output tokens of the i-th block are: It is expressed as:

[0055]

[0056] Among them, Concat(·) represents concatenation, B represents position embedding, and W O It is possible to learn parameters and merge the outputs of all blocks to obtain the final output feature map X out .

[0057] Preferably, the discrimination module D is designed based on the Darknet model, as follows:

[0058] Darknet layers include convolution, normalization, and leaky ReLU as activation functions. The discriminator D distinguishes the generated fake images from the real ones and applies the regularized adversarial loss term L adv :

[0059]

[0060] Where KL(·) is the Kullback-Leibler divergence, D srcis the loss between the real image and the generated image, D src (A) and D src (A ′ ) represent the probability that image A is a real image and the probability that image A is generated. ′ The probability of a fake image is given by introducing a regularization term to maintain the separability between distributions when training the generator G. When training the generator, the regularization term is minimized to maintain the difference between the source distribution and the target distribution. In the initial stage of training the discriminator, the regularization term is maximized to ensure that there is more overlap between the two distributions. Minimizing the feature loss is achieved by the pre-trained VGG-19 network. The output image A obtained from the generator G ′ The real image A is input into the pre-trained VGG network for feature extraction, and the loss error is back-propagated to update the weight of G. The formula is as follows:

[0061]

[0062] Among them, R(A) is the size of the feature space, VGG(·) is the VGG feature extraction network;

[0063] The objective loss function D of the discriminator is as follows:

[0064] D=-L adv +L VGG .

[0065] Preferably, in step S5:

[0066] The regression analysis process is as follows:

[0067] According to the lesion images at n time nodes, the images and corresponding times are sent to the Resnet network for regression, and the functions y1=f(x1,t) and y2=f(x2,t) are constructed. x1 represents the image of the initial modality A input, x2 represents the image of the initial modality B input, y1 represents the lesion image of modality A at time t, and y2 represents the lesion image of modality B at time t. According to the lesion image of modality A or modality B at a certain moment, it is substituted into the model, the time t is inferred, and the evolution image of the other modality at this time is obtained. The image of modality A or modality B at a certain moment is input, and the lesion images of modality A and modality B at any time t are predicted.

[0068] According to a computer-readable storage medium storing a computer program provided by the present invention, when the computer program is executed by a processor, the steps of the method for temporal evolution of lesions based on multimodal medical images are implemented.

[0069] According to the present invention, a lesion time evolution device based on multimodal medical images includes: a processor, and a memory and a network interface connected to the processor; the network interface is connected to a non-volatile memory in a server; when the processor is running, it calls a computer program from the non-volatile memory through the network interface, and runs the computer program through the memory to execute the steps of the lesion time evolution method based on multimodal medical images.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] 1. The present invention provides a method, device and storage medium for monitoring the temporal evolution of lesions based on multimodal medical images. The lesion image generation model based on the generative adversarial network is used to generate lesion images at various time nodes of animals, and correspond to the human timeline, thereby obtaining the temporal evolution law of human lesions, which can help doctors better understand the growth trend of lesions, facilitate the formulation of more effective treatment plans, and thus improve medical efficiency and treatment effects;

[0072] 2. The present invention matches two modalities one by one, and through matching modeling, data of one modality can be obtained from another modality. Some imaging modalities require a long imaging time or a high dose of radiation, and the model can obtain the modality information through a more efficient or safer imaging modality;

[0073] 3. By establishing a lesion evolution model, the present invention can monitor the course of the patient's disease and help doctors understand the development trend of the disease so as to adjust the treatment strategy in time. For chronic and recurrent diseases, this method can improve the effect of disease management. In addition, by comparing medical images before and after treatment, doctors can more accurately evaluate the treatment effect. This helps guide clinical decision-making, such as adjusting treatment plans or terminating ineffective treatments early;

[0074] 4. The present invention is conducive to individualized treatment. By monitoring the temporal evolution of individual lesions, it helps us better understand the heterogeneity of disease evolution at the individual level. Doctors can formulate more accurate treatment plans based on the specific conditions of patients. It can also help doctors judge the prognosis and accurately indicate the time for review, paving the way for personalized medicine.

[0075] 5. The present invention accurately constructs the evolution law of lesions of animals in special environments through experiments, which can evaluate whether the special environment is suitable for long-term survival and activities of humans, and effectively avoid damage to humans caused by special environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0077] Figure 1 It is a schematic flow chart of the steps for implementing the method of the present invention;

[0078] Figure 2 This is a schematic diagram of step 1 of the present invention;

[0079] Figure 3 This is a schematic diagram of the process of step 2 of the present invention;

[0080] Figure 4 This is a schematic diagram of step 3 of the present invention;

[0081] Figure 5 This is a schematic diagram of step 4 of the present invention;

[0082] Figure 6 This is a schematic diagram of step 5 of the present invention;

[0083] Figure 7 This is a schematic diagram of the process of step 6 of the present invention. DETAILED DESCRIPTION

[0084] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0085] Embodiment 1:

[0086] The present invention belongs to the field of multimodal medical image evolution generation prediction technology, and relates to a lesion time evolution method, device and storage medium based on multimodal medical images. The lesion time evolution method based on multimodal medical images includes: an animal multimodal image acquisition module, a preprocessing module, a conditional input module, a lesion time evolution module, result visualization and modeling analysis, and mapping of animal models and human models. Through a lesion image generation model based on a generative adversarial network with a multi-attention mechanism, lesion images of various time nodes of different animal modes are generated, and correspond to the human timeline, so as to obtain the time evolution law of human lesions, and can obtain modal information that is not easy to obtain through more efficient or safer imaging modalities, and help doctors better understand the growth trend of lesions, which is conducive to formulating more effective treatment plans, and can help doctors judge the prognosis, accurately prompt the review time, and pave the way for personalized medicine.

[0087] The present invention aims to explore the evolution law of damage caused by special environments to humans, evaluate whether the special environment is suitable for long-term survival and activities of humans, and effectively avoid damage caused by special environments to humans. By changing the time parameters, predicting the imaging changes of modal A and modal B at each time node, and forming the evolution process of multimodal medical image lesions, we can better understand the heterogeneity of the individual level of disease evolution, and can help doctors judge the prognosis, accurately indicate the time for review, and provide an accurate evolution prediction model for personalized intervention and prevention. At the same time, it is beneficial for operators to intuitively see the changes in body lesions, lesion sites, and changes in the progression of lesions that may affect other organs, thereby enhancing protection and health awareness and preventing the occurrence of diseases in advance.

[0088] According to a lesion time evolution method based on multimodal medical images provided by the present invention, Figure 1-Figure 7 As shown, including:

[0089] Step S1: The animal multimodal image acquisition module acquires the associated modality A and modality B medical image data of the animal at different time points under a preset environment;

[0090] Step S2: The preprocessing module preprocesses the image data to obtain a training set and a test set after preprocessing, so that the spatial and grayscale proximity of images of different periods and different modalities meets the preset standards;

[0091] Specifically, in step S2:

[0092] The preprocessing steps include image alignment, denoising and pairing, image standardization, and partitioning into training, validation and test sets in proportion;

[0093] All images are standardized using the following formula:

[0094]

[0095] Among them, μ is the mean of the image, x represents the image matrix, σ and N represent the standard deviation and the number of image voxels respectively.

[0096] Step S3: The conditional input module processes the modality A and modality B medical images at time node t to obtain two sets of lesion abnormality images, calculates the lesion areas respectively and normalizes them to obtain conditional input;

[0097] Specifically, in step S3:

[0098] The modality A and modality B medical images at time node n are A n and B n , A n Subtract A0 from B nSubtract it from B0 to get two sets of lesion abnormality maps. The lesion areas of the two sets of lesion abnormality maps are calculated and normalized to get the conditional input and As the input of the n+1 time node modal B lesion time evolution module, As the input of the time evolution module of the modal A lesion at the n+1 time node;

[0099] Conditional input module, the specific process is as follows:

[0100] A n Subtract it from A to get the abnormal lesion map;

[0101] Use OpenCV to convert the lesion abnormality image into a grayscale image;

[0102] Perform a binarization operation on the grayscale image to convert the target lesion in the image into a binary image;

[0103] Perform contour detection on the binarized image and use the findContours function of OpenCV to obtain the contour information of the lesion;

[0104] According to the contour information, the contourArea function of OpenCV is used to calculate the area of ​​the lesion. The lesion area is calculated and normalized to obtain the conditional input Serves as conditional input to the modality B lesion temporal evolution module.

[0105] Step S4: The lesion time evolution module inputs the modal data into the lesion image generation model based on the multi-attention generative adversarial network, adjusts the parameters and saves them, and the test set tests the quality of the generated image;

[0106] Specifically, in step S4:

[0107] For modality A, input A0 and A1 into the lesion image generation model based on the multi-attention generative adversarial network, adjust the parameters and save them, and use the test set to test the quality of the image generated at t1; input A1 and A2 into the network, adjust the parameters and save them, and test the results, and so on, until A n-1 With A n Input the network, adjust the parameters, save them, and test the results; the processing of mode B is the same as that of mode A;

[0108] The lesion image generation model based on multi-attention generative adversarial network, the specific process is as follows:

[0109] For mode A, the resulting conditional vector and A0 as input images, and A1 as target image, and input them into the network; the network includes a generator G based on a multi-attention mechanism, a discriminator D based on Darknet, and a feature loss calculation module;

[0110] By synchronously training the generator and discriminator in the generative adversarial network, the target loss function is reduced until the error between the output image and the target image is less than a preset value, and the model from A0 to A1 is obtained and saved; similarly, A is obtained n-1 To A n and B n-1 To B n Model.

[0111] Specifically, the generator G includes: a lesion attention module and a feature attention module, as follows:

[0112] The lesion attention module generates an attention map of the lesion, detects the lesion area, and keeps other parts unchanged; the lesion attention module inputs A0 and A1, and uses the attention localization loss function AL to minimize the difference between the generated lesion mask and the true mask:

[0113]

[0114] Among them, A is the original input image, A ′ is the generated image, A M is the true lesion mask, balanced by parameters λ1 and λ2, is the maximum likelihood estimate;

[0115] AL is back-propagated to the lesion attention module until the difference between the generated nodule mask and the true nodule mask is minimized; the lesion attention module is F s (·), the attention map uses m = F s (A0) is generated, where m∈[1,0]; the regions with non-zero values ​​are lesion-specific regions, and the rest are lesion-independent regions;

[0116] The feature attention module transforms image A0 into A ′ (G(A0,c)→A ′ ), G(·,c) is the generator output function, the feature attention module modifies the lesion area, and the rest of the image remains unchanged. In order to train the feature attention module, the reconstruction loss function RL is used:

[0117]

[0118] ‖IG(A ′ ,c)‖1 and ‖IG(A,c)‖1 represent the consistency loss and reconstruction loss respectively. For maximum likelihood estimation, the consistency loss ensures that when the image A generated using the conditional vector c ′ The translation back is the same as the original image A, while the reconstruction loss ensures that the input image remains unchanged when it is modified.

[0119] Specifically, the lesion attention module and the feature attention module are based on a U-shaped structure based on the attention mechanism, as follows:

[0120] The U-net architecture consists of an encoder and a decoder: the encoder has n stages for deep feature extraction, and the size of n is determined according to the input medical image; each stage consists of a downsampling layer; the previous layer in the encoder The image features are first converted into two feature spaces f and g to calculate the attention. It is a C×N dimensional vector space, where C is the number of channels and N is the number of feature positions from the previous hidden layer;

[0121] f(x)=W f x,g(x)=W g x, where W f is the weight of f, W g is the weight of g. According to the U-Net network structure, the decoder is symmetrical with the encoder, which also contains n stages. Each stage of the decoder consists of two multi-head self-attention modules and an upsampling layer. The upsampling layer includes bilinear interpolation and convolution layers. Skip connections are used for feature fusion between the encoder and decoder.

[0122] The self-attention module consists of a block multi-head self-attention plus normalization layer and a feed-forward network plus normalization layer, as follows:

[0123] F′=MSA(Norm(F in ))+F in

[0124] F o =FFN(Norm(F′))+F′

[0125] Among them, F in represents the input feature map of the self-attention module, Norm(·) represents the normalization layer, F′ and F o They represent the output features of block MSA and FFN respectively, MSA is a multi-head self-attention module, and FFN is a feed-forward network;

[0126] The specific process of block multi-head self-attention is as follows:

[0127] The input feature map X in Divided into L×L non-overlapping blocks, the i-th block is where i∈{1,2,…,N}, is flattened and transposed to X i , for X i Apply MSA, first X i Linear mapping to query:Q i , key value: K i And value item value:V i .

[0128] Q i =X i W Q

[0129] K i =X i W K

[0130] V i =X i W v

[0131] Among them, W Q , W K and W v are learnable parameters, representing the projection matrices of query, key value, and value item; Q i , K i and V i Divide into k groups along the channel dimension: For the jth group of self-attention, it is expressed as:

[0132]

[0133] in, and are the query, key, and value items of the jth group, respectively, and d k For each group dimension, the output tokens of the i-th block are: It is expressed as:

[0134]

[0135] Among them, Concat(·) represents concatenation, B represents position embedding, and W O It is possible to learn parameters and merge the outputs of all blocks to obtain the final output feature map X out .

[0136] Specifically, the discriminant module D is designed based on the Darknet model, as follows:

[0137] Darknet layers include convolution, normalization, and leaky ReLU as activation functions. The discriminator D distinguishes the generated fake images from the real ones and applies the regularized adversarial loss term L adv :

[0138]

[0139] Where KL(·) is the Kullback-Leibler divergence, D src is the loss between the real image and the generated image, D src (A) and D src (A ′ ) represent the probability that image A is a real image and the probability that image A is generated. ′ The probability of a fake image is given by introducing a regularization term to maintain the separability between distributions when training the generator G. When training the generator, the regularization term is minimized to maintain the difference between the source distribution and the target distribution. In the initial stage of training the discriminator, the regularization term is maximized to ensure that there is more overlap between the two distributions. Minimizing the feature loss is achieved by the pre-trained VGG-19 network. The output image A obtained from the generator G ′ The real image A is input into the pre-trained VGG network for feature extraction, and the loss error is back-propagated to update the weight of G. The formula is as follows:

[0140]

[0141] Among them, R(A) is the size of the feature space, VGG(·) is the VGG feature extraction network;

[0142] The objective loss function D of the discriminator is as follows:

[0143] D=-L adv +L VGG .

[0144] Step S5: Result visualization and modeling analysis, visualizing the results of the time evolution of the lesion and performing regression modeling analysis;

[0145] Specifically, in step S5:

[0146] The regression analysis process is as follows:

[0147] According to the lesion images at n time nodes, the images and corresponding times are sent to the Resnet network for regression, and the functions y1=f(x1,t) and y2=f(x2,t) are constructed. x1 represents the image of the initial modality A input, x2 represents the image of the initial modality B input, y1 represents the lesion image of modality A at time t, and y2 represents the lesion image of modality B at time t. According to the lesion image of modality A or modality B at a certain moment, it is substituted into the model, the time t is inferred, and the evolution image of the other modality at this time is obtained. The image of modality A or modality B at a certain moment is input, and the lesion images of modality A and modality B at any time t are predicted.

[0148] Step S6: Mapping of animal model and human model: Match the time lines of the lesion features of a certain modality of human image with the lesion features of animal image, explore the correspondence between human and animal based on examples, and obtain the evolution law of multimodal lesion images of human under preset environment.

[0149] According to a computer-readable storage medium storing a computer program provided by the present invention, when the computer program is executed by a processor, the steps of the method for temporal evolution of lesions based on multimodal medical images are implemented.

[0150] According to the present invention, a lesion time evolution device based on multimodal medical images includes: a processor, and a memory and a network interface connected to the processor; the network interface is connected to a non-volatile memory in a server; when the processor is running, it calls a computer program from the non-volatile memory through the network interface, and runs the computer program through the memory to execute the steps of the lesion time evolution method based on multimodal medical images.

[0151] Embodiment 2:

[0152] Embodiment 2 is a preferred example of Embodiment 1, and is used to illustrate the present invention in more detail.

[0153] The purpose of the present invention is to provide a method, device and storage medium for the time evolution of lesions based on multimodal medical images, which can provide doctors with quantitative lesion information in the time dimension, which is very helpful for medical image analysis and clinical diagnosis.

[0154] The purpose of the present invention can be achieved through the following technical solutions:

[0155] To achieve the above object, the present invention provides a method for temporal evolution of lesions based on multimodal medical images in a first aspect, the method comprising the following steps:

[0156] Step 1, multimodal image acquisition module: acquiring the associated modality A and modality B medical image data of the human body at different time nodes under a special environment;

[0157] Step 2: Preprocessing module: preprocess the data to obtain the training set and test set after preprocessing, so that the images of different periods and different modes are as consistent as possible in space and grayscale, so as to facilitate the subsequent time evolution analysis;

[0158] Step 3: Conditional input module: define the modality A and modality B medical images at time node t as A t and B t . t Subtract A0 from B tSubtract it from B0 to get two sets of lesion abnormality maps. Calculate the lesion area and normalize it to get the conditional input and As the input of the (t+1) time node modal B lesion time evolution module, As the input of the time evolution module of the modal A lesion at the (t+1) time node, it is used to fuse multimodal information, learn the relationship between modalities, and make the output more realistic.

[0159] Step 4, lesion time evolution module: establish a mathematical model of lesion evolution. For modality A, input A0, A1 and c into the lesion image generation model based on the multi-attention generative adversarial network, adjust the parameters and save them, and use the test set to test the image quality generated at t1; similarly, input A1 and A2 into the network, adjust the parameters and save them, and test the results, and so on, until A is n-1 With A N Input the network, adjust the parameters, save them, and test the results. The processing of mode B is the same as that of mode A;

[0160] Step 5: Result visualization and modeling analysis: Visualize the results of the time evolution of the lesions and perform regression modeling analysis. This model can be used to simulate the laws of human body damage caused by special environments.

[0161] In the step 2:

[0162] The preprocessing steps include image alignment, denoising and pairing. Secondly, the images are standardized to enhance the features of the target area of ​​the image, and the training set, validation set and test set are divided in a ratio of 6:2:2.

[0163] All images are standardized using the following formula:

[0164]

[0165] Among them, μ is the mean of the image, x represents the image matrix, σ and N represent the standard deviation and the number of image voxels respectively.

[0166] Optionally, the lesion image generation model based on the multi-attention generative adversarial network in step 4 is specifically constructed as follows (taking the model from A0 to A1 as an example):

[0167] 1. Use the conditional vector c obtained in step 3 and A0 as the input image and A1 as the target image to input the network;

[0168] 2. The network includes a generator G based on a multi-attention mechanism, a discriminator D based on Darknet, and a feature loss calculation module;

[0169] 3. By synchronously training the generator and discriminator in the generative adversarial network, the target loss function is reduced until the error between the output image and the target image is less than a preset value, and the model from A0 to A1 is obtained and saved.

[0170] Optionally, the generator G includes a lesion attention module and a feature attention module, as follows:

[0171] The main goal of the lesion attention module is to generate an attention map of the lesion, detect the lesion area, and keep other parts unchanged. The input of the lesion attention module is A0 and A1. The difference between the generated lesion mask and the true mask is minimized using the attention localization loss function AL.

[0172]

[0173] Among them, A0 is the original input image, A ′ is the generated image, I M is the true lesion mask, and these terms are balanced by parameters λ1 and λ2. AL is back-propagated to the lesion attention module until the difference between the generated nodule mask and the true nodule mask is minimized. The lesion attention module is denoted as F s (·), the attention map uses m = F s (A0) is generated, where m∈[1,0]. Ideally, m should be 1 for the lesion area and 0 for the rest of the area. In practice, since F s Due to the continuous nature of , pixel values ​​vary from 0 to 1. Regions with non-zero values ​​are considered lesion-specific regions, and the remaining regions are considered lesion-independent regions.

[0174] The goal of the feature attention module is to transform a given image A0 into A ′ (G(A0,c)→A ′ ), this module modifies the lesion area, and the rest of the image remains unchanged. To train the feature attention module, a reconstruction loss function is used.

[0175]

[0176] ‖IG(A ′ ,c)‖1 and ‖IG(A,c)‖1 represent the consistency loss and reconstruction loss respectively. The consistency loss ensures that when the image A generated using the conditional vector c ′ The translation back is the same as the original image A, while the reconstruction loss ensures that the input image remains unchanged when it is modified.

[0177] Optionally, the lesion attention module and the feature attention module are based on a U-shaped structure based on an attention mechanism, as follows:

[0178] The U-net architecture consists of an encoder and a decoder. The encoder has n stages for deep feature extraction, and the size of n is determined based on the input medical image. Each stage consists of a downsampling layer. The image features are first converted into two feature spaces f, g to calculate the attention, where f(x) = W f x and g(x) = W g x. Where C is the number of channels and N is the number of feature positions from the previous hidden layer. According to the U-Net network structure, the decoder is symmetrical with the encoder, which also contains n stages. Each stage of the decoder also consists of two multi-head self-attention modules and an upsampling layer. The upsampling layer includes bilinear interpolation and convolution layers. In order to alleviate the information loss caused by encoder downsampling, jump connections are used for feature fusion between the encoder and decoder.

[0179] Optionally, the discriminant module D is designed based on the Darknet model, as follows:

[0180] Darknet layers include convolution, normalization, and leaky ReLU as activation functions. The goal of the discriminator D is to distinguish the generated fake images from the real ones. As a discriminator, in order to successfully distinguish the fake images from the real ones, a regularized adversarial loss term L is applied. adv .

[0181]

[0182] Among them, D src Represents the loss between the real image and the generated image. src (A) and D src (A ′ ) represent the probability that image A is a real image and the probability that image A is generated. ′ The probability that the image is fake. If there is a significant overlap between the distributions of two images from different domains, they may look similar. To prevent this, we introduce a regularization term (the third term) to maintain the separability between the distributions when training the generator G. When training the generator, the regularization term is minimized to maintain the difference between the source distribution and the target distribution. Conversely, the function is maximized to ensure more overlap between the two distributions in the initial stage of training the discriminator. In addition, in order to make the generated images more semantically similar to the target images, we minimize the feature loss, which is implemented by the pre-trained VGG-19 network. The output image A obtained from the generator G ′ The real image A is input into the pre-trained VGG network for feature extraction. The loss error is then back-propagated to update the weights of G. The formula is as follows:

[0183]

[0184] Therefore, the objective loss function of the discriminator is as follows:

[0185] D=-L adv +L VGG

[0186] Optionally, the regression analysis process in step 5 is as follows:

[0187] According to the lesion images of n time nodes obtained in step 4, the images and corresponding times are sent to the Resnet network for regression, which is equivalent to constructing the functions y1=f(x1,t) and y2=f(x2,t), where x1 represents the image of the initial modality A input, x2 represents the image of the initial modality B input, y1 represents the lesion image of modality A at time t, and y2 represents the lesion image of modality B at time t. At the same time, the lesion image of modality A or modality B at a certain time can be substituted into the model, and the time t can be inferred inversely, and the evolution image of the other modality at this time can be obtained. That is, by inputting the image of modality A or modality B at a certain time, the lesion images of modality A and modality B at any time t can be accurately predicted.

[0188] The second aspect of the present invention provides a lesion time evolution device based on multimodal medical images, comprising: a processor, and a memory and a network interface connected to the processor; the network interface is connected to a non-volatile memory in a server; when the processor is running, it calls a computer program from the non-volatile memory through the network interface, and runs the computer program through the memory to execute the lesion time evolution method based on multimodal medical images described in the present invention.

[0189] The third aspect of the present invention provides a storage medium for the time evolution of lesions based on multimodal medical images, wherein the storage medium for the time evolution of lesions based on multimodal medical images is burned with a computer program, and when the computer program is run in the memory of a server, it implements the method for the time evolution of lesions based on multimodal medical images described in the present invention.

[0190] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0191] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for temporal evolution of lesions based on multimodal medical images, characterized in that: include: Step S1: The animal multimodal image acquisition module acquires the associated modality A and modality B medical image data of the animal at different time points under a preset environment; Step S2: The preprocessing module preprocesses the image data to obtain a training set and a test set after preprocessing, so that the spatial and grayscale proximity of images of different periods and different modalities meets the preset standards; Step S3: The conditional input module processes the modality A and modality B medical images at the n time nodes to obtain two sets of lesion abnormality images, calculates the lesion areas respectively and normalizes them to obtain conditional input; Step S4: The lesion time evolution module inputs the data of modality A and modality B into the lesion image generation model based on the multi-attention generative adversarial network, adjusts the parameters and saves them, and tests the quality of the generated image with the test set; Step S5: Result visualization and modeling analysis, visualizing the results of the time evolution of the lesion and performing regression modeling analysis; Step S6: Mapping of animal model and human model: Match the time lines of the lesion features of a certain modality of human image with the lesion features of animal image, explore the correspondence between human and animal based on examples, and obtain the evolution law of multimodal lesion images of human under preset environment; In step S4: For modality A, input A0 and A1 into the lesion image generation model based on the multi-attention generative adversarial network, adjust the parameters and save them, and use the test set to test the quality of the image generated at t1; input A1 and A2 into the network, adjust the parameters and save them, and test the results, and so on, until A n-1 With A n Input the network, adjust the parameters, save them, and test the results; the processing of mode B is the same as that of mode A; The lesion image generation model based on multi-attention generative adversarial network, the specific process is as follows: For mode A, enter the resulting condition into and A0 as input, and A1 as the target image, and input the network; the network includes a generator G based on a multi-attention mechanism, a discriminator D based on Darknet, and a feature loss calculation module; By synchronously training the generator and discriminator in the generative adversarial network, the target loss function is reduced until the error between the output image and the target image is less than the preset value, and the model from A0 to A1 is obtained and saved; similarly, A is obtained n-1 To A n and B n-1 To B n Model.

2. The method for temporal evolution of lesions based on multimodal medical images according to claim 1, characterized in that: In step S2: The preprocessing steps include image alignment, denoising and pairing, image standardization, and partitioning into training, validation and test sets in proportion; All images are standardized using the following formula: Among them, μ is the mean of the image, x represents the image matrix, σ and N represent the standard deviation and the number of image voxels respectively.

3. The method for temporal evolution of lesions based on multimodal medical images according to claim 1, characterized in that: In step S3: The modality A and modality B medical images at time node n are A n and B n , A n With A n-1 Subtract, B n With B n-1 Subtract them to get two sets of lesion abnormality maps. The lesion areas of the two sets of lesion abnormality maps are calculated and normalized to get the conditional input and As the input of the n+1 time node modal B lesion time evolution module, As the input of the time evolution module of the modal A lesion at the n+1 time node; Conditional input module, the specific process is as follows: For mode A, n With A n-1 Subtract them to get the abnormal lesion map; Use OpenCV to convert the lesion abnormality image into a grayscale image; Perform a binarization operation on the grayscale image to convert the target lesion in the image into a binary image; Perform contour detection on the binarized image and use the findContours function of OpenCV to obtain the contour information of the lesion; According to the contour information, the contourArea function of OpenCV is used to calculate the area of ​​the lesion. The lesion area is calculated and normalized to obtain the conditional input The processing of mode B is the same as that of mode A.

4. The method for temporal evolution of lesions based on multimodal medical images according to claim 1, characterized in that: The generator G includes: a lesion attention module and a feature attention module, as follows: The lesion attention module generates an attention map of the lesion, detects the lesion area, and keeps other parts unchanged; the lesion attention module inputs A0 and A1, and uses the attention localization loss function AL to minimize the difference between the generated lesion mask and the true mask: Among them, A is the original input image, A ′ is the generated image, A M is the true lesion mask, balanced by parameters λ1 and λ2, is the maximum likelihood estimate; AL is back-propagated to the lesion attention module until the difference between the generated nodule mask and the true nodule mask is minimized; the lesion attention module is F s (·), the attention map uses m = F s (A0) is generated, where m∈[1,0]; the regions with non-zero values ​​are lesion-specific regions, and the rest are lesion-independent regions; The feature attention module transforms image A0 into A ′ (G(A0,c)→A ′ ), G(·,c) is the generator output function, the feature attention module modifies the lesion area, and the rest of the image remains unchanged. In order to train the feature attention module, the reconstruction loss function RL is used: ‖AG(A ′ ,c)‖1 and ‖AG(A,c)‖1 represent the consistency loss and reconstruction loss respectively. For maximum likelihood estimation, the consistency loss ensures that when the image A generated using the conditional vector c ′ The translation back is the same as the original image A, while the reconstruction loss ensures that the input image remains unchanged when it is modified.

5. The method for temporal evolution of lesions based on multimodal medical images according to claim 4, characterized in that: The lesion attention module and feature attention module are U-shaped structures based on the attention mechanism, as follows: The U-net architecture consists of an encoder and a decoder: the encoder has n stages for deep feature extraction, and the size of n is determined according to the input medical image; each stage consists of a downsampling layer; the previous layer in the encoder The image features are first converted into two feature spaces f and g to calculate the attention. It is a C×N dimensional vector space, where C is the number of channels and N is the number of feature positions from the previous hidden layer; f(x)=W f x,g(x)=W g x, where W f is the weight of f, W g is the weight of g. According to the U-Net network structure, the decoder is symmetrical with the encoder, which also contains n stages. Each stage of the decoder consists of two multi-head self-attention modules and an upsampling layer. The upsampling layer includes bilinear interpolation and convolution layers. Skip connections are used for feature fusion between the encoder and decoder. The self-attention module consists of a block multi-head self-attention plus normalization layer and a feed-forward network plus normalization layer, as follows: F′=MSA(Norm(F in ))+F in F o =FFN(Norm(F′))+F′ Among them, F in represents the input feature map of the self-attention module, Norm(·) represents the normalization layer, F′ and F o They represent the output features of block MSA and FFN respectively, MSA is a multi-head self-attention module, and FFN is a feed-forward network; The specific process of block multi-head self-attention is as follows: The input feature map X in Divided into L×L non-overlapping blocks, the i-th block is where i∈{1,2,…,N}, is flattened and transposed to X i , for X i Apply MSA, first X i Linear mapping to query:Q i , key value: K i And value item value:V i ; Q i =X i W Q K i =X i W K V i =X i W V Among them, W Q , W K and W V are learnable parameters, representing the projection matrices of query, key value, and value item; Q i , K i and V i Divide into k groups along the channel dimension: For the jth group of self-attention, it is expressed as: in, and are the query, key, and value items of the jth group, respectively, and d k For each group dimension, the output tokens of the i-th block are: It is expressed as: Among them, Concat(·) represents concatenation, B represents position embedding, and W O It is possible to learn parameters and merge the outputs of all blocks to obtain the final output feature map X out .

6. The method for temporal evolution of lesions based on multimodal medical images according to claim 1, characterized in that: The discriminant module D is designed based on the Darknet model, as follows: Darknet layers include convolution, normalization, and leaky ReLU as activation functions. The discriminator D distinguishes the generated fake images from the real ones and applies the regularized adversarial loss term L adv : Where KL(·) is the Kullback-Leibler divergence, D src is the loss between the real image and the generated image, D src (A) and D src (A ′ ) represent the probability that image A is a real image and the probability that image A is generated. ′ The probability of a fake image is given by introducing a regularization term to maintain the separability between distributions when training the generator G. When training the generator, the regularization term is minimized to maintain the difference between the source distribution and the target distribution. In the initial stage of training the discriminator, the regularization term is maximized to ensure that there is more overlap between the two distributions. Minimizing the feature loss is achieved by the pre-trained VGG-19 network. The output image A obtained from the generator G ′ The real image A is input into the pre-trained VGG network for feature extraction, and the loss error is back-propagated to update the weight of G. The formula is as follows: Among them, R(A) is the size of the feature space, VGG(·) is the VGG feature extraction network; The objective loss function D of the discriminator is as follows: D-L adv +L VGG .

7. The method for temporal evolution of lesions based on multimodal medical images according to claim 1, characterized in that: In step S5: The regression analysis process is as follows: According to the lesion images at n time nodes, the images and corresponding times are sent to the Resnet network for regression to construct y1=f y (x1, t) and y2 = f y (x2, t), x1 represents the input image of the initial modality A, x2 represents the input image of the initial modality B, y1 represents the lesion image of modality A at time t, and y2 represents the lesion image of modality B at time t; according to the lesion image of modality A or modality B at a certain moment, substitute it into the model, infer the time t, and obtain the evolution image of the other modality at this time, input the image of modality A or modality B at a certain moment, and predict the lesion images of modality A and modality B at the corresponding moment.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for time evolution of lesions based on multimodal medical images described in any one of claims 1 to 7 are implemented.

9. A device for temporal evolution of lesions based on multimodal medical images, characterized in that: include: A processor, and memory and network interfaces connected to the processor; The network interface is connected to a non-volatile memory in the server; When running, the processor retrieves the computer program from the non-volatile memory through the network interface, and runs the computer program through the memory to execute the steps of the lesion time evolution method based on multimodal medical images described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal medical image processing

    CN109844808A

  • Apparatus, method, and storage medium for predicting evolution of fundus disease

    CN115376698A

  • Multi-modal magnetic resonance image generation method, system and device based on generative adversarial network and medium

    CN117710754A