A medical image generation method, system, electronic device and storage medium

By constructing a medical image generation model including 3D encoder and 3D decoder, the problem of low image conversion quality from sMRI to PET in the prior art is solved, efficient and accurate image conversion is achieved, and the accuracy of generated images is improved.

CN119919527BActive Publication Date: 2025-06-24ZHIXING TECHNOLOGY (CHANGSHA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510415174.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-24
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately realize high-quality conversion from structural magnetic resonance imaging (sMRI) to positron emission tomography (PET) images, and cannot meet the needs of clinical diagnosis.

Method used

A medical image generation model is built that includes a 3D encoder and a 3D decoder. The 3D encoder includes a depth-separable convolution and a multiple feature extraction modules containing MVM blocks and downsampling. The 3D decoder includes multiple residual blocks, multiple upsampling and prediction heads. The model is trained through the training data set to generate high-quality PET images.

Benefits of technology

By aggregating and extracting information at different scales, the model can better learn image features, realize high-quality conversion from sMRI images to PET images, and improve the accuracy of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919527B_ABST
    Figure CN119919527B_ABST
Patent Text Reader

Abstract

The present application discloses a medical image generation method, system, electronic device and storage medium. The method includes obtaining a training data set constructed from paired sMRI images and PET images; constructing a medical image generation model including a 3D encoder and a 3D decoder, wherein the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules each including an MVM block and downsampling, the 3D decoder includes multiple residual blocks, multiple upsamplings and a prediction head, and the MVM block is used to aggregate the input information of the MVM block and extract features of different scales; training the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image. The present application can efficiently and accurately achieve high-quality conversion from sMRI images to PET images, thereby improving the accuracy of the generated images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical image processing, and in particular, to a medical image generation method, system, electronic device, and storage medium. Background Art

[0002] In clinical practice, positron emission tomography (PET) imaging is of great significance for the early diagnosis of various diseases, such as Alzheimer's disease and cardiovascular diseases. However, PET scans have many limitations, including high cost and the risk of radiation exposure, which makes it impossible for many medical centers, especially in underdeveloped areas, to widely provide PET scan services. In contrast, structural magnetic resonance imaging (sMRI) is less costly and more widespread. Therefore, generating PET images from sMRI data has become a research direction to solve the problem of insufficient PET data.

[0003] Currently, there have been various methods attempting to achieve the conversion from sMRI to PET images. Although early traditional 3D convolutional neural networks (3D CNNs) can perform modality conversion, the quality of the generated PET images is low. Subsequently, methods such as the 3D U-Net network have used skip connections to improve the local details and texture information of the synthesized PET images, but there are still deficiencies in the balance between local details and global contours of the images. The 3D-Cycle GAN model preserves the basic structure of the image through cycle consistency, but due to the large domain gap between the sMRI and PET modalities, the actual conversion effect is still not ideal.

[0004] These existing methods above are all difficult to achieve high-quality conversion from sMRI to PET images efficiently and accurately, and cannot meet the needs of clinical diagnosis. Summary of the Invention

[0005] This application aims to propose a medical image generation method, system, electronic device, and storage medium, which can efficiently and accurately achieve high-quality conversion from sMRI images to PET images, thereby improving the accuracy of the generated images.

[0006] In a first aspect, an embodiment of this application provides a medical image generation method, and the method includes:

[0007] Obtain the sMRI image to be generated, and obtain a training data set constructed by paired sMRI images and PET images;

[0008] Construct a medical image generation model including a 3D encoder and a 3D decoder. Among them, the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules each containing an MVM block and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head. The MVM block is used to aggregate the input information of the MVM block and extract features of different scales;

[0009] Use the training dataset to train the medical image generation model to obtain a trained medical image generation model, so as to input the to-be-generated sMRI image into the trained medical image generation model to generate a PET image.

[0010] Compared with the prior art, the first aspect of this application has the following beneficial effects:

[0011] This method obtains the to-be-generated sMRI image and a training dataset constructed from paired sMRI images and PET images; constructs a medical image generation model including a 3D encoder and a 3D decoder. Among them, the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules each containing an MVM block and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head. The MVM block is used to aggregate the input information of the MVM block and extract features of different scales; uses the training dataset to train the medical image generation model to obtain a trained medical image generation model, so as to input the to-be-generated sMRI image into the trained medical image generation model to generate a PET image. In this way, by aggregating and extracting information of different scales, the constructed medical image generation model can better learn image features, can efficiently and accurately achieve high-quality conversion from sMRI images to PET images, and thus can improve the accuracy of the generated images.

[0012] In some embodiments, the step of inputting the to-be-generated sMRI image into the trained medical image generation model to generate a PET image includes:

[0013] Process the to-be-generated sMRI image through the depthwise separable convolution to obtain a depthwise separable convolution result;

[0014] Process the depthwise separable convolution result through the multiple feature extraction modules each containing an MVM block and downsampling to obtain the output result of each feature extraction module;

[0015] Process the output result of each feature extraction module through a residual block to obtain the output result of the residual block corresponding to each feature extraction module;

[0016] Upsample the output result of the residual block corresponding to the last feature extraction module to obtain a first upsampling result; and add and average the first upsampling result with the output result of the residual block corresponding to the previous feature extraction module to obtain a first calculation result;

[0017] Upsample the first calculation result to obtain a second upsampling result; and add and average the second upsampling result with the output result of the residual block corresponding to the previous feature extraction module adjacent to the current feature extraction module to obtain a second calculation result, and so on, until the output result of the residual block corresponding to the first feature extraction module is added and averaged with the calculation result corresponding to the next adjacent feature extraction module to obtain the calculation result corresponding to the first feature extraction module;

[0018] Process the depthwise separable convolution result through a residual block to obtain a first convolutional residual result;

[0019] Upsample the calculation result corresponding to the first feature extraction module to obtain a first feature upsampling result;

[0020] Add and average the first feature upsampling result with the first convolutional residual result to obtain a first average result;

[0021] Upsample the first average result to obtain a second feature upsampling result;

[0022] Process the sMRI image to be generated through multiple residual blocks to obtain a second convolutional residual result;

[0023] Add and average the second feature upsampling result with the second convolutional residual result to obtain a second average result;

[0024] Input the second average result into the prediction head to generate a PET image.

[0025] In some embodiments, the process of passing the depthwise separable convolution result through the multiple feature extraction modules including MVM blocks and downsampling to obtain the output result of each feature extraction module includes:

[0026] Input the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain an output result of the first MVM block;

[0027] Downsample the output result of the first MVM block to obtain the output result of the first feature extraction module;

[0028] Use the output result of the first feature extraction module as the input to the MVM block in the next feature extraction module, and after all feature extraction modules are processed, obtain the output result of each feature extraction module.

[0029] In some embodiments, the MVM block includes an MVM sub-block and an ASC sub-block. The MVM sub-block is used to extract sequence features in the forward domain, reverse domain, and hybrid domain of the input MVM sub-block features. The ASC sub-block is used to aggregate the feature information of the input ASC sub-block. Inputting the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain the first MVM block output result includes:

[0030] Processing the depthwise separable convolution result through the ASC sub-block to obtain the ASC sub-block output result;

[0031] Processing the ASC sub-block output result through the first normalization to obtain the first normalization result;

[0032] Processing the first normalization result through the MVM sub-block to obtain the MVM sub-block output result;

[0033] Adding the ASC sub-block output result and the MVM sub-block output result to obtain the first addition result;

[0034] Processing the first addition result through the second normalization to obtain the second normalization result;

[0035] Processing the second normalization result through a multi-layer perceptron to obtain the multi-layer perceptron output result;

[0036] Adding the multi-layer perceptron output result and the first addition result to obtain the first MVM block output result.

[0037] In some embodiments, the processing the depthwise separable convolution result through the ASC sub-block to obtain the ASC sub-block output result includes:

[0038] Processing the depthwise separable convolution result through the first normalization, the first convolutional layer, and the first activation function to obtain the first processing result;

[0039] Processing the depthwise separable convolution result through the second normalization, the second convolutional layer, and the second activation function to obtain the second processing result;

[0040] Multiplying the first processing result and the second processing result to obtain the multiplication result;

[0041] Processing the multiplication result through the third normalization, the third convolutional layer, and the third activation function to obtain the third processing result;

[0042] Adding the third processing result and the depthwise separable convolution result to obtain the ASC sub-block output result.

[0043] In some embodiments, training the medical image generation model using the training data set to obtain a trained medical image generation model includes:

[0044] Constructing a perceptual loss function for the medical image generation model;

[0045] Based on the perceptual loss function, training the medical image generation model using the training data set to obtain a trained medical image generation model.

[0046] In some embodiments, constructing the perceptual loss function for the medical image generation model includes:

[0047] ;

[0048] where represents the perceptual loss function, represents the number of channels of the feature map, represents the height of the feature map, represents the width of the feature map, represents the feature extraction function of the th layer in the pre-trained 3D-ResNet10 network, represents the generated PET image, represents the real PET image, represents the element-wise squared error.

[0049] In a second aspect, an embodiment of the present application further provides a medical image generation system, and the system includes:

[0050] A data acquisition unit, configured to acquire the sMRI image to be generated, and acquire a training data set constructed by paired sMRI images and PET images;

[0051] A model construction unit, configured to construct a medical image generation model including a 3D encoder and a 3D decoder, where the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules including MVM blocks and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head, and the MVM block is used to aggregate the input information of the MVM block and extract features of different scales;

[0052] An image generation unit, configured to train the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image.

[0053] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and when the instructions are executed by the at least one control processor, the at least one control processor is enabled to execute a medical image generation method as described above.

[0054] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a medical image generation method as described above.

[0055] It can be understood that the beneficial effects of the above second to fourth aspects compared with the related art are the same as those of the above first aspect compared with the related art. For relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0057] Figure 1 is a schematic flowchart of an embodiment of the medical image generation method provided by the present application;

[0058] Figure 2 is a schematic diagram of the overall structure of the GenMamba model in the best embodiment of the medical image generation method provided by the present application;

[0059] Figure 3 is a schematic diagram of the structure of the MVM block in the best embodiment of the medical image generation method provided by the present application;

[0060] Figure 4 is a schematic diagram of the MVM sub-block extracting sequence features in the best embodiment of the medical image generation method provided by the present application;

[0061] Figure 5 is a schematic diagram of the structure of the ASC sub-block in the best embodiment of the medical image generation method provided by the present application;

[0062] Figure 6 is a schematic diagram of the structure of an embodiment of the medical image generation system provided by the present application;

[0063] Figure 7 is a schematic diagram of the structure of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0065] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0066] It should also be understood that in the description of the embodiments of the present application, referring to "one embodiment" or "some embodiments" etc. means that in one or more embodiments of the embodiments of the present application, specific features, structures or characteristics described in connection with that embodiment are included. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0067] In the description of the present application, it should be understood that with regard to the orientation description, such as up, down, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application.

[0068] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0069] In clinical practice, positron emission tomography (PET) imaging is of great significance for the early diagnosis of various diseases, such as Alzheimer's disease and cardiovascular diseases. However, PET scans have many limitations, including high cost and the risk of radiation exposure, which makes it impossible for many medical centers, especially in underdeveloped regions, to widely provide PET scan services. In contrast, structural magnetic resonance imaging (sMRI) is less costly and more widespread. Therefore, generating PET images from sMRI data has become a research direction to solve the problem of insufficient PET data.

[0070] Currently, there have been various methods attempting to achieve the conversion from sMRI to PET images. Although early traditional 3D convolutional neural networks (3D CNNs) were able to perform modality conversion, the quality of the generated PET images was low. Subsequently, networks such as 3D U-Net emerged, which utilized skip connections to improve the local details and texture information of the synthesized PET images, but there were still deficiencies in the balance between local details and global contours of the images. The 3D-Cycle GAN model preserved the basic structure of the images through cycle consistency, but due to the large domain gap between the sMRI and PET modalities, the actual conversion effect was still not ideal.

[0071] These existing methods mentioned above are all difficult to achieve high-quality conversion from sMRI to PET images efficiently and accurately, and cannot meet the needs of clinical diagnosis.

[0072] To solve the problem that the existing methods mentioned above are difficult to achieve high-quality conversion from sMRI to PET images efficiently and accurately, this application proposes a medical image generation method, system, electronic device, and storage medium.

[0073] Referring to Figure 1 , an embodiment of this application provides a medical image generation method, which includes the following steps:

[0074] Step S100: Obtain the sMRI image to be generated, and obtain the training data set constructed from the paired sMRI images and PET images;

[0075] Step S200: Construct a medical image generation model including a 3D encoder and a 3D decoder. Among them, the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules containing MVM blocks and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head. The MVM block is used to aggregate the input information of the MVM block and extract features of different scales;

[0076] Step S300: Train the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image.

[0077] In this embodiment, by obtaining the sMRI image to be generated and the training data set constructed from the paired sMRI images and PET images; constructing a medical image generation model including a 3D encoder and a 3D decoder, where the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules each containing an MVM block and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head, and the MVM block is used to aggregate the input information of the MVM block and extract features of different scales; training the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image. In this way, by aggregating and extracting information of different scales, the constructed medical image generation model can better learn image features, and can efficiently and accurately achieve high-quality conversion from sMRI images to PET images, thereby improving the accuracy of the generated images.

[0078] The above-mentioned training of the medical image generation model using the training data set to obtain a trained medical image generation model can be to set appropriate hyperparameters, such as the learning rate and the number of iterations, etc., during the model training process. Use the training data set to train the model; during the training process, the medical image generation model continuously adjusts the model parameters through backpropagation to optimize the model performance.

[0079] The above-mentioned inputting the sMRI image to be generated into the trained medical image generation model to generate a PET image can be to input the sMRI image to be generated into the trained medical image generation model, extract the feature information of the sMRI image to be generated through the 3D encoder, and then gradually restore the details and structure of the image through the 3D decoder, and convert the features extracted by the 3D encoder into image information in the PET modality.

[0080] In some embodiments, inputting the sMRI image to be generated into the trained medical image generation model to generate a PET image includes:

[0081] Performing depthwise separable convolution processing on the sMRI image to be generated to obtain a depthwise separable convolution result;

[0082] Passing the depthwise separable convolution result through multiple feature extraction modules each containing an MVM block and downsampling to obtain the output result of each feature extraction module;

[0083] Performing residual block processing on the output result of each feature extraction module to obtain the output result of the residual block corresponding to each feature extraction module;

[0084] Upsample the output result of the residual block corresponding to the last feature extraction module to obtain the first upsampling result; and add the first upsampling result to the output result of the residual block corresponding to the previous feature extraction module and take the average to obtain the first calculation result;

[0085] Upsample the first calculation result to obtain the second upsampling result; and add the second upsampling result to the output result of the residual block corresponding to the previous feature extraction module adjacent to the current feature extraction module and take the average to obtain the second calculation result, and so on, until the output result of the residual block corresponding to the first feature extraction module is added to the calculation result corresponding to the next adjacent feature extraction module and take the average to obtain the calculation result corresponding to the first feature extraction module;

[0086] Process the depthwise separable convolution result through a residual block to obtain the first convolutional residual result;

[0087] Upsample the calculation result corresponding to the first feature extraction module to obtain the first feature upsampling result;

[0088] Add the first feature upsampling result to the first convolutional residual result and take the average to obtain the first average result;

[0089] Upsample the first average result to obtain the second feature upsampling result;

[0090] Process the sMRI image to be generated through multiple residual blocks to obtain the second convolutional residual result;

[0091] Add the second feature upsampling result and the second convolutional residual result and take the average to obtain the second average result;

[0092] Input the second average result into the prediction head to generate the PET image.

[0093] In this embodiment, the output result of the residual block corresponding to the last feature extraction module is upsampled to obtain a first upsampling result; the first upsampling result is added to the output result of the residual block corresponding to the previous feature extraction module and averaged to obtain a first calculation result; the first calculation result is upsampled to obtain a second upsampling result; the second upsampling result is added to the output result of the residual block corresponding to the previous feature extraction module adjacent to the current feature extraction module and averaged to obtain a second calculation result, and so on, until the output result of the residual block corresponding to the first feature extraction module is added to the calculation result corresponding to the next adjacent feature extraction module and averaged to obtain the calculation result corresponding to the first feature extraction module; the depthwise separable convolution result is processed through a residual block to obtain a first convolutional residual result; the calculation result corresponding to the first feature extraction module is upsampled to obtain a first feature upsampling result; the first feature upsampling result is added to the first convolutional residual result and averaged to obtain a first average result; the first average result is upsampled to obtain a second feature upsampling result; the sMRI image to be generated is processed through multiple residual blocks to obtain a second convolutional residual result; the second feature upsampling result and the second convolutional residual result are added and averaged to obtain a second average result; the second average result is input into a prediction head to generate a PET image. In this way, by using multiple feature extraction modules including MVM blocks and downsampling to extract feature information at multiple scales, and then combining the output results of the residual blocks and the upsampling results, the details and structure of the image are gradually restored, so that high-quality conversion from the sMRI image to the PET image can be efficiently and accurately achieved.

[0094] In some embodiments, the depthwise separable convolution result is passed through multiple feature extraction modules including MVM blocks and downsampling to obtain the output result of each feature extraction module, including:

[0095] The depthwise separable convolution result is input into the MVM block in the first feature extraction module to obtain the output result of the first MVM block;

[0096] The output result of the first MVM block is downsampled to obtain the output result of the first feature extraction module;

[0097] The output result of the first feature extraction module is used as the input to the MVM block in the next feature extraction module, and after all feature extraction modules are processed, the output result of each feature extraction module is obtained.

[0098] In this embodiment, by inputting the depthwise separable convolution result into the MVM block in the first feature extraction module, the output result of the first MVM block is obtained; the output result of the first MVM block is downsampled to obtain the output result of the first feature extraction module; the output result of the first feature extraction module is used as the input of the MVM block in the next feature extraction module. After all feature extraction modules are processed, the output results of each feature extraction module are obtained. In this way, different-scale feature information is extracted by multiple feature extraction modules, which can lay a good data foundation for generating high-quality images in the later stage.

[0099] In some embodiments, the MVM block includes an MVM sub-block and an ASC sub-block. The MVM sub-block is used to extract the sequence features of the forward domain, reverse domain, and mixed domain of the input MVM sub-block features. The ASC sub-block is used to aggregate the feature information of the input ASC sub-block. Inputting the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain the output result of the first MVM block includes:

[0100] The depthwise separable convolution result is processed by the ASC sub-block to obtain the output result of the ASC sub-block;

[0101] The output result of the ASC sub-block is processed by the first layer of normalization to obtain the first layer of normalization result;

[0102] The first layer of normalization result is processed by the MVM sub-block to obtain the output result of the MVM sub-block;

[0103] The output result of the ASC sub-block and the output result of the MVM sub-block are added together to obtain the first addition result;

[0104] The first addition result is processed by the second layer of normalization to obtain the second layer of normalization result;

[0105] The second layer of normalization result is processed by a multi-layer perceptron to obtain the output result of the multi-layer perceptron;

[0106] The output result of the multi-layer perceptron and the first addition result are added together to obtain the output result of the first MVM block.

[0107] In this embodiment, the depthwise separable convolution result is processed through the ASC sub-block to obtain the ASC sub-block output result; the ASC sub-block output result is processed through the first normalization to obtain the first normalization result; the first normalization result is processed through the MVM sub-block to obtain the MVM sub-block output result; the ASC sub-block output result and the MVM sub-block output result are added together to obtain the first addition result; the first addition result is processed through the second normalization to obtain the second normalization result; the second normalization result is processed through a multi-layer perceptron to obtain the multi-layer perceptron output result; the multi-layer perceptron output result and the first addition result are added together to obtain the first MVM block output result. In this way, the sequence features of the forward domain, reverse domain, and mixed domain of the input MVM sub-block features are extracted through the MVM sub-block, and the feature information of the input ASC sub-block is aggregated through the ASC sub-block, which can enhance the feature extraction ability and information fusion ability of the medical image generation model. And by extracting the feature information of each domain from different spatial perspectives, the learning ability of the medical image generation model for 3D information is enhanced.

[0108] In some embodiments, processing the depthwise separable convolution result through the ASC sub-block to obtain the ASC sub-block output result includes:

[0109] The depthwise separable convolution result is processed through the first normalization, the first convolutional layer, and the first activation function to obtain the first processing result;

[0110] The depthwise separable convolution result is processed through the second normalization, the second convolutional layer, and the second activation function to obtain the second processing result;

[0111] The first processing result and the second processing result are multiplied to obtain the multiplication result;

[0112] The multiplication result is processed through the third normalization, the third convolutional layer, and the third activation function to obtain the third processing result;

[0113] The third processing result is added to the depthwise separable convolution result to obtain the ASC sub-block output result.

[0114] In this embodiment, aggregating the feature information of the input ASC sub-block through the ASC sub-block can enhance the information fusion ability of the medical image generation model and lay a good data foundation for generating high-quality images later.

[0115] In some embodiments, training the medical image generation model with a training data set to obtain a trained medical image generation model includes:

[0116] Constructing the perceptual loss function of the medical image generation model;

[0117] Based on the perceptual loss function, the medical image generation model is trained using the training data set to obtain a trained medical image generation model.

[0118] In this embodiment, the perceptual loss function of the medical image generation model is constructed; based on the perceptual loss function, the medical image generation model is trained using the training data set to obtain a trained medical image generation model. In this way, the generated PET image guided by the perceptual loss function can focus on details and improve the image quality.

[0119] In some embodiments, constructing the perceptual loss function of the medical image generation model includes:

[0120] ;

[0121] where, represents the perceptual loss function, represents the number of channels of the feature map, represents the height of the feature map, represents the width of the feature map, represents the feature extraction function of the th layer in the pre-trained 3D-ResNet10 network, represents the generated PET image, represents the real PET image, represents the element-wise squared error.

[0122] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments:

[0123] In clinical practice, positron emission tomography (PET) imaging is of great significance for the early diagnosis of various diseases, such as Alzheimer's disease and cardiovascular diseases. However, PET scans have many limitations, including high cost and the risk of radiation exposure, which makes it impossible for many medical centers, especially in underdeveloped regions, to widely provide PET scan services. In contrast, structural magnetic resonance imaging (sMRI) is less costly and more widespread. Therefore, generating PET images from sMRI data has become a research direction to solve the problem of insufficient PET data.

[0124] Currently, there have been various methods attempting to achieve the conversion from sMRI to PET. Although early traditional 3D convolutional neural networks (3D CNNs) were able to perform modality conversion, the quality of the generated PET images was low. Subsequently, networks such as 3D U-Net emerged, which utilized skip connections to improve the local details and texture information of the synthesized PET images, but there were still deficiencies in the balance between local details and global contours of the images. The 3D-Cycle GAN model retained the basic structure of the images through cycle consistency, but due to the large domain gap between the sMRI and PET modalities, the actual conversion effect was still not ideal. These methods are all difficult to achieve high-quality conversion from sMRI to PET efficiently and accurately, and cannot meet the needs of clinical diagnosis.

[0125] To solve the problems existing in the conversion from sMRI to PET in the prior art, this embodiment proposes a deep learning method for generating the corresponding PET modality from the sMRI modality based on the GenMamba model (i.e., a medical image generation model), aiming to improve the quality of the generated PET images, enhance the feature extraction ability and information fusion ability of the medical image generation model, and provide more reliable imaging support for clinical diagnosis.

[0126] To achieve the above object, the technical solution of this embodiment is as follows:

[0127] The deep learning method for generating the corresponding PET modality from the sMRI modality based on the GenMamba model in this embodiment includes the following steps:

[0128] Step S1: Pair the sMRI images of the subject with the corresponding PET images. Collect imaging data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database, where the sMRI images are T1-weighted MP-RAGE sequences, acquired from a 1.5T scanner, usually composed of individual voxels, and the voxel size is approximately , and the image resolution is ; process the PET images into a unified format.

[0129] Step S2: Construct a GenMamba model (i.e., a medical image generation model), which includes data preprocessing (i.e., processing through depthwise separable convolution) of the data input in Step S1. Normalize the sMRI images and map the pixel values to the interval to eliminate the influence of data scale differences on model training. Among them, the overall structural schematic diagram of the GenMamba model is as shown in Figure 2As shown in the figure, the GenMamba model includes a 3D encoder and a 3D decoder. The 3D encoder includes depthwise separable convolutions, multiple Multi-View Mamba (MVM) blocks, and multiple downsamplings. The number of MVM blocks is the same as the number of downsamplings. One MVM block and one downsampling are combined into a feature extraction module, and multiple feature extraction modules are obtained (in this embodiment, four feature extraction modules are constructed); the 3D decoder includes multiple residual blocks, multiple upsamplings (i.e., transposed convolutions), and a prediction head.

[0130] Step S3: Use the Multi-View Mamba (MVM) sub-block in the 3D encoder to perform multi-scale modeling on the preprocessed data, and aggregate slice information through the Aggregated Spatial Convolution (ASC) sub-block. Specifically:

[0131] Obtain the sMRI image (i.e., 3D input). After preprocessing the sMRI image through depthwise separable convolutions (Stem), obtain the depthwise separable convolution result; input the depthwise separable convolution result into the first feature extraction module to obtain the output result of the first feature extraction module (i.e., the extracted feature). Among them, the feature extraction module first processes the obtained input information (which can be the depthwise separable convolution result) through the MVM block to obtain the output result of the MVM block, and then downsamples the output result of the MVM block to obtain the output result of the feature extraction module.

[0132] Specifically, referring to Figure 3 , the MVM block includes an MVM sub-block and an ASC sub-block. The processing method of the MVM block is: process the input information obtained by the MVM block (i.e., the extracted image feature) through the ASC sub-block to obtain the output result of the ASC sub-block, process the output result of the ASC sub-block through the first layer of normalization (LayerNorm) to obtain the first layer of normalization result, process the first layer of normalization result through the MVM sub-block to obtain the output result of the MVM sub-block. Add the output result of the ASC sub-block and the output result of the MVM sub-block to obtain the first addition result. Process the first addition result through the second layer of normalization to obtain the second layer of normalization result, process the second layer of normalization result through a multi-layer perceptron (MLP) to obtain the output result of the multi-layer perceptron, and add the output result of the multi-layer perceptron and the first addition result to obtain the output result of the MVM block.

[0133] The MVM sub-block models each domain from three different spatial perspectives. The formula is:

[0134] ;

[0135] Among them, respectively represent the sequence features from the forward domain, reverse domain, and hybrid domain. Since the image features received by the Multi-View Mamba (MVM) sub-block have multiple feature information elements, the multiple feature information elements are arranged in a column in sequence to obtain a column of feature information elements. Refer to Figure 4 , the forward domain refers to extracting feature information elements in sequence starting from the front of a column of feature information elements to obtain the sequence features of the forward domain; the reverse domain refers to extracting feature information elements in sequence starting from the back of a column of feature information elements to obtain the sequence features of the reverse domain; the hybrid domain refers to extracting one from the front of a column of feature information elements and then one from the back of a column of feature information elements in such a cross-extraction manner to obtain the sequence features of the hybrid domain, so as to enhance the learning ability of 3D information.

[0136] Specifically, the processing method of the ASC sub-block is as follows: Refer to Figure 5 , the input information obtained by the ASC sub-block is passed through the first normalization (Norm), the first convolutional layer (3×3×3 Conv3D), and the first activation function (ReLU) to obtain the first processing result; the input information obtained by the ASC sub-block is passed through the second normalization (Norm), the second convolutional layer (1×1×1 Conv3D), and the second activation function (ReLU) to obtain the second processing result; the first processing result and the second processing result are multiplied to obtain the multiplication result, and the multiplication result is passed through the third normalization (Norm), the third convolutional layer (3×3×3 Conv3D), and the third activation function (ReLU) to obtain the third processing result; finally, the third processing result is added to the input information obtained by the ASC sub-block to obtain the output result of the ASC sub-block.

[0137] The ASC sub-block performs specific convolution operations and fusion operations, and the formula is:

[0138] ;

[0139] Among them, represents the input 3D feature information, represents the aggregation space module. The ASC sub-block can effectively aggregate the spatio-temporal features between slices.

[0140] Take the output result of the first feature extraction module as the input of the second feature extraction module, so that the second feature extraction module processes the input information to obtain the output result of the second feature extraction module , and so on until the output result of the last feature extraction module is obtained. In this way, multi-scale features can be extracted through the encoder .

[0141] Step S4: Decode using a 3D decoder. The 3D decoder includes multiple residual blocks, multiple upsamplings (i.e., transposed convolutions), and a prediction head (Generator Head, the network layer structure of the prediction head is well-known to those skilled in the art and will not be specifically described in this embodiment). Through convolutional operations (i.e., transposed convolutions) and combined with skip connections (i.e., residual blocks), PET modality information is predicted and generated. The decoder gradually restores the details and structure of the image, converting the features extracted by the encoder into image information in the PET modality.

[0142] The specific process of the 3D decoder part is as follows:

[0143] The output result of the last feature extraction module After being processed by a residual block, the first residual output result is obtained. The output result of the previous feature extraction module of the last feature extraction module After being processed by a residual block, the second residual output result is obtained. The first residual output result is upsampled to obtain the first upsampling result. The first upsampling result and the second residual output result are added and averaged to obtain the first calculation result;

[0144] The first calculation result is upsampled to obtain the second upsampling result. The output result of the previous feature extraction module After being processed by a residual block, the third residual output result is obtained. The third residual output result and the second upsampling result are added and averaged to obtain the second calculation result;

[0145] The second calculation result is upsampled to obtain the third upsampling result. The output result of the previous feature extraction module After being processed by a residual block, the fourth residual output result is obtained. The fourth residual output result and the third upsampling result are added and averaged to obtain the third calculation result (i.e., the calculation result corresponding to the first feature extraction module);

[0146] The third calculation result is upsampled to obtain the fourth upsampling result (i.e., the first feature upsampling result). The depthwise separable convolution result After being processed by a residual block, the fifth residual output result (i.e., the first convolutional residual result) is obtained. The fifth residual output result and the fourth upsampling result are added and averaged to obtain the fourth calculation result (i.e., the first average result);

[0147] The fourth calculation result is upsampled to obtain the fifth upsampling result (i.e., the second feature upsampling result). The sMRI image is processed by a residual block to obtain the sixth residual output result , the sixth residual output result is further processed through a residual block to obtain a seventh residual output result (i.e., the second convolutional residual result). The seventh residual output result and the fifth upsampling result are added and averaged to obtain a fifth calculation result (i.e., the second average result);

[0148] The feature information in the fifth calculation result is input into the prediction head for prediction, and a PET image is predicted, that is, a PET image is generated.

[0149] Step S5: During the generation process, the perceptual loss function is used to guide the generated PET image to focus on details and improve the image quality. Using the 3D-ResNet10 pre-trained on 23 medical data sets as the feature extractor, calculate the difference between the features of the real PET image and the synthetic image, and obtain the perceptual loss by averaging in the spatial dimension, so that the synthetic image is closer to the features of the real PET image. Among them, the calculation formula of the perceptual loss function is:

[0150] ;

[0151] Among them, represents the feature extraction function of the th layer in the pre-trained 3D-ResNet10 network; and represent the generated PET image and the real PET image respectively; represent the number of channels, height and width of the feature map respectively; represents the element-wise squared error, and finally the average is taken over all positions of the feature map. The pre-trained 3D-ResNet-10 is a model that has been trained on 23 medical data sets; the real PET image and the generated PET image are respectively input into this feature extractor, the features extracted from each layer are normalized on specific channels, the difference between the input features and the target features is calculated, and the perceptual loss is obtained by averaging in the spatial dimension.

[0152] Step S6: Verify the effectiveness of the MVM sub-block, and compare the SSIM, PSNR and MAE indicators under different Mamba domain modellings; verify the effectiveness of the ASC sub-block, and compare the changes in the SSIM, PSNR and MAE indicators before and after using the ASC sub-block; adjust the weight of the perceptual loss function , and compare the SSIM, PSNR and MAE indicators of the synthetic PET image under different values to determine the optimal value. Quantitatively evaluate the similarity and quality difference between the generated PET image and the real PET image through these indicators.

[0153] In this embodiment, through the innovative GenMamba model, combined with the MVM and ASC sub-blocks and the perceptual loss function, the quality and accuracy of generating PET modality images from sMRI modality are effectively improved, providing more valuable imaging data for clinical diagnosis.

[0154] Referring Figure 6 , the embodiment of the present application also provides a medical image generation system, which includes a data acquisition unit 100, a model construction unit 200, and an image generation unit 300, where:

[0155] The data acquisition unit 100 is configured to acquire the sMRI image to be generated, and acquire a training data set constructed from paired sMRI images and PET images;

[0156] The model construction unit 200 is configured to construct a medical image generation model including a 3D encoder and a 3D decoder. Among them, the 3D encoder includes depthwise separable convolutions and multiple feature extraction modules including MVM blocks and downsampling, and the 3D decoder includes multiple residual blocks, multiple upsamplings, and a prediction head. The MVM block is used to aggregate the input information of the MVM block and extract features at different scales;

[0157] The image generation unit 300 is configured to train the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image.

[0158] In some embodiments, the image generation unit 300 may specifically be configured to:

[0159] Perform depthwise separable convolution processing on the sMRI image to be generated to obtain a depthwise separable convolution result;

[0160] Pass the depthwise separable convolution result through multiple feature extraction modules including MVM blocks and downsampling to obtain the output result of each feature extraction module;

[0161] Perform residual block processing on the output result of each feature extraction module to obtain the output result of the residual block corresponding to each feature extraction module;

[0162] Upsample the output result of the residual block corresponding to the last feature extraction module to obtain a first upsampling result; and add and average the first upsampling result and the output result of the residual block corresponding to the previous feature extraction module to obtain a first calculation result;

[0163] Upsample the first calculation result to obtain a second upsampled result; add the second upsampled result to the output result of the residual block corresponding to the previous feature extraction module adjacent to the current feature extraction module and take the average to obtain a second calculation result, and so on, until the output result of the residual block corresponding to the first feature extraction module is added to the calculation result corresponding to the next adjacent feature extraction module and the average is taken to obtain the calculation result corresponding to the first feature extraction module;

[0164] Process the depthwise separable convolution result through a residual block to obtain a first convolutional residual result;

[0165] Upsample the calculation result corresponding to the first feature extraction module to obtain a first feature upsampled result;

[0166] Add the first feature upsampled result to the first convolutional residual result and take the average to obtain a first average result;

[0167] Upsample the first average result to obtain a second feature upsampled result;

[0168] Process the sMRI image to be generated through multiple residual blocks to obtain a second convolutional residual result;

[0169] Add the second feature upsampled result to the second convolutional residual result and take the average to obtain a second average result;

[0170] Input the second average result into the prediction head to generate a PET image.

[0171] In some embodiments, the image generation unit 300 may specifically be used for:

[0172] Input the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain an output result of the first MVM block;

[0173] Downsample the output result of the first MVM block to obtain an output result of the first feature extraction module;

[0174] Use the output result of the first feature extraction module as the input to the MVM block in the next feature extraction module. After all feature extraction modules are processed, obtain the output results of each feature extraction module.

[0175] In some embodiments, the image generation unit 300 may specifically be used for:

[0176] Process the depthwise separable convolution result through the ASC sub-block to obtain an output result of the ASC sub-block;

[0177] Process the output result of the ASC sub-block through the first layer of normalization to obtain a first layer of normalization result;

[0178] Process the first - layer normalization result through the MVM sub - block to obtain the output result of the MVM sub - block;

[0179] Add the output result of the ASC sub - block and the output result of the MVM sub - block to obtain the first addition result;

[0180] Process the first addition result through the second - layer normalization to obtain the second - layer normalization result;

[0181] Process the second - layer normalization result through a multi - layer perceptron to obtain the output result of the multi - layer perceptron;

[0182] Add the output result of the multi - layer perceptron and the first addition result to obtain the output result of the first MVM block.

[0183] In some embodiments, the image generation unit 300 may be specifically configured to:

[0184] Process the depth - wise separable convolution result through the first normalization, the first convolutional layer, and the first activation function to obtain the first processing result;

[0185] Process the depth - wise separable convolution result through the second normalization, the second convolutional layer, and the second activation function to obtain the second processing result;

[0186] Multiply the first processing result and the second processing result to obtain the multiplication result;

[0187] Process the multiplication result through the third normalization, the third convolutional layer, and the third activation function to obtain the third processing result;

[0188] Add the third processing result to the depth - wise separable convolution result to obtain the output result of the ASC sub - block.

[0189] In some embodiments, the image generation unit 300 may be specifically configured to:

[0190] Construct the perceptual loss function of the medical image generation model;

[0191] Based on the perceptual loss function, use the training data set to train the medical image generation model to obtain the trained medical image generation model.

[0192] In some embodiments, the image generation unit 300 may be specifically configured to:

[0193] ;

[0194] Wherein, represents the perceptual loss function, represents the number of channels of the feature map, represents the height of the feature map, represents the width of the feature map, represents the feature extraction function of the layer in the pre-trained 3D-ResNet10 network, represents the generated PET image, represents the real PET image, represents the element-wise squared error.

[0195] It should be noted that since a medical image generation system in this embodiment and the above-mentioned medical image generation method are based on the same inventive concept, the corresponding content in the method embodiment also applies to the system embodiment of this system and will not be elaborated here.

[0196] Referring to Figure 7 , an electronic device is further provided in an embodiment of the present application. This electronic device includes:

[0197] at least one memory;

[0198] at least one processor;

[0199] at least one program;

[0200] The program is stored in the memory, and the processor executes at least one program to implement the medical image generation method described above in the present disclosure.

[0201] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0202] The electronic device in the embodiment of the present application will be introduced in detail below.

[0203] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;

[0204] The memory 1700 can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the medical image generation method of the embodiments of the present disclosure.

[0205] The input / output interface 1800 is used to implement information input and output;

[0206] The communication interface 1900 is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.), or can also achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0207] The bus 2000 transmits information between the various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900);

[0208] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0209] The embodiments of the present disclosure also provide a storage medium. This storage medium is a computer-readable storage medium. This computer-readable storage medium stores computer-executable instructions, and these computer-executable instructions are used to cause a computer to execute the above-mentioned medical image generation method.

[0210] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0211] The embodiments described in the embodiments of the present disclosure are to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0212] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0213] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0214] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0215] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0216] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0217] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0218] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0220] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The embodiments of this application have been described in detail above with reference to the accompanying drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can also be made without departing from the purpose of this application.

[0221] The embodiments of this application have been described in detail above with reference to the accompanying drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can also be made without departing from the purpose of this application.

Claims

1. A medical image generation method, characterized in that: The method comprises: Acquire an sMRI image to be generated, and acquire a training data set constructed by paired sMRI images and PET images; Constructing a medical image generation model including a 3D encoder and a 3D decoder, wherein the 3D encoder includes a depthwise separable convolution and a plurality of feature extraction modules including an MVM block and downsampling, and the 3D decoder includes a plurality of residual blocks, a plurality of upsampling and prediction heads, and the MVM block is used to aggregate input information of the MVM block and extract features of different scales; The medical image generation model is trained using the training data set to obtain a trained medical image generation model, so that the sMRI image to be generated is input into the trained medical image generation model to generate a PET image, including: inputting the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain a first MVM block output result, wherein the depthwise separable convolution result is a result obtained by subjecting the sMRI image to be generated to the depthwise separable convolution processing, the MVM block includes an MVM sub-block and an ASC sub-block, the MVM sub-block is used to extract sequence features of the forward domain, the reverse domain and the mixed domain of the input MVM sub-block feature, and the ASC sub-block is used to aggregate feature information of the input ASC sub-block; Processing the depthwise separable convolution result through the ASC sub-block to obtain an ASC sub-block output result; The ASC sub-block output result is subjected to first-layer normalization processing to obtain a first-layer normalized result; Processing the first layer normalization result through the MVM sub-block to obtain an MVM sub-block output result; Adding the ASC sub-block output result and the MVM sub-block output result to obtain a first addition result; Subjecting the first addition result to a second-layer normalization process to obtain a second-layer normalization result; Processing the normalized result of the second layer through a multi-layer perceptron to obtain an output result of the multi-layer perceptron; Adding the multilayer perceptron output result and the first addition result to obtain a first MVM block output result; Downsampling the output result of the first MVM block to obtain an output result of the first feature extraction module; The output result of the first feature extraction module is used as the input of the MVM block in the next feature extraction module until all feature extraction modules are processed and the output result of each feature extraction module is obtained.

2. The medical image generation method according to claim 1, characterized in that: The step of inputting the sMRI image to be generated into the trained medical image generation model to generate a PET image comprises: Processing the sMRI image to be generated by the depthwise separable convolution to obtain a depthwise separable convolution result; Passing the depthwise separable convolution result through the plurality of feature extraction modules including MVM blocks and downsampling to obtain an output result of each feature extraction module; Processing the output results of each feature extraction module by a residual block to obtain a residual block output result corresponding to each feature extraction module; Upsampling the residual block output result corresponding to the last feature extraction module to obtain a first upsampling result; and adding the first upsampling result to the residual block output result corresponding to the previous feature extraction module and taking the average value to obtain a first calculation result; The first calculation result is upsampled to obtain a second upsampled result; and the second upsampled result is added to the residual block output result corresponding to the previous feature extraction module adjacent to the current feature extraction module and the average value is taken to obtain a second calculation result, and so on, until the residual block output result corresponding to the first feature extraction module is added to the calculation result corresponding to the next feature extraction module adjacent to the current feature extraction module and the average value is taken to obtain the calculation result corresponding to the first feature extraction module; Processing the depth-wise separable convolution result through a residual block to obtain a first convolution residual result; Upsampling the calculation result corresponding to the first feature extraction module to obtain a first feature upsampling result; Adding the first feature upsampling result and the first convolution residual result and taking the average value to obtain a first average result; Upsampling the first average result to obtain a second feature upsampling result; Processing the sMRI image to be generated through a plurality of residual blocks to obtain a second convolution residual result; Add the second feature upsampling result and the second convolution residual result and take the average value to obtain a second average result; The second average result is input into the prediction head to generate a PET image.

3. The medical image generation method according to claim 1, characterized in that: The step of processing the depthwise separable convolution result through the ASC sub-block to obtain an ASC sub-block output result includes: The depth-separable convolution result is processed by a first normalization, a first convolution layer and a first activation function to obtain a first processing result; The depth-separable convolution result is processed by a second normalization, a second convolution layer, and a second activation function to obtain a second processing result; multiplying the first processing result and the second processing result to obtain a multiplication result; The multiplication result is processed by a third normalization, a third convolution layer and a third activation function to obtain a third processing result; The third processing result is added to the depthwise separable convolution result to obtain an ASC sub-block output result.

4. The medical image generation method according to claim 1, characterized in that: The step of training the medical image generation model using the training data set to obtain a trained medical image generation model comprises: Constructing perceptual loss functions for medical image generation models; Based on the perceptual loss function, the medical image generation model is trained using the training data set to obtain a trained medical image generation model.

5. The medical image generation method according to claim 4, characterized in that: The perceptual loss function of constructing the medical image generation model includes: ; in, represents the perceptual loss function, Indicates the number of channels of the feature map, represents the height of the feature map, represents the width of the feature map, Represents the first The feature extraction function of the layer, represents the generated PET image, represents the real PET image, Represents the element-wise squared error.

6. A medical image generation system, characterized in that: The system comprises: A data acquisition unit, used to acquire the sMRI image to be generated, and to acquire a training data set constructed by paired sMRI images and PET images; A model construction unit, configured to construct a medical image generation model comprising a 3D encoder and a 3D decoder, wherein the 3D encoder comprises a depthwise separable convolution and a plurality of feature extraction modules comprising an MVM block and downsampling, the 3D decoder comprises a plurality of residual blocks, a plurality of upsampling and prediction heads, and the MVM block is configured to aggregate input information of the MVM block and extract features of different scales; An image generation unit, used to train the medical image generation model using the training data set to obtain a trained medical image generation model, so as to input the sMRI image to be generated into the trained medical image generation model to generate a PET image, including: inputting the depthwise separable convolution result into the MVM block in the first feature extraction module to obtain a first MVM block output result, wherein the depthwise separable convolution result is a result obtained by subjecting the sMRI image to be generated to the depthwise separable convolution processing, the MVM block includes an MVM sub-block and an ASC sub-block, the MVM sub-block is used to extract sequence features of the forward domain, the reverse domain and the mixed domain of the input MVM sub-block feature, and the ASC sub-block is used to aggregate feature information of the input ASC sub-block; Processing the depthwise separable convolution result through the ASC sub-block to obtain an ASC sub-block output result; The ASC sub-block output result is subjected to first-layer normalization processing to obtain a first-layer normalized result; Processing the first layer normalization result through the MVM sub-block to obtain an MVM sub-block output result; Adding the ASC sub-block output result and the MVM sub-block output result to obtain a first addition result; Subjecting the first addition result to a second-layer normalization process to obtain a second-layer normalization result; Processing the normalized result of the second layer through a multi-layer perceptron to obtain an output result of the multi-layer perceptron; Adding the multilayer perceptron output result and the first addition result to obtain a first MVM block output result; Downsampling the output result of the first MVM block to obtain an output result of the first feature extraction module; The output result of the first feature extraction module is used as the input of the MVM block in the next feature extraction module until all feature extraction modules are processed and the output result of each feature extraction module is obtained.

7. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the medical image generation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the medical image generating method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-mode three-dimensional medical image fusion method and system and electronic equipment

    CN110580695A