Method, device and medium for training multi-task prediction model for medical images

By performing equalization processing and feature fusion on multi-center, multi-device, and multi-tracer medical images, the accuracy and generalization performance issues of multi-task prediction models for multimodal images are solved, and stable multi-task prediction under different conditions is achieved.

CN115187841BActive Publication Date: 2026-07-24SHANGHAI UNIV OF MEDICINE & HEALTH SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV OF MEDICINE & HEALTH SCI
Filing Date
2022-07-01
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, the inconsistency of medical images from multiple centers, multiple devices, and multiple tracers limits the accuracy and generalization performance of lung cancer auxiliary diagnosis and prognosis prediction models. Moreover, most existing studies are single-task models, which are difficult to meet the multi-task requirements of multimodal images.

Method used

By equalizing medical images from multiple centers, multiple devices, and multiple tracers, and combining deep learning with traditional machine learning, deep learning features of PET and CT images are extracted and weighted to establish a multi-task prediction model, including lesion segmentation, feature extraction, and feature post-processing. Convolutional neural networks and feature adjustment methods are used to achieve multi-scale cascading of features.

Benefits of technology

It improves the accuracy of image processing and the reliability of the model, enhances the multi-task prediction capability under different centers, devices and tracers, and achieves stable multi-task prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187841B_ABST
    Figure CN115187841B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of training method, equipment and medium for the multi-task prediction model of medical image, the method includes the following steps: obtaining multi-center, multiple devices, multiple tracer initial medical image and clinical risk factor feature, the initial medical image includes PET image and CT image;First equalization processing is carried out to initial medical image corresponding lesion segmentation gold standard, obtain equalized 3D image sequence;Resampling is carried out to the 3D image sequence, extract deep learning feature and pre-defined image group feature;Second equalization processing is carried out to the deep learning feature, and feature post-processing is carried out to the equalized feature;The deep learning feature, pre-defined image group feature and clinical risk factor feature obtained are spliced, and final feature is obtained;With final feature as input, multi-task prediction model is trained and obtained.Compared with prior art, the present application has the advantages of high medical image processing accuracy, suitable for multi-modal image and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing, and in particular to a training method, device, and medium for a multi-task prediction model for medical images. Background Technology

[0002] Currently, radiomics research on lung cancer is increasingly trending towards multi-center, multi-modal, multi-task, and even multi-tracer approaches. Traditional machine learning and deep learning are two common methods in radiomics research. The traditional machine learning-based radiomics workflow includes image acquisition, lesion segmentation, feature extraction, feature dimensionality reduction and selection, and model building. Deep learning-based radiomics, on the other hand, automatically learns features and builds models after inputting images and labels. However, both lesion feature extraction in traditional machine learning and image analysis in deep learning are sensitive to image contrast, resolution, and noise. Inconsistencies in images and variations in radiomics features caused by different imaging protocols from different hospitals, equipment, and manufacturers pose a significant obstacle to building multi-center, big data-based models for lung cancer auxiliary diagnosis and prognosis prediction, making it difficult to guarantee the accuracy of medical image processing. Furthermore, in terms of task selection, most existing studies build predictive models based on single tasks, resulting in models that are limited in scope and have limited generalization performance. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a training method, device and medium for a multi-task prediction model for medical images that has high accuracy in medical image processing and is applicable to multimodal images.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] A training method for a multi-task prediction model for medical images, characterized by the following steps:

[0006] Acquire initial medical images and clinical risk factor characteristics from multiple centers, multiple devices, and multiple tracers, wherein the initial medical images include PET images and CT images;

[0007] The initial medical image corresponding to the gold standard for lesion segmentation is subjected to a first equalization process to obtain an equalized 3D image sequence;

[0008] The 3D image sequence is resampled to extract deep learning features and predefined image omics features;

[0009] The deep learning features are subjected to a second equalization process, which refers to the equalization between different centers, different devices and different tracers, and the equalized features are then subjected to feature post-processing.

[0010] The obtained deep learning features, predefined radiomics features, and clinical risk factor features are concatenated to obtain the final features;

[0011] Using the final features as input, a multi-task prediction model is trained to obtain the model.

[0012] Furthermore, the first equalization process includes:

[0013] Read the initial medical images and the corresponding gold standard for lesion segmentation, sort each slice separately, and normalize them using body weight and tracer intake values;

[0014] Determine the center point of each slice, and obtain multiple 2D slices around that center point;

[0015] Features of the 2D slices are extracted using a 2D convolutional neural network, and a 2D equalized image is obtained through feature adjustment and feature multi-scale cascading.

[0016] The first equalization process is performed on the PET image and the CT image respectively.

[0017] Furthermore, the feature adjustment and feature multi-scale cascading are performed using bilinear representation.

[0018] Furthermore, the extraction of deep learning features specifically involves:

[0019] Multiple PET and CT slices were obtained by resampling the 3D image sequence;

[0020] Deep learning features of the PET and CT slices were extracted using convolutional neural networks, respectively.

[0021] The predefined radiomics features were obtained using Pyradiomics.

[0022] Furthermore, for PET images, MN-Combat is used to implement the second equalization process; for CT images, MM-Combat is used to implement the second equalization process.

[0023] Furthermore, the feature post-processing includes feature dimensionality reduction and feature filtering.

[0024] Furthermore, the deep learning features obtained from PET and CT images are weighted and summed, and then concatenated with the predefined radiomics features and clinical risk factor features to obtain the final features.

[0025] The present invention also provides an electronic device, comprising:

[0026] One or more processors;

[0027] Memory; and

[0028] One or more programs stored in memory, the one or more programs including instructions for performing the training method for a multi-task prediction model for medical images as described above.

[0029] The present invention also provides a computer-readable storage medium including one or more programs executable by one or more processors of an electronic device, said one or more programs including instructions for performing the training method for a multi-task prediction model for medical images as described above.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. This invention performs equalization processing on PET / CT images from multiple centers / multi-devices / multi-tracers to obtain more accurate and reliable image features, thereby improving the reliability of model training and the generalization ability of the model.

[0032] 2. This invention constructs a convolutional neural network for deep learning feature extraction of PET and CT, extracts deep learning features of PET and CT respectively, and performs weighted fusion, resulting in high computational accuracy.

[0033] 3. This invention introduces convolutional neural networks for PET / CT image preprocessing, uses the preprocessed image features to build a prediction model, introduces task transfer, and establishes a multi-task prediction model with good robustness. It can obtain stable multi-task prediction results for PET / CT images from different centers / different devices / different tracers. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the process of the present invention;

[0035] Figure 2 This is a schematic diagram of the feature extraction network and multi-task prediction model constructed in this invention;

[0036] Figure 3 for Figure 2 A schematic diagram of the ResBlock structure. Detailed Implementation

[0037] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0038] This study proposes preprocessing methods for multi-center, multi-device, multi-modal, and multi-tracer medical images, as well as post-processing methods for lesion feature analysis, to effectively equalize the data.

[0039] Based on PET / CT image equalization, a convolutional neural network for feature extraction is constructed to extract the fused features of CT and PET images as deep learning features. Regarding data acquisition, most lung cancer-related prediction models are based on CT images, while multimodal PET / CT and PET / MR, due to their comprehensive functionality, have become current research hotspots. Therefore, for multi-center, multi-device, multimodal, and multi-tracer lung cancer medical images, this invention combines deep learning and traditional machine learning to propose a multi-task prediction model for lung cancer auxiliary diagnosis and prognosis.

[0040] Example 1

[0041] like Figure 1 As shown, this embodiment provides a training method for a multi-task prediction model for medical images, characterized by the following steps:

[0042] Step 1: Obtain initial medical images and clinical risk factor characteristics from multiple centers, multiple devices, and multiple tracers. The initial medical images include PET images and CT images.

[0043] Step 2: Perform a first equalization process on the gold standard for lesion segmentation corresponding to the initial medical image to obtain an equalized 3D image sequence;

[0044] Step 3: Resample the 3D image sequence to extract deep learning features and predefined image omics features;

[0045] Step 4: Perform a second equalization process on the deep learning features. This second equalization process refers to the equalization between different centers, different devices and different tracers, and perform feature post-processing on the equalized features.

[0046] Step 5: Concatenate the obtained deep learning features, predefined radiomics features, and clinical risk factor features to obtain the final features;

[0047] Step 6: Using the final features as input, train to obtain a multi-task prediction model.

[0048] The aforementioned method targets multi-center, multi-device, multi-modal, and multi-tracer lung cancer medical images. Combining deep learning and traditional machine learning, it performs a series of feature extractions on medical images and proposes a training method for a multi-task prediction model, enabling more accurate multi-task prediction using image features. These multi-tasks include diagnostic assistance and prognostic assessment.

[0049] In the above method, the first equalization process includes: reading the initial medical image and the corresponding gold standard for lesion segmentation, sorting each slice separately, and normalizing it using body weight and tracer intake value; determining the center point of each slice, and obtaining multiple 2D slices around the center point; extracting the features of the 2D slices through a 2D convolutional neural network, and obtaining a 2D equalized image through feature adjustment and feature multi-scale cascading.

[0050] The first equalization process is performed on both the PET image and the CT image. Specifically, the SUV value of the PET image is equalized, and the CT value of the CT image is equalized. In this embodiment, a multimodal medical image convolutional neural network is used for equalization, including the following steps:

[0051] (21) Read the original CT DICOM file and the corresponding gold standard for lesion segmentation, sort each slice separately, and fix the window width and window level (-600HU~1600HU);

[0052] (22) Determine the center point of each slice in the CT tumor region segmentation file, and obtain 64 slices around the center point;

[0053] (23) Read the original PET DICOM file and the corresponding gold standard for lesion segmentation, sort each slice separately, and normalize it using body weight and tracer intake value;

[0054] (24) Determine the center point of each slice in the PET tumor region segmentation file, and obtain 64 slices around the center point;

[0055] (25) Input the CT 2D slices into 2D-SS-CNN_CT and obtain the results. Input PET 2D slices into 2D-SS-CNN_PET to obtain the results.

[0056] (26) will Input 2D-SE-CNN_CT to get the results. Will Input 2D-SE-CNN_PET to get the results.

[0057] (27) will Input Generator_CT, and Inputting Generator_PET will use Bilinear Representation (BRP) for feature conditioning and multi-scale feature concatenation;

[0058] (28) Acquire 2D equalized PET and CT images respectively;

[0059] (29) Resample 2D CT and PET images to obtain 3D sequence images of patients.

[0060] like Figure 2 As shown, in this embodiment, the structure of 2D-SS-CNN_CT includes:

[0061] (31) A convolutional layer with a stride of 2;

[0062] (32) A 3*3 max pooling layer with a step size of 2;

[0063] (33) One Dense Block and one Transition layer;

[0064] (34) One Dense Block and one Transition layer;

[0065] (35) Five upsampling operations;

[0066] (36) A 1*1 convolutional layer;

[0067] (37) Obtain the output

[0068] The Dense Block consists of a Batch Normalization, a ReLU function, a 1*1 convolutional layer, a Batch Normalization, a ReLU, and a 3*3 convolutional layer; the Transition layer consists of a 1*1 convolutional layer and a 2*2 max pooling layer.

[0069] like Figure 2 As shown, in this embodiment, the structure of 2D-SS-CNN_PET includes:

[0070] (41) A 1*1 convolutional layer;

[0071] (42) A convolutional layer with a stride of 4 is connected to a ReLU activation function, and this is repeated three times;

[0072] (43) One ResBlock(256), repeated three times;

[0073] (44) A 3*3 transformed convolutional layer, connected to a Batch Normalization and ReLU activation function, repeated twice;

[0074] (45) A 1*1 convolutional layer;

[0075] (46) Obtain the output

[0076] The ResBlock consists of Batch Normalization repeated twice, a ReLU activation function, and a Weight layer, which performs an addition operation with the original input. Its structure is as follows: Figure 3 As shown.

[0077] like Figure 2 As shown, in this embodiment, the specific structure of 2D-SE-CNN_CT is as follows:

[0078] (51) A 3*3 transformed convolutional layer, a batch normalization and ReLU activation function;

[0079] (52) A 3*3 transformed convolutional layer, an average pooling layer with ReLU activation function, repeated three times;

[0080] (53) An average pooling layer, the output vector is flattened;

[0081] (54) A fully connected layer, a Leaky ReLU;

[0082] (55) One Structure Embedding, repeated twice;

[0083] (56) Obtain the output

[0084] like Figure 2 As shown, in this embodiment, the specific structure of 2D-SE-CNN_PET is as follows:

[0085] (61) A 3*3 transformed convolutional layer, a batch normalization and ReLU activation function;

[0086] (62) A 3*3 transformed convolutional layer, an average pooling layer with ReLU activation function, repeated three times;

[0087] (63) An average pooling layer, the output vector is flattened;

[0088] (64) A fully connected layer, a Leaky ReLU;

[0089] (65) One Structure Embedding, repeated twice;

[0090] (66) Obtain the output

[0091] In this embodiment, the BRP process is specifically as follows:

[0092]

[0093] Among them, I f ∈R D Indicates the output of the previous layer, I c ∈R D′ Let D and D' represent the conditional vectors, where D and D' are the feature dimensions, [I f I c ]∈R D+D′ For connection representation, W i As the weight, I O To output a tensor.

[0094] like Figure 2 As shown, in this embodiment, the specific structures of Generator_CT and Generator_PET are as follows:

[0095] (81) A 4x4 convolutional layer;

[0096] (82) One ResBlock layer with two downsampling operations, repeated three times;

[0097] (83) One ResBlock layer, repeated three times;

[0098] (84) One ResBlock layer with two downsampling operations, repeated three times;

[0099] (85) A 3*3 convolutional layer and a Tanh activation function;

[0100] (86) Obtain a balanced medical image.

[0101] The ResBlock structure is as follows: Figure 3 As shown.

[0102] In step 3, the extraction of deep learning features specifically involves: resampling the 3D image sequence to obtain multiple PET and CT slices. In a specific implementation, 64 slices are obtained for each case. Deep learning features of the PET and CT slices are extracted using a convolutional neural network.

[0103] In step 3, the preprocessed lesion segmentation gold standard file is resampled, with 64 slices per case. Pyradiomics is used to obtain the predefined radiomics features, including features calculated based on the original image and filtered images (specifically, 12 filters).

[0104] In step 4, the second equalization process is implemented using MN-Combat for PET images and MM-Combat for CT images. The feature post-processing includes feature dimensionality reduction and feature selection.

[0105] Among them, MN-Combat and MM-Combat are:

[0106]

[0107]

[0108] Where i represents different centers / different equipment / different tracer types, j represents different samples, and g represents different features; This represents the average value. γ represents the potential noncentrally correlated covariates and coefficients in the model. ig δ represents the central effect parameter of the positive gamma prior distribution. ig Let F represent the central effect parameter of the inverse gamma prior distribution, F represent the original eigenvalue, f represent the eigenvalue after batch effect removal, and k∈[1,…,1000].

[0109] The deep learning features obtained from PET and CT images are weighted and summed, and then concatenated with the predefined radiomics features and clinical risk factor features to obtain the final features.

[0110] The weighted summation of deep learning features is specifically as follows:

[0111]

[0112] in, Deep learning features of PET images and CT images for each patient were analyzed separately.

[0113] In this embodiment, a multi-task prediction model is constructed using a 1D CNN. The preprocessed image features are input into the 1D CNN to train the network. If there are multiple tasks, task transfer is performed, parameters are optimized, and the model is retrained.

[0114] like Figure 2 As shown, the specific structure of the 1D CNN in this embodiment includes:

[0115] (12_1) One 1D convolutional layer, one ReLU, one average pooling layer, repeated four times;

[0116] (12_2) One 1D convolutional layer, one ReLU, and one global max pooling layer;

[0117] (12_3) Two fully connected layers;

[0118] (12_4) The Softmax function is used to obtain the prediction results.

[0119] If the above methods are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0120] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A training method for a multi-task prediction model for medical images, characterized in that, Includes the following steps: Acquire initial medical images and clinical risk factor characteristics from multiple centers, multiple devices, and multiple tracers, wherein the initial medical images include PET images and CT images; The initial medical image corresponding to the gold standard for lesion segmentation is subjected to a first equalization process to obtain an equalized 3D image sequence; The 3D image sequence is resampled to extract deep learning features and predefined image omics features; The deep learning features are subjected to a second equalization process, which refers to the equalization between different centers, different devices and different tracers, and the equalized features are then subjected to feature post-processing. The obtained deep learning features, predefined radiomics features, and clinical risk factor features are concatenated to obtain the final features; Using the final features as input, a multi-task prediction model is trained to obtain the model. The first equalization process includes: Read the initial medical images and the corresponding gold standard for lesion segmentation, sort each slice separately, and normalize them using body weight and tracer intake values; Determine the center point of each slice, and obtain multiple 2D slices around that center point; Features of the 2D slices are extracted using a 2D convolutional neural network, and a 2D equalized image is obtained through feature adjustment and feature multi-scale concatenation. Bilinear representation is used for feature adjustment and feature multi-scale concatenation. The first equalization process is performed on the PET image and the CT image respectively; For PET images, MN-Combat is used to implement the second equalization process; for CT images, MM-Combat is used to implement the second equalization process.

2. The training method for a multi-task prediction model for medical images according to claim 1, characterized in that, The extraction of deep learning features specifically involves: Multiple PET and CT slices were obtained by resampling the 3D image sequence; Deep learning features of the PET and CT slices were extracted using convolutional neural networks, respectively.

3. The training method for a multi-task prediction model for medical images according to claim 1, characterized in that, The predefined radiomics features were obtained using Pyradiomics.

4. The training method for a multi-task prediction model for medical images according to claim 1, characterized in that, The feature post-processing includes feature dimensionality reduction and feature filtering.

5. The training method for a multi-task prediction model for medical images according to claim 1, characterized in that, The deep learning features obtained from PET and CT images are weighted and summed, and then concatenated with the predefined radiomics features and clinical risk factor features to obtain the final features.

6. An electronic device, characterized in that, include: One or more processors; Memory; and One or more programs stored in memory, the one or more programs including instructions for performing the training method for a multi-task prediction model for medical images as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, It includes one or more programs that are executed by one or more processors of an electronic device, the one or more programs including instructions for performing the training method for a multi-task prediction model for medical images as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Pulmonary nodule benign and malignant prediction method and device

    CN111915596A

  • Epidermal growth factor receptor mutation state judgment method, medium and electronic equipment

    CN112488992A