Cancer image processing method and system based on Mamba structure, and medium
By applying deep learning methods based on Mamba structure in cancer diagnosis, we build an improved neural network model for multi-task learning, solving the problem of image analysis accuracy and inefficiency in cancer diagnosis, and achieving more efficient cancer image processing and diagnosis.
Patent Information
- Application Number
- CN202510127983.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In the prior art, cancer diagnosis depends on the subjective judgment of the doctor, and the application of deep learning algorithms in medical image intelligent analysis has not been sufficient, resulting in inadequate cancer-assisted diagnosis accuracy and efficiency.
Using a deep learning method based on Mamba structure, an improved neural network model is constructed, and a multi-modal image feature extraction network and multi-task learning module are extracted, multi-scale features are performed and multi-task learning is performed to realize the lesion positioning and grading of cancer images.
It improves the detection and classification performance and efficiency of cancer images, enhances the robustness and accuracy of the cancer lesion grading module, and solves the problems of multimodal information fusion and multi-objective grading.
Smart Images

Figure CN120070355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a cancer image processing method, system and medium based on the Mamba structure. Background Art
[0002] In related technologies, the diagnosis of cancer mainly relies on the judgment of physicians. However, the diagnosis of radiologists is subjective and will be affected by the experience and knowledge level of physicians. Therefore, how to achieve automated and interactive diagnosis of cancer by combining medicine and artificial intelligence has become a major scientific problem currently faced. In order to meet the increasingly complex requirements of medical image processing, how to apply deep learning algorithms to the intelligent analysis of cancer images to improve the accuracy and efficiency of cancer auxiliary diagnosis has become a technical problem to be solved urgently. Summary of the Invention
[0003] The purpose of the present invention is to provide a cancer image processing method, system and medium based on the Mamba structure, which can improve the detection and classification performance and efficiency of cancer images.
[0004] To achieve the above purpose, the present invention provides the following technical solutions:
[0005] In a first aspect, an embodiment of the present invention provides a cancer image processing method based on the Mamba structure. The method includes the following steps:
[0006] S100, obtain the MRI image of prostate cancer, and mark the lesion area and grade in the MRI image;
[0007] S200, preprocess the MRI image, perform image expansion processing on the preprocessed MRI image to form a sample data set including multiple image modalities, and divide the sample data set into a training set and a test set; the sample data set includes multiple cancer image samples, and each cancer image sample has a corresponding image modality;
[0008] S300, construct an improved neural network model, input the cancer image samples in the training set into the improved neural network model, and extract multi-scale features of each image modality; wherein, the improved neural network model includes multiple multi-modal image feature extraction networks, and the number of multi-modal image feature extraction networks is the same as the number of image modalities; the multi-modal image feature extraction network includes an improved V-Net network combined with SSM, and the improved V-Net network includes a CNN-SSM block mixed with a convolutional neural network and a state space model;
[0009] S400. Feed the multi-scale features extracted by the multi-modal image feature extraction network through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder.
[0010] S500. Construct a multi-task learning module, input the multi-modal output features of the encoder into the multi-task learning module for multi-task learning to obtain a multi-task learning loss. Among them, the multi-task learning module includes a lesion detection module and a lesion grading module. The lesion detection module is used to locate the lesions in the image, and the lesion grading module is used to classify the lesions in the image.
[0011] S600. Input the cancer image samples in the test set into the improved neural network model, calculate the multi-task learning loss of the improved neural network model, and update the parameters of the improved neural network model according to the multi-task learning loss until convergence to obtain a trained model.
[0012] S700. Input the newly acquired MRI images into the trained model to obtain the lesion localization result of the cancer image and the classification result of the cancer image.
[0013] Preferably, in S300, the step of inputting the cancer image samples in the training set into the improved neural network model to extract the multi-scale features of each image modality includes:
[0014] S310. Input the cancer image samples of each image modality into the improved V-Net model in each multi-modal image feature extraction network one by one for M×M×M convolution operations to obtain a first intermediate feature map. Among them, the number of channels of the improved V-Net model is N 1 , the stride is K, and it is followed by a ReLU function with a decay rate of 0.85;
[0015] S320. Input the first intermediate feature map into the convolution block in the improved V-Net network for convolution operations to obtain a second intermediate feature map. Among them, the convolution block includes 2 M×M×M convolution operations, and the number of channels is N 1 , and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85;
[0016] S330. Input the second intermediate feature map into the max-pooling layer in the improved V-Net network for downsampling operations to obtain a third intermediate feature map. Among them, the stride of the max-pooling layer is K, and the feature size of the third intermediate feature map is 1 / 2 of the second intermediate feature map;
[0017] S340, loop and execute S320 and S330 until the size of the third intermediate feature map is less than the set value P, record the number of loops L and the number of channels N of the third intermediate feature map L+1 ;
[0018] S350, input the third intermediate feature map into the residual block in the improved V-Net network to obtain an output feature map; wherein, the residual block includes 2 convolution operations of M×M×M, and the number of channels is N L+1 , and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85;
[0019] S360, form the multi-scale features extracted by the improved neural network model from the output feature maps corresponding to each image modality.
[0020] Preferably, in S400, feeding the multi-scale features extracted by the multi-modal image feature extraction network through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder includes:
[0021] S410, input the output features corresponding to each image modality Figure 1 one-to-one into a plurality of Mamba blocks, and extract the 3D features of the output feature map;
[0022] S420, for the output feature maps corresponding to each image modality, use a flattening operation to reshape the 3D features of the output feature map into a 1D long sequence to obtain the multi-modal features corresponding to each image modality;
[0023] S430, perform feature fusion processing on the multi-modal features corresponding to each of the image modalities to obtain the fused features;
[0024] S440, input the fused features into the Mamba module to obtain the multi-modal output features of the encoder:
[0025] Preferably, in S500, constructing the multi-task learning module, inputting the multi-modal output features of the encoder into the multi-task learning module for multi-task learning to obtain the multi-task learning loss, including:
[0026] S510, input the multi-modal output features of the encoder into the lesion detection module and the lesion grading module respectively for multi-task learning, and the multi-task learning includes a lesion localization task and a lesion grading task;
[0027] S520, for each cancer lesion localization, input the multi-modal output features into the attention module for lesion localization processing to obtain the output features of the attention module;
[0028] S530. Determine the cross - entropy loss for cancer lesion grading based on the multi - modal output features and the output features of the attention module;
[0029] S540. Determine the loss for each classification task based on the grading probabilities output by the lesion grading module in each image modality;
[0030] S550. Determine the multi - task learning loss based on the cross - entropy loss for cancer lesion grading and the losses for each classification task.
[0031] Preferably, the output features of the attention module are:
[0032]
[0033] where ψ(·) represents a convolution operation with a stride of 1 and a convolution kernel of 1, σ(·) is used for activation, χ spatical represents spatial attention, α dec represents channel attention, represents the output features of the attention module.
[0034] Preferably, the cross - entropy loss for cancer lesion grading is:
[0035]
[0036] where y img represents the multi - modal output features, L dec represents the cross - entropy loss for cancer lesion grading.
[0037] Preferably, the losses for each classification task are expressed as:
[0038]
[0039] where K represents the number of data sets, g k represents the true label of grading, p k represents the probability of grading output by the lesion grading module, n represents the number of the classification task, represents the loss of the nth classification task.
[0040] Preferably, the expression of the multi - task learning loss is:
[0041]
[0042] where α dec represents the uncertain weight, α 1 ...α n represent the weights of n classification tasks and are obtained by network learning; L represents the multi - task learning loss.
[0043] Second aspect, an embodiment of the present invention provides a cancer image processing system based on the Mamba structure, and the system includes:
[0044] At least one processor;
[0045] At least one memory for storing at least one program;
[0046] When the at least one program is executed by the at least one processor, the at least one processor implements the method described in any one of the above.
[0047] Third aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the method described in any one of the above when executed by the processor.
[0048] The beneficial effects of the present invention are as follows: In view of the problems of multi-modal information fusion and multi-target grading in the cancer diagnosis process, the present invention introduces a multi-modal deep multi-task self-attention network to interactively fuse multi-modal information and construct multi-task correlation features, improving the robustness and accuracy of the cancer lesion grading module. In addition, in order to overcome the problem of low operating efficiency of the multi-modal long sequence data model, a deep network model based on the Mamba structure is constructed, and a state space model is introduced to reduce the time complexity of the algorithm operation and improve the detection and classification performance and efficiency of cancer images. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0050] Figure 1 is a schematic flowchart of the cancer image processing method based on the Mamba structure in the embodiment of the present invention;
[0051] Figure 2 is a schematic framework diagram of the cancer image processing method based on the Mamba structure in the embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the prostate cancer detection result provided by the embodiment of the present invention;
[0053] Figure 4 is a schematic structural diagram of the cancer image processing system based on the Mamba structure in the embodiment of the present invention. Detailed Embodiments
[0054] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the accompanying drawings, so as to fully understand the purpose, solution and effects of the present invention. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0055] Generally, mpMRI (Multi-parametric MRI) includes both anatomical sequences (T1-weighted imaging (T1WI) and T2-weighted imaging (T2WI)), and functional sequences (diffusion weighted imaging (DWI) and dynamic contrast enhanced (DCE)). For example, DWI images of prostate cancer are mainly used to detect tumors in the peripheral zone, T2WI images are mainly used to detect tumors in the transitional zone, and the ADC map combined with DWI images with different b values is one of the main modalities for diagnosing PCa. Therefore, it is far from enough to rely solely on single-modal image information for the intelligent auxiliary diagnosis of PCa.
[0056] Currently, the usual cancer image classification is mainly through doctors observing single-modal lesion information or observing by combining each image modality one by one, which is very complex and cumbersome, and both the accuracy and efficiency are very low.
[0057] Referring to Figure 1 , the present invention provides a cancer image processing method based on the Mamba structure, and the method includes the following steps:
[0058] S100, obtaining MRI images of prostate cancer, and marking the lesion areas and grades in the MRI images;
[0059] Specifically, the data of the present invention was collected from the People's Hospital of Haikou, Central South University and the Affiliated Hospital of Xiangya Medical College, and included a total of 98 prostate cancer patients from 2013 to 2016. Doctors performed magnetic resonance imaging on all patients and determined suspected cancer. 98 MRI images of patients with prostate cancer were collected. The three-dimensional DICOM format data was processed into two-dimensional slices and the image format was converted to the BMP format, and the size of the two-dimensional image was 512×512. Then, referring to relevant literature and under the guidance of experts, with the help of the Photoshop drawing tool, the lesion areas of the prostate were manually marked in the MRI images to obtain a.shp vector file composed of points, lines and surfaces as the ground truth map, and the PI-RADS grade and tumor IISUP grade of the prostate in the MRI images were marked, and the obtained classification label file was used as the gold standard.
[0060] In this embodiment, all examinations were performed on a 3T scanner (Achieva 3T; Philips Healthcare, Eindhoven, the Netherlands) using a 32-channel phased array coil. During this period, prostate biopsies were performed and diagnosed as prostate cancer. The pathological diagnosis was made by a pathologist certified by the hospital board according to the Gleason grading system. The data of the present invention were the MRI corresponding to the initial diagnosis of prostate cancer in 98 patients, and the voxel size of the images was 512×512×22. The actual scan field of view (FOV) of the patients was 400mm×400mm, and the thickness was 4mm. Note that the dataset used has passed the ethical review of the relevant hospital and obtained the consent of the informed patients.
[0061] PI-RADS (Prostate Imaging Reporting and Data System) is a standardized grading system for the imaging evaluation of prostate cancer, mainly used for the interpretation of the results of multi-parametric MRI (mpMRI) examinations. The PI-RADS grading system classifies prostate lesions into grades 1 to 5, and the higher the grade, the higher the likelihood of prostate cancer.
[0062] The ISUP (International Society of Urological Pathology) grading system is a grading method for pathological evaluation, mainly used for the prognostic evaluation of prostate cancer and other urological tumors. The ISUP grading system is simplified and optimized based on the Gleason scoring system, and classifies prostate cancer into grades 1 to 5.
[0063] S200, preprocess the MRI images, perform image expansion processing on the preprocessed MRI images to form a sample dataset containing multiple image modalities, and divide the sample dataset into a training set and a test set; the sample dataset contains multiple cancer image samples, and each cancer image sample has a corresponding image modality;
[0064] In some embodiments, due to the large acquisition field of view of multi-modal MRI images and the small area occupied by the prostate and tumor regions in the images, it is necessary to crop the original MRI images to extract the central region containing the prostate and tumor; since deep learning requires a large number of training samples, therefore, perform rigid transformation processing on the images before and after cropping, including flipping up and down, left and right, and rotating a certain angle to expand the sample size. Use the cropped MRI images and the expanded sample data as cancer image samples, and divide the cancer image samples into a training set and a test set according to a certain ratio. The cancer image samples include three image modalities: T2WI images, DWI images, and ADC images.
[0065] S300. Build an improved neural network model, input the cancer image samples in the training set into the improved neural network model, and extract multi-scale features of each image modality. Among them, the improved neural network model includes multiple multi-modal image feature extraction networks, and the number of multi-modal image feature extraction networks is the same as the number of image modalities. The multi-modal image feature extraction network includes an improved V-Net network combined with SSM, and the improved V-Net network includes a CNN-SSM block that mixes a convolutional neural network and a state space model.
[0066] It should be noted that the T2WI images, DWI images, and ADC images in the dataset are respectively input into the multi-modal image feature extraction network. Cancer image samples of different image modalities show different information, and multi-modal fusion extraction can more accurately represent the prostate cancer region.
[0067] The CNN-SSM block improves multi-modal image feature extraction and classification by combining the advantages of a convolutional neural network (CNN) in local feature extraction and the efficiency of a state space model (SSM) in capturing long-range dependencies.
[0068] S400. Feed the multi-scale features extracted by the multi-modal image feature extraction network through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder.
[0069] S500. Build a multi-task learning module, input the multi-modal output features of the encoder into the multi-task learning module for multi-task learning to obtain a multi-task learning loss. Among them, the multi-task learning module includes a lesion detection module and a lesion grading module. The lesion detection module is used to locate the lesions in the image, and the lesion grading module is used to classify the lesions in the image.
[0070] S600. Input the cancer image samples in the test set into the improved neural network model, calculate the multi-task learning loss of the improved neural network model, and update the parameters of the improved neural network model according to the multi-task learning loss until convergence to obtain a trained model.
[0071] S700. Input the newly acquired MRI images into the trained model to obtain the lesion localization result of the cancer image and the classification result of the cancer image.
[0072] Input the test data into the trained model to obtain the cancer image detection results and cancer image classification results; compare the results obtained in step S700 with the annotations in step S100, and evaluate the performance of the proposed model according to the evaluation metrics. The evaluation metrics for the cancer detection method include the correlation coefficient (CC), overlap rate (Overlap), Dice similarity coefficient (DSC), and accuracy (ACC).
[0073] Are defined as follows:
[0074]
[0075] Where A i and B i respectively represent the annotation map of the cancer region of the i-th scanned slice and the output of the model.
[0076]
[0077] Where A i and B i respectively represent the annotation map of the cancer region of the i-th scanned slice and the output of the model.
[0078]
[0079] The evaluation method for the cancer classification results uses 5 performance evaluation metrics to quantitatively evaluate the automatic grading results of prostate PI-RADS, including calculating ACC, Precision (PRE), Recall, Specificity (SPE), and F1 score (F1score, F1) to evaluate the performance.
[0080]
[0081] Where TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative respectively.
[0082] The present invention aims at the problems of multi-modal information fusion and multi-object grading in the cancer diagnosis process, introduces a multi-modal deep multi-task self-attention network, interactively fuses multi-modal information and constructs multi-task correlation features, and improves the robustness and accuracy of the cancer lesion grading module. In addition, in order to break through the problem of low operating efficiency of the multi-modal long sequence data model, a deep network model based on the Mamba structure is constructed, a state space model is introduced, the time complexity of the algorithm is reduced, and the detection and classification performance and efficiency of cancer images are improved.
[0083] Reference Figure 2, in some improved embodiments, in S300, inputting the cancer image samples in the training set into the improved neural network model to extract multi-scale features of each image modality includes:
[0084] S310, inputting the cancer image samples of each image modality into the improved V-Net model in each multi-modal image feature extraction network one by one for M×M×M convolution operation to obtain a first intermediate feature map; wherein, the number of channels of the improved V-Net model is N 1 , with a step size of K, and followed by a ReLU function with a decay rate of 0.85;
[0085] The input of the improved neural network model is the MRI voxels of the cancer image samples. Let the MRI voxels be V = {S 0 ,…S i ,…S n}, where S i represents the i-th slice, d l and d w represent the size of the slices of the MRI image, d l represents the length of the slice, d w represents the width of the slice.
[0086] S320, inputting the first intermediate feature map into the convolutional block in the improved V-Net network for convolution operation to obtain a second intermediate feature map; wherein, the convolutional block includes 2 M×M×M convolution operations, and the number of channels is N 1 , and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85;
[0087] S330, inputting the second intermediate feature map into the max-pooling layer in the improved V-Net network for downsampling operation to obtain a third intermediate feature map; wherein, the step size of the max-pooling layer is K, and the feature size of the third intermediate feature map is 1 / 2 of that of the second intermediate feature map;
[0088] S340, repeatedly execute S320 and S330 until the size of the third intermediate feature map is less than the set value P, record the number of loops L and the number of channels N of the third intermediate feature map L+1 ;
[0089] S350, inputting the third intermediate feature map into the residual block in the improved V-Net network to obtain an output feature map; wherein, the residual block includes 2 M×M×M convolution operations, and the number of channels is N L+1 , and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85;
[0090] S360, form the multi-scale features extracted by the improved neural network model from the output feature maps corresponding to each image modality.
[0091] Specifically, denote the output feature map corresponding to the T2WI image as the first feature map, the output feature map corresponding to the DWI image as the second feature map, and the output feature map corresponding to the ADC image as the third feature map; form the multi-scale features extracted by the improved neural network model from the first feature map, the second feature map, and the third feature map.
[0092] In some improved embodiments, in S400, feeding the multi-scale features extracted by the multi-modal image feature extraction network through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder includes:
[0093] S410, input the output features corresponding to each image modality Figure 1 one-to-one into multiple Mamba blocks, and extract the 3D features of the output feature map;
[0094] S420, for the output feature map corresponding to each image modality, use a flattening operation to reshape the 3D features of the output feature map into a 1D long sequence to obtain the multi-modal features corresponding to each image modality;
[0095] Specifically, feed the first multi-scale 3D feature z 0 extracted by the backbone layer through each Mamba block and the corresponding downsampling layer. For the multi-modal images corresponding to the same image modality, a flattening operation φ is used before the Mamba block to reshape the 3D features of each multi-modal image into multiple 1D long sequences for fusion, thereby achieving efficient sequence modeling with less inductive bias.
[0096] Input the first feature map, the second feature map, and the third feature map into each Mamba block respectively; for the 3D features in the cancer imaging samples, through the formula z l-1 = φ(z l-1 ) reshape the 3D features into 1D long sequences respectively to obtain the multi-modal features
[0097] S430, perform feature fusion processing on the multi-modal features corresponding to each image modality to obtain the fused features;
[0098] The feature fusion method is to input the multi-modal features of each image modality separately, and concatenate the multi-modal features of each image modality into the same sequence according to the corresponding serial numbers at the node generation stage. The formula for feature fusion processing is:
[0099]
[0100] Among them, is the fused feature;
[0101] S440, input the fused feature into the Mamba module to obtain the multi-modal output feature of the encoder:
[0102]
[0103] Among them, is the multi-modal output feature of the encoder.
[0104] In some improved embodiments, in S500, the multi-task learning module is constructed, and the multi-modal output feature of the encoder is input into the multi-task learning module for multi-task learning to obtain the multi-task learning loss, including:
[0105] S510, input the multi-modal output feature of the encoder into the lesion detection module and the lesion grading module respectively for multi-task learning, and the multi-task learning includes a lesion localization task and a lesion grading task;
[0106] Specifically, the modal output features of the multi-encoder are input into the lesion detection module and the lesion grading module respectively for multi-task learning. The multi-task learning module learns through multiple tasks, which are the lesion localization task and the lesion grading task respectively. The lesion grading task in this embodiment includes the PI-RADS grading task and the tumor ISUP grading task.
[0107] S520, for each cancer lesion localization, input the multi-modal output feature into the attention module for lesion localization processing to obtain the output feature of the attention module;
[0108] Specifically, for each cancer lesion localization, input the multi-modal output feature into the attention module for processing to obtain the output feature of the attention module; the output feature of the attention module is:
[0109]
[0110] Among them, ψ(·) represents a convolution operation with a step size of 1 and a convolution kernel of 1, σ(·) is used for activation, χ spatical represents spatial attention, α dec represents channel attention, represents the output feature of the attention module.
[0111] S530, determine the cross-entropy loss of cancer lesion grading based on the multi-modal output feature and the output feature of the attention module;
[0112] For cancer lesion grading, the cross-entropy loss of the cancer lesion grading is:
[0113]
[0114] Among them, y img represents the multi-modal output feature, and L dec represents the cross-entropy loss for cancer lesion grading.
[0115] At the tail of each branch of the joint grading task, a fully connected layer is used to classify each lesion. After generating lesion-specific features, they are concatenated to obtain multi-modal features and a classifier is trained for the final prediction.
[0116] S540, determining the loss of each classification task based on the grading probabilities output by the lesion grading module in each image modality;
[0117] The loss of each classification task is expressed as:
[0118]
[0119] Among them, K represents the number of datasets, and g k represents the true label of the grading, p k represents the probability of the grading output by the lesion grading module, and n represents the number of the classification task, representing the loss of the nth classification task.
[0120] Specifically, the loss functions in the PI-RADS grading task are respectively expressed as:
[0121]
[0122] The loss functions in the ISUP grading task are respectively expressed as:
[0123]
[0124] Among them, K represents the number of datasets, and g k represents the true label of the PI-RADS grading, p k represents the probability of the PI-RADS grading output by the lesion grading module. m k represents the true label of the ISUP grading, n k represents the probability of the ISUP grading output by the lesion grading module, and n represents the n tasks of cancer classification.
[0125] S550, determining the multi-task learning loss based on the cross-entropy loss of cancer lesion grading and the losses of each classification task.
[0126] The expression of the multi-task learning loss is:
[0127]
[0128] Among them, α dec represents an uncertain weight, and α 1 ...α n represent the weights of n classification tasks and are obtained through network learning; L represents the multi-task learning loss.
[0129] Specifically, for the PI-RADS grading and ISUP grading in the lesion grading task, the finally obtained multi-task learning loss is:
[0130]
[0131] Among them, α dec , α pi-rads and α isup represent uncertain weights, the weight of PI-RADS grading, and the weight of ISUP grading, and are obtained through network learning. The final model outputs the prostate lesion detection result, the prostate PI-RADS grading, and the ISUP grading of the tumor.
[0132] To verify the accuracy of this method, the method in the cancer medical record image embodiment is used to analyze the prostate cancer medical record image, and the detection result of prostate cancer is as Figure 3 shown, and the classification result of prostate cancer is shown in Table 1. Compared with the baseline network, adding only the SSM module (B+SSM), adding only the Mamba module (B+Mamba), and adding the SSM and Mamba modules but without multi-task learning processing (B+SSM+Mamba), the proposed method has greatly improved in the classification result.
[0133] Table 1:
[0134]
[0135]
[0136] Corresponding to the Figure 1 method, referring to Figure 4 , the embodiment of the present invention provides a cancer image processing system based on the Mamba structure, including:
[0137] At least one processor;
[0138] At least one memory for storing at least one program;
[0139] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0140] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented in the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0141] In addition, the embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program is stored in a computer-readable storage medium. The processor of the computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program to enable the computer device to execute the above method. Similarly, the content in the above method embodiments is applicable to the storage medium embodiments of the present invention. The functions specifically implemented in the storage medium embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0142] Those of ordinary skill in the art can understand that all or some of the methods and systems disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0143] The above is a specific description of the preferred embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A cancer image processing method based on Mamba structure, characterized in that: The method comprises the following steps: S100, acquiring an MRI image of prostate cancer, and marking the lesion area and grade in the MRI image; S200, preprocessing the MRI image, performing image expansion processing on the preprocessed MRI image to form a sample data set including multiple image modalities, and dividing the sample data set into a training set and a test set; the sample data set includes multiple cancer image samples, each of which has a corresponding image modality; S300, constructing an improved neural network model, inputting the cancer image samples in the training set into the improved neural network model, and extracting multi-scale features of each image modality; wherein the improved neural network model includes a plurality of multimodal image feature extraction networks, and the number of the multimodal image feature extraction networks is consistent with the number of image modalities; the multimodal image feature extraction network includes an improved V-Net network combined with SSM, and the improved V-Net network includes a CNN-SSM block that is a mixture of a convolutional neural network and a state space model; S400, feeding the multi-scale features extracted by the multimodal image feature extraction network through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder; S500, constructing a multi-task learning module, inputting the multimodal output features of the encoder into the multi-task learning module for multi-task learning, and obtaining a multi-task learning loss; wherein the multi-task learning module includes a lesion detection module and a lesion grading module, the lesion detection module is used to locate the lesions in the image, and the lesion grading module is used to classify the lesions in the image; S600, inputting the cancer image samples in the test set into the improved neural network model, calculating the multi-task learning loss of the improved neural network model, and updating the parameters of the improved neural network model according to the multi-task learning loss until convergence, thereby obtaining a trained model; S700, inputting the newly acquired MRI image into the trained model to obtain the lesion localization result of the cancer image and the classification result of the cancer image.
2. The method according to claim 1, characterized in that In S300, the step of inputting the cancer image samples in the training set into the improved neural network model to extract multi-scale features of each image modality includes: S310, inputting the cancer image samples of each image modality into the improved V-Net model in each multimodal image feature extraction network one by one, performing an M×M×M convolution operation, and obtaining a first intermediate feature map; wherein the improved V-Net model has N1 channels, a step size of K, and follows a ReLU function with a decay rate of 0.85; S320, inputting the first intermediate feature map into the convolution block in the improved V-Net network for convolution operation to obtain a second intermediate feature map; wherein the convolution block includes 2 M×M×M convolution operations, the number of channels is N1, and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85; S330, inputting the second intermediate feature map into the maximum pooling layer in the improved V-Net network for downsampling operation to obtain a third intermediate feature map; wherein the step size of the maximum pooling layer is K, and the feature size of the third intermediate feature map is 1 / 2 of the second intermediate feature map; S340, loop through S320 and S330 until the size of the third intermediate feature map is smaller than the set value P, and record the number of loops L and the number of channels N of the third intermediate feature map L+1 ; S350, inputting the third intermediate feature map into the residual block in the improved V-Net network to obtain an output feature map; wherein the residual block includes 2 M×M×M convolution operations, and the number of channels is N L+1 , and each convolution operation is followed by a ReLU function operation with a decay rate of 0.85; S360, forming the output feature graphs corresponding to each image modality into multi-scale features extracted by the improved neural network model.
3. The method according to claim 2, characterized in that In S400, the multi-scale features extracted by the multi-modal image feature extraction network are fed through each Mamba block and the corresponding downsampling layer to obtain the modal output features of the encoder, including: S410, inputting the output feature maps corresponding to each image modality into multiple Mamba blocks one by one, and extracting 3D features of the output feature maps; S420, for the output feature graph corresponding to each image modality, reshape the 3D features of the output feature graph into a 1D long sequence by using a flattening operation to obtain multimodal features corresponding to each image modality; S430, performing feature fusion processing on the multimodal features corresponding to each of the image modalities to obtain fused features; S440, input the fused features into the Mamba module to obtain the multimodal output features of the encoder.
4. The method according to claim 3, characterized in that In S500, the multi-task learning module is constructed, and the multi-modal output features of the encoder are input into the multi-task learning module for multi-task learning to obtain the multi-task learning loss, including: S510, inputting the multimodal output features of the encoder into a lesion detection module and a lesion grading module respectively to perform multi-task learning, wherein the multi-task learning includes a lesion localization task and a lesion grading task; S520, for each cancer lesion location, input the multimodal output features into the attention module for lesion location processing to obtain output features of the attention module; S530, determining a cross entropy loss for cancer lesion grading based on the multimodal output features and the output features of the attention module; S540, determining the loss of each classification task based on the classification probability output by the lesion grading module in each image modality; S550, determining a multi-task learning loss based on a cross entropy loss of cancer lesion grading and losses of each classification task.
5. The method according to claim 4, characterized in that The output features of the attention module are: Among them, ψ(·) represents the convolution operation with a step size of 1 and a convolution kernel of 1, σ(·) is used for activation, and χ spatical represents spatial attention, α dec represents channel attention, Represents the output features of the attention module.
6. The method according to claim 5, characterized in that The cross entropy loss of the cancer lesion classification is: Among them, y img represents the multimodal output feature, L dec Represents the cross entropy loss for cancer lesion grading.
7. The method according to claim 6, characterized in that The loss of each classification task is expressed as: Among them, K represents the number of data sets, g k represents the true value label of the classification, p k represents the probability of classification output by the lesion grading module, n represents the number of classification tasks, represents the loss of the n-th classification task.
8. The method according to claim 7, characterized in that The expression of the multi-task learning loss is: Among them, α dec represents uncertain weights, α1...α n represents the weights of n classification tasks and is obtained through network learning; L represents the multi-task learning loss.
9. A cancer image processing system based on Mamba structure, characterized in that: The system comprises: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 8 when executed by the processor.
Citation Information
Patent Citations
Glioma image classification method and system based on multi-task learning
CN116152560A
Multi-modal data fusion method and device
CN118395391A
Brain tumor image segmentation method based on multi-scale convolution and Mama structure
CN118447244A
Prostate tumor image analysis method based on bimodal multi-task learning model
CN118840308A
GFE-Mama neural network-based interpretable Alzheimer's progress classification method
CN118918363A