Tumor benign and malignant classification method, system and device based on multi-modal data fusion and deep learning and storage medium thereof

Through the multimodal data fusion and deep learning methods, combined with global-hybrid-local networks and integrated learning, the problem of underutilization of multimodal data in the existing technology is solved, efficient classification of benign and malignant tumors is achieved, and the credibility of diagnosis is improved through interpretability analysis.

CN120014321APending Publication Date: 2025-05-16SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510010130.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the complementarity of multimodal data in tumor detection and diagnosis, resulting in bottlenecks in diagnostic accuracy, and the "black box" problem of deep learning models limits the interpretability of diagnostic results.

Method used

The tumor benign and malignant classification method based on multimodal data fusion and deep learning is adopted. By acquiring patient image data and laboratory data, intelligent segmentation and feature extraction are performed, combined with global-hybrid-local networks and integrated learning methods, efficient fusion and feature extraction of multimodal data are achieved, and interpretability analysis is provided through SHAP and occlusion methods.

Benefits of technology

It improves the efficiency of medical imaging feature extraction, fully integrates medical information from multiple modalities, provides a basis for explanation of decisions, and improves the accuracy and credibility of tumor benign and malignant classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014321A_ABST
    Figure CN120014321A_ABST
Patent Text Reader

Abstract

According to the tumor benign and malignant classification method based on multi-modal data fusion and deep learning, image data of a patient and laboratory data are fused, and high-precision detection of benign and malignant tumors is achieved; the method comprises the following steps: intelligently segmenting a patient image through a U-shaped network, and screening out interested area slices; performing feature extraction and classification on the image data by using the designed global-hybrid-local network to obtain a tumor benign and malignant prediction result of the image data; inputting the image prediction result and the clinical features into an integrated learning classification model, and combining the advantages of various machine learning classifiers to obtain a multi-modal tumor benign and malignant prediction result; the SHAP method and the occlusion-based method are adopted to carry out visual result interpretation on different parts of input elements and images, so that the interpretability and transparency of the model are improved. The invention provides powerful support for early discovery and treatment of tumors, and has remarkable clinical application value and popularization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information processing technology, and specifically to a method, system, device and storage medium for classifying benign and malignant tumors based on multimodal data fusion and deep learning. Background Art

[0002] In the medical field, tumors are growths formed by abnormal overgrowth of body cells and tissues, which seriously threaten human life and health. Early detection of tumors is crucial. The deep learning architecture based on convolutional neural networks has realized the detection of imaging modality tumors on single modality data. F.et al. "Brain tumor detectionbased on Convolutional Neural Network with neutrosophic expert maximum fuzzysure entropy," Measurement, 2019, 147: 106830). However, a single form of data limits the ability of deep networks to mine and integrate multivariate data information, resulting in a bottleneck in diagnostic accuracy. This limitation is particularly evident in current technologies because tumor detection and diagnosis often rely on clinical data from multiple modalities.

[0004] The key challenge facing the development of intelligent medical diagnosis technology is how to efficiently process heterogeneous and complementary clinical data of different modalities and make objective quantitative decisions. Although some studies have attempted to apply multimodal deep learning technology to the extraction and fusion of clinical data, most of the existing methods remain at the level of simple data splicing and fail to give full play to the complementary advantages of multimodal data. This splicing method ignores the complex associations between different modal data, resulting in insufficient information integration and limited improvement in diagnostic accuracy. Although researchers have tried to splice multiple modal visual features extracted from convolutional neural networks with other data sources to improve the accuracy of disease detection, this splicing method still lacks in-depth data fusion and feature extraction mechanisms, making it difficult to fully tap the potential value of multimodal data. (Y.Yang.et al "Skin lesion classification based on two-modal images using a multi-scale fully-shared fusion network," Comput.Methods Programs Biomed, 2023, 107315)

[0005] Among the various modalities of medical diagnosis, images provide the most abundant and special information. However, existing artificial intelligence and medical imaging technologies are still insufficient in distinguishing benign and malignant tumors. Although some studies have proposed deep learning-based models to process image data of specific modalities. (Wang, Z. et al. "Detection of COVID-19 Cases Based on Deep Learning with X-ray Images," Electronics 2022, 11, 3511), this work proposes a custom ResNet model MHSA-ResNet, which extracts texture features from images and uses a multi-head self-attention mechanism to classify lung images into three categories (normal, pneumonia, and COVID-19), but these models are often limited to single-modality data and fail to fully utilize the complementarity of multimodal data. In addition, although some researchers have tried to build multi-channel input deep learning network models based on different modal imaging data (such as dermoscopy and clinical images) to improve the detection of skin diseases (Omeroglu AN. et al. "A novel soft attention-based multi-modal deep learning framework for multi-label skin lesion classification," Engineering Applications of Artificial Intelligence, 2023, 120: 105897), the application of these models in the field of tumor detection is still insufficient, and there is a lack of in-depth exploration of the complex relationship between different modal data.

[0006] In clinical practice, there are huge differences in the number of dimensions and features (multiple orders of magnitude) between different modalities of data, such as tumor image features and metadata. Multimodal data fusion and tumor detection aim to integrate clinical data of different distributions, sources, and types into a global space that can uniformly represent inter-modal and cross-modal information, which increases the difficulty of multimodal data fusion. Ordinary multi-data source splicing methods cannot effectively integrate multimodal semantic information, resulting in information loss and decreased diagnostic accuracy. Therefore, it is necessary to explore efficient multimodal data fusion methods to achieve feature complementarity and global spatial representation between different modal data.

[0007] There are large intra-group differences and small inter-group differences in the expression of tumor imaging data. Most existing methods use single-scale images as input, which easily miss important lesion information. In order to overcome this limitation, a multi-scale method is introduced to take global images and local lesions as inputs and perform efficient information fusion. However, there is no effective application of such multi-scale methods in the current research on disease diagnosis, resulting in room for improvement in diagnostic accuracy.

[0008] In addition, existing deep learning models often have a "black box" problem when building neural network-assisted intelligent diagnosis, that is, the diagnostic results output by the model lack interpretability. This limits clinicians' trust and understanding of the diagnostic results. Summary of the invention

[0009] The purpose of the present invention is to provide a multimodal deep learning method and system for benign and malignant tumor detection, which can not only improve the efficiency of medical image feature extraction, but also integrate medical information of multiple modalities and provide corresponding interpretable basis for decision-making, thereby improving the accuracy and credibility of classification.

[0010] The technical solution of the present invention is as follows:

[0011] The method for classifying benign and malignant tumors based on multimodal data fusion and deep learning includes the following steps:

[0012] Obtain patient imaging and laboratory data;

[0013] Segment the organs and tumor areas of patient images to select slices of interest;

[0014] Extract and classify the image slice data of the screened patients to obtain the prediction results of the benign and malignant tumors of the image data;

[0015] The benign and malignant prediction results of the images and the clinical characteristics of the patients are input into the integrated learning classification model to obtain the benign and malignant prediction results of the multimodal tumors;

[0016] It is used to interpret the visual results of different parts of multimodal input elements and images respectively.

[0017] The patient's clinical scale data refers to tumor markers and various blood tests, which are stored in the form of a table.

[0018] The intelligent segmentation of organs and tumor areas in the patient's image is as follows:

[0019] The tumor image and its corresponding benign or malignant tumor label are input into a U-shaped network; the U-shaped network is a convolutional neural network model for image segmentation, which captures contextual information and achieves accurate pixel-level prediction through a symmetrical encoder and decoder structure; the trained U-shaped network model is used for intelligent segmentation to obtain the region of interest for tumor diagnosis.

[0020] The feature extraction and classification of patient image data to obtain the benign and malignant tumor prediction results of the image data specifically includes input data processing and sending the image data to the designed global-hybrid-local network for benign and malignant tumor classification.

[0021] The input data processing is specifically to select the best N slices according to the principle that the region of interest accounts for the largest proportion in the image based on the obtained intelligent segmentation results, and then cut these N slices into global and local slices according to the region of interest, and then send both to the designed imaging classification network at the same time.

[0022] The designed global-hybrid-local imaging classification network is as follows:

[0023] The network contains three network branches, namely the global branch, the mixed branch, and the local branch. The three branches are used to extract features from the image data respectively, and finally feature fusion and classifier classification are performed to obtain the benign and malignant classification results of tumor imaging.

[0024] The global branch and local branch use the residual connection network as the backbone network for feature extraction. Specifically, both branches contain a shallow module and four deep modules, where the number of groups in each module is 3, 3, 4, 6 and 3, respectively, which is consistent with the main architecture of the residual connection network. The modules are connected through an average pooling layer. The group in each shallow module consists of "a 3×3 convolution layer, a batch normalization (BN) layer and a ReLu layer". The subsequent deep modules introduce multi-scale modules and spatial-channel attention modules based on the residual block to optimize the feature extraction effect.

[0025] The residual connection network introduces jump connections between network layers, so that the deep network can effectively alleviate the gradient vanishing problem during training. The global branch and the local branch respectively use two types of data of different sizes, namely the global image and the local image, as inputs of the residual connection network.

[0026] The multi-scale module introduces a multi-scale feature fusion mechanism, and by fusing feature maps of multiple scales in each residual block, the network's ability to express multi-scale information of tumor images can be enhanced.

[0027] The described added spatial-channel attention module is a method that combines spatial attention and channel attention; spatial attention uses the feature map after channel compression to perform convolution operations to generate spatial weights, emphasize important areas in the input feature map, and suppress unimportant parts; channel attention generates channel weights through global average pooling and global maximum pooling to reflect the importance of each channel, and then generates weights through the Sigmoid activation function to proportionally adjust the channel response of the input feature map.

[0028] The hybrid branch uses the attention feature fusion method to fuse the low-dimensional features of the global branch and the local fourth module, and further extracts features to integrate semantic information of different scales.

[0029] The described attention feature fusion method aggregates multi-scale contextual information along the channel dimension, which can simultaneously focus on the global and local semantic information of tumor image distribution, alleviating the problems caused by scale changes and small object instances.

[0030] The feature fusion of the three branches is specifically to transform the feature maps obtained by each of the three branches into one-dimensional vectors and splice them, and finally send them to the fully connected classifier to obtain the prediction probability value of the benign and malignant tumor of the imaging.

[0031] The multimodal tumor benign and malignant detection method is to add the imaging benign and malignant prediction probability value as a new dimension to the scale data, and then use multiple machine learning classifiers and ensemble learning methods to obtain the tumor benign and malignant discrimination results of the multimodal data. Among them, the ensemble learning method uses linear regression to train based on the prediction values ​​obtained by multiple machine learning classifiers to obtain the final ensemble prediction result.

[0032] The interpretable method mainly includes: using the SHAP method to perform decision weight analysis on different input elements of different modalities; using an occlusion-based method to perform interpretable attribution on different parts of the image.

[0033] Among them, the SHAP method is a type of technology used to explain the output of machine learning models. It understands the importance of each feature by assigning the contribution value of the feature to the model prediction.

[0034] The occlusion-based method uses a sliding window to occlude different parts of the image, and then iteratively inputs them into the imaging classification model to determine the contribution of different parts of the image to the decision of whether the tumor is benign or malignant by comparing the differences in the outputs.

[0035] The present invention also discloses a tumor benign and malignant classification system based on multimodal data fusion and deep learning, which includes the following modules:

[0036] A data acquisition module, used to acquire patient imaging data and laboratory data;

[0037] Image segmentation module, used to segment organs and tumor areas in patient images to select slices of regions of interest;

[0038] The imaging classification module is used to extract features and classify the image slice data of the screened patients to obtain the prediction results of the tumor benign or malignant;

[0039] The multimodal detection module is used to input the benign and malignant prediction results of the image and the patient's clinical characteristics into the integrated learning classification model to obtain the multimodal tumor benign and malignant prediction results;

[0040] The interpretable module is used to interpret the visualization results of each multimodal input element and different parts of the image. The present invention also discloses an electronic device, which is characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the above method steps when executing the program stored in the memory. The present invention also discloses a computer-readable storage medium on which a computer program is stored, characterized in that the computer program implements the above method steps when executed by a processor.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1. Based on multimodal clinical data, the present invention realizes an integrated and robust intelligent medical auxiliary diagnosis system, which can effectively integrate imaging information and clinical examination information to obtain prediction results of the benign and malignant nature of tumors, and can effectively deal with information acquisition, fusion and benign and malignant diagnosis decisions in different tumor scenarios, and assist doctors in making more accurate judgments on the benign and malignant nature of tumors;

[0043] 2. The present invention constructs a multi-scale intelligent disease detection network based on a deep learning algorithm to address the problem of large intra-group differences and small inter-group differences in the expression of tumor image data. It combines the spatial and channel attention mechanisms to fully extract and fuse the semantic information contained in the global and local scales of the image. A feature fusion module is designed based on the attention mechanism to aggregate multi-scale contextual information along the channel dimension. It can simultaneously focus on large targets distributed globally and small targets distributed locally, alleviate the interference caused by scale changes and small object instances, and further improve the accuracy of disease diagnosis.

[0044] 3. The present invention uses the SHAP method to perform decision weight analysis on different input factors of different modalities, and uses an occlusion-based method to perform interpretable attribution on different parts of the image, providing visual auxiliary analysis for the results obtained by the diagnostic network, thereby improving the reliability and credibility of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flow chart of a method for classifying benign and malignant tumors based on multimodal data fusion and deep learning according to an embodiment of the present invention;

[0046] Figure 2 This is a schematic diagram of a global-hybrid-local network architecture according to an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of the composition of the MS-scSE Block in the image classification network provided by an embodiment of the present invention;

[0048] Figure 4 A schematic diagram of the composition of a network attention feature fusion module provided in an embodiment of the present invention;

[0049] Figure 5 A schematic diagram of a multimodal clinical data fusion architecture provided by an embodiment of the present invention;

[0050] Figure 6 A schematic diagram of an explainable framework for benign and malignant tumor detection provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.

[0052] See also Figure 1 , Figure 1 Flow chart of a method for classifying benign and malignant tumors based on multimodal data fusion and deep learning according to an embodiment of the present invention. Figure 1 As shown, the following steps are included:

[0053] Step 1. Obtain patient imaging data and laboratory data; In this embodiment, the patient's imaging data is abdominal enhanced CT scan data, and the laboratory data refers to tumor markers and various blood test data, as shown in Table 1. The combination of imaging data and clinical data provides a comprehensive view of the patient. Table 1: Patient laboratory data indicators

[0054] Step 2. Input the tumor image and its corresponding benign or malignant tumor label into the U-net, and use the U-net in deep learning to intelligently segment the organs and tumor areas in the patient image to screen out the area slices of interest. U-Net is a convolutional neural network model for image segmentation. It captures contextual information and achieves accurate pixel-level prediction through a symmetrical encoder and decoder structure. The encoder part gradually downsamples the input image to extract features; the decoder part gradually upsamples the extracted features to restore them to the size of the original image for accurate segmentation. Specifically, the encoder part contains four downsampling modules, each of which consists of two 3x3 convolutional layers and a 2x2 maxpooling layer. The decoder part is similarly composed of four upsampling modules, each of which contains a deconvolution layer, a feature concatenation layer, and two 3x3 convolutional layers in sequence. This structure enables the U-shaped network to effectively capture key features when processing image data.

[0055] Step 3. Extract features and classify the image slice data of the screened patients to obtain the prediction results of the benign and malignant tumors of the image data. This process includes the input data processing stage, which will select the best N slices based on the intelligent segmentation results to ensure that these slices can fully represent the region of interest. Subsequently, the processed image data is sent to the designed global-hybrid-local network for tumor benign and malignant classification. In this embodiment, since the Z-axis voxel spacing of the original three-dimensional imaging data is 0.5 mm, the number of slices N selected is 15. This selection ensures that the selected slices can remove most of the background and cover most of the region of interest. See also Figure 2 , Figure 2 Schematic diagram of the global-hybrid-local network architecture of an embodiment of the present invention. Figure 2 The global-hybrid-local imaging classification network shown in the figure includes three branches: global branch, hybrid branch and local branch. These three branches extract features from the imaging data respectively, and finally perform feature fusion and classification to obtain the benign and malignant classification results of tumor imaging. The global-hybrid-local network design is specifically introduced as follows: Global branches: In this embodiment, the input image data is 256*256 size; A residual connection network (similar to the ResNet-50 architecture) is used as the backbone network, which includes one shallow module and four deep modules. The number of groups in each module is 3, 3, 4, 6, and 3, respectively, which is consistent with the main architecture of the residual connection network ResNet-50. The group in each shallow module consists of "a 3×3 convolution layer, a batch normalization (BN) layer, and a ReLu layer", which helps to improve the learning ability of the model. The deep module is based on the residual connection network. By introducing jump connections between network layers, the gradient vanishing problem that may be encountered by the deep network during training is effectively alleviated. The global branch focuses on extracting the global features of the image, providing a macro perspective for the classification of benign and malignant tumors. Local branch: In this embodiment, the input image data is 128*128 size. The residual connection network is also used as the backbone, and the structure is similar to the global branch, but the input size and feature extraction focus are different. The partial branch aims to capture the local features of the image, especially the subtle changes in the tumor area. Hybrid Branch : The multi-scale module and the spatial-channel attention module are introduced. Specifically, the multi-scale module introduces a multi-scale feature fusion mechanism. By fusing feature maps of multiple scales in each residual block, the network's ability to express multi-scale information of tumor images is enhanced. This ability is especially important for detecting tumors of different sizes. The spatial-channel attention module combines the spatial attention and channel attention methods. Spatial attention generates spatial weights by performing convolution operations on the feature maps after channel compression to emphasize important areas in the input feature map and suppress unimportant parts; channel attention generates channel weights through global average pooling and global maximum pooling to reflect the importance of each channel, and then generates weights through the Sigmoid activation function to proportionally adjust the channel response of the input feature map. The attention feature fusion method (iAFF) is used to fuse the low-dimensional features of the global branch and the local branch, and further extract features to integrate semantic information at different scales. Specifically, the MS-CAM module is first introduced to aggregate multi-scale contextual information along the channel dimension, which can simultaneously focus on the global and local semantic information of tumor images, thereby alleviating the problems caused by scale changes and small object instances. iAFF reuses the MS-CAM module twice to generate the weights when the two feature maps are fused, thereby achieving efficient feature fusion. The feature maps obtained by each of the three network branches are transformed into one-dimensional vectors and concatenated to form a fused feature vector. The fused feature vector is sent to the fully connected classifier to obtain the prediction results of benign and malignant tumors in imaging. Step 4. Input the image benign and malignant prediction results and the patient's clinical characteristics into the integrated learning classification model to obtain the multimodal tumor benign and malignant prediction results. Specifically include: Step 4.1: Add the benign or malignant tumor prediction results of imaging as a new dimension to the scale data (as shown in Table 1) to form a multimodal dataset. Step 4.2 uses a variety of machine learning classifiers (such as CatBoost, SVM (Support Vector Machine), Decision Tree, Random Forest, KNN (K Nearest Neighbor), MLP (Multi-layer Perceptron), etc.) to make preliminary predictions. Step 4.3 uses an ensemble learning method (based on linear regression) to combine the prediction results of multiple classifiers in step 4.2 to obtain the final ensemble prediction results. Among them, the ensemble learning method improves the prediction performance and reduces the defects that may exist in a single model by combining multiple learning models. These learning models can be homogeneous (such as multiple decision trees) or heterogeneous (such as a combination of decision trees and SVM). By integrating their prediction results, the advantages of each model can be fully utilized while reducing the deviations and errors that may be caused by a single model. Step 5. Visualize and interpret the results of each multimodal input element and different parts of the image, which helps doctors better understand the prediction results and make more accurate diagnostic decisions, such as Figure 6 As shown, specifically including: Step 5.1 Use the SHAP (SHapley Additive exPlanations) method to analyze the decision weights of different input factors of different modalities and understand the importance of features. Step 5.2 uses an occlusion-based method to interpretably attribute different parts of the image. By comparing the differences in model output before and after occlusion, the contribution of each part of the image to the decision of whether the tumor is benign or malignant is determined. Step 5.3 shows the impact of different parts of the image data on the final decision result in the form of a heat map to improve the interpretability and transparency of the model. This embodiment not only considers the patient's imaging data (such as abdominal enhanced CT scan), but also combines laboratory data (such as tumor markers and various blood tests). This fusion of multimodal data provides a comprehensive patient view, allowing doctors to more accurately assess the benign and malignant nature of tumors. The image data is intelligently segmented through the U-Net network (U-Net) in deep learning to screen out the image slices of interest. And combined with global branches, mixed branches and local branches, feature extraction and classification of image data from different scales and angles, the network's ability to express multi-scale information of tumor images is enhanced, and the accuracy of classification is improved. Further, by introducing a multi-scale feature fusion mechanism and a space-channel attention module, the key features in tumor images can be more effectively captured, and the robustness and generalization ability of the model can be improved. At the same time, the prediction results of multiple classifiers are integrated by linear regression to improve the overall stability and accuracy, and the contribution value of features to model prediction is calculated to help understand the importance of each feature. By blocking different parts of the image and observing the changes in the model output, the method can determine which parts of the image contribute the most to the benign and malignant decision of the tumor. The present invention combines the advantages of multimodal data, adopts image classification network and ensemble learning method, and provides detailed interpretability, which provides strong support for the early detection and treatment of tumors.

Claims

1. A method for classifying benign and malignant tumors based on multimodal data fusion and deep learning, characterized in that: include: Obtaining patient imaging data and laboratory data, wherein the imaging data is abdominal enhanced CT scan data, and the laboratory data includes tumor markers and various blood test data; The U-Net model is used to segment the organs and tumor regions of patient image data, and to select the slices of the region of interest containing tumor diagnostic information; The screened patient imaging data is subjected to feature extraction and classification through a global-hybrid-local network to obtain a tumor benign or malignant prediction result of the imaging data; wherein the global-hybrid-local network includes three network branches, namely a global branch, a hybrid branch, and a local branch, and the three network branches are used to extract features from the imaging data, fuse the features, and perform classification; The benign and malignant prediction results of the imaging data and the clinical characteristics of the patient are input into the integrated learning classification model to obtain the multimodal tumor benign and malignant prediction results; The visual results are interpreted for each multimodal input element and different parts of the image, including the SHAP method for decision weight analysis of different input elements of different modalities, and the occlusion-based multi-scale method for interpretable attribution of different parts of the image, which are displayed in the form of a heat map.

2. The method for classifying benign and malignant tumors according to claim 1, characterized in that: The U-Net model is a convolutional neural network model for image segmentation. It captures contextual information through the symmetrical structure of the encoder and decoder to achieve pixel-level prediction. The input of the U-Net model is the tumor image and its corresponding benign or malignant tumor label.

3. The method for classifying benign and malignant tumors according to claim 2, characterized in that: The U-Net model includes four downsampling modules in the encoder part, each module consists of two 3x3 convolutional layers and a 2x2 maxpooling layer; the decoder part includes four upsampling modules, each module includes a deconvolution layer, a feature concatenation layer and two 3x3 convolutional layers in sequence.

4. The method for classifying benign and malignant tumors according to claim 1, characterized in that: The feature extraction and classification of the screened patient image data specifically includes the following steps: According to the intelligent segmentation results obtained above, the optimal N slices are selected according to the principle that the area of ​​interest accounts for the largest proportion in the image. These N slices are cut into global slices and local slices according to the regions of interest; The local slices and the partial slices are simultaneously sent to the global-hybrid-local network to classify the benign and malignant tumors, so as to obtain a benign and malignant tumor prediction result based on the image data.

5. The method for classifying benign and malignant tumors according to claim 1, characterized in that: The global branch and the local branch use a residual connection network as the backbone network, and use a multi-scale module and a space-channel attention module, wherein the multi-scale module is used to enhance the network's ability to express multi-scale information of tumor images, and the space-channel attention module is used to improve the accuracy of feature extraction; the hybrid branch uses an attention feature fusion method to fuse the features of the global branch and the local branch.

6. The method for classifying benign and malignant tumors according to claim 5, characterized in that: The residual connection network introduces jump connections between network layers, so that the deep network can effectively alleviate the gradient vanishing problem during training. The global branch and the local branch respectively use two types of data of different sizes, namely global image and local image, as inputs of the residual connection network; the multi-scale module uses feature maps of multiple scales in each residual block for fusion, thereby enhancing the network's ability to express multi-scale information of tumor images.

7. The method for classifying benign and malignant tumors according to claim 5, characterized in that: The spatial-channel attention module combines the methods of spatial attention and channel attention; the spatial attention uses the feature map after channel compression to perform convolution operations to generate spatial weights, emphasize the important areas in the input feature map, and suppress the unimportant parts; the channel attention generates channel weights through global average pooling and global maximum pooling to reflect the importance of each channel, and then generates weights through the Sigmoid activation function to proportionally adjust the channel response of the input feature map.

8. The method for classifying benign and malignant tumors according to claim 5, characterized in that: The attention feature fusion method aggregates multi-scale contextual information along the channel dimension, which can simultaneously focus on the global and local semantic information of tumor image distribution and alleviate the problems caused by scale changes and small object instances.

9. The method for classifying benign and malignant tumors according to claim 1, characterized in that: The feature fusion of the global-hybrid-local network is specifically as follows: the feature maps obtained from each of the three branches are transformed into one-dimensional vectors and concatenated, and finally sent to a fully connected classifier to obtain the prediction probability value of benign and malignant tumors in imaging.

10. A tumor benign and malignant classification system based on multimodal data fusion and deep learning, characterized by ,, characterized in that it includes: A data acquisition module, used to acquire patient imaging data and laboratory data; The image segmentation module uses the U-Net model to segment the organs and tumor areas of patient image data and screen out the slices of the region of interest containing tumor diagnostic information; An imaging classification module is used to extract and classify features of the screened patient image slice data to obtain a tumor benign or malignant prediction result. The imaging classification module includes a global-hybrid-local network for extracting and classifying features of imaging data at different scales. The multimodal detection module is used to input the benign and malignant prediction results of the image and the patient's clinical characteristics into the integrated learning classification model to obtain the multimodal tumor benign and malignant prediction results; The interpretable module is used to interpret the visual results of various multimodal input elements and different parts of the image, including attribution analysis using the SHAP method and occlusion-based methods.

11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 9 when executing a program stored in a memory.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method steps according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Prediction system and method for postoperative cognitive impairment of gastrointestinal tumor surgery patient

    CN120356679A

  • Multi-modal medical data fusion prediction and interpretable report generation system, method and device, processor and storage medium thereof

    CN121641324A