SMRI image classification method, system and terminal based on multi-modal chart fusion

Through the multimodal chart fusion method, an sMRI image classification model is constructed, and the image characteristics and clinical table characteristics are deeply integrated, which solves the problem of low classification accuracy of sMRI image in the prior art and achieves higher classification accuracy.

CN120088528APending Publication Date: 2025-06-03SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002300.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the complex relationship between clinical data and neuroimaging data, resulting in insufficient accuracy of sMRI image classification.

Method used

The sMRI image classification model is constructed through the data condition enhancement framework, the table Transformer embedding module, the image table cross-fusion module and the full connection layer to achieve the deep fusion of image features and clinical table features.

Benefits of technology

Effectively integrate imaging and clinical table data, improve the accuracy of sMRI image classification, and significantly improve multimodal learning and feature extraction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088528A_ABST
    Figure CN120088528A_ABST
Patent Text Reader

Abstract

The invention discloses an sMRI image classification method and system based on multi-modal chart fusion, and a terminal. The method comprises the steps of obtaining target sMRI image data, target PET image data and target clinical data; the method comprises the following steps: constructing an sMRI image classification model, training the sMRI image classification model to obtain a target model, comprising a data condition enhancement framework, a table Transform embedding module, an image table cross fusion module and a full connection layer; respectively inputting the target sMRI image data and the target PET image data into a data condition enhancement framework for feature extraction and fusion to obtain image features, and inputting the target clinical data into a table Transform embedding module for feature extraction to obtain clinical table features; and inputting the image features and the clinical table features into an image table cross fusion module for feature fusion, and obtaining a classification result of the sMRI image data after the fused features pass through a full connection layer. According to the method, imaging and clinical table data are effectively integrated, and the sMRI image classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to an sMRI image classification method, system, terminal and computer-readable storage medium based on multi-modal chart fusion. Background Art

[0002] Multi-modal data refers to data obtained from different fields or perspectives, which can come from different senses or information presentation forms, such as text, images, audio, video, etc. Multi-modal data fusion is the core of this technology, which can be achieved through data-level fusion, feature-level fusion or decision-level fusion. With the development of technology, various multi-modal data have begun to be used to guide the classification of sMRI (structural magnetic resonance imaging) images to improve the accuracy of sMRI image classification.

[0003] However, traditional methods often fail to fully utilize the complex relationship between clinical data and neuroimaging, thus limiting the effective utilization of complementary information. In structural and functional magnetic resonance imaging sMRI and positron emission tomography PET data, common spatial and channel feature extraction methods are difficult to capture deeper interactions. In addition, these methods often ignore the key relationships between different modalities.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0005] The main object of the present invention is to provide an sMRI image classification method, system, terminal and computer-readable storage medium based on multi-modal chart fusion, aiming to solve the problem that the existing technology cannot achieve robust interaction and deep integration between clinical data and imaging data, resulting in insufficient accuracy of sMRI image classification.

[0006] To achieve the above object, the present invention provides an sMRI image classification method based on multi-modal chart fusion, and the sMRI image classification method based on multi-modal chart fusion includes the following steps:

[0007] Obtain sMRI image data, PET image data and clinical data of a target object, and preprocess the sMRI image data, the PET image data and the clinical data to obtain target sMRI image data, target PET image data and target clinical data;

[0008] Construct an sMRI image classification model, train and test the sMRI image classification model to obtain a target model, where the target model includes: a data conditional augmentation framework, a tabular Transformer embedding module, an image-table cross-fusion module, and a fully connected layer;

[0009] Input the target sMRI image data and the target PET image data into different branches of the data conditional augmentation framework respectively for feature extraction and fusion to obtain image features, and input the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features;

[0010] Input the image features and the clinical tabular features into the image-table cross-fusion module for feature fusion to obtain fusion features, and after passing the fusion features through the fully connected layer, obtain the classification result of the sMRI image data.

[0011] Optionally, in the sMRI image classification method based on multi-modal chart fusion, before training and testing the sMRI image classification model, it further includes:

[0012] Obtain historical data, where the historical data includes historical sMRI image data, historical PET image data, historical clinical data, and the classification result corresponding to the historical sMRI image data;

[0013] Use the historical sMRI image data, the historical PET image data, and the historical clinical data as training samples, and use the classification result corresponding to the historical sMRI image data as a label to construct a data set;

[0014] Divide the data set into a training set, a test set, and a validation set according to a preset ratio. The training set is used to train the sMRI image classification model, the test set is used to evaluate the sMRI image classification model in each round of training, and the validation set is used to evaluate the trained sMRI image classification model.

[0015] Optionally, in the sMRI image classification method based on multi-modal chart fusion, the data conditional augmentation framework is a multi-modal framework based on 3DResNet50, and the data conditional augmentation framework includes an sMRI feature extraction branch, a PET feature extraction branch, and a feature fusion module.

[0016] Optionally, in the sMRI image classification method based on multi-modal chart fusion, the step of inputting the target sMRI image data and the target PET image data into different branches of the data conditional augmentation framework respectively for feature extraction and fusion to obtain image features specifically includes:

[0017] Input the target sMRI image data and the target PET image data into the sMRI feature extraction branch and the PET feature extraction branch respectively for upsampling. After reducing the dimensions of the upsampling results, perform feature extraction through a residual layer to obtain sMRI features and PET features;

[0018] Input the sMRI features and the PET features into the feature fusion module for global average pooling and max pooling operations to obtain the first average pooling result and the first max pooling result of the sMRI features, as well as the second average pooling result and the second max pooling result of the PET features;

[0019] Concatenate the first max pooling result and the second average pooling result to obtain a first aggregated feature. Input the first aggregated feature into a one-dimensional convolution for calculation to obtain a first sMRI spatial weight and a first PET spatial weight

[0020] Concatenate the second max pooling result and the first average pooling result to obtain a second aggregated feature. Input the second aggregated feature into a two-dimensional convolution for calculation to obtain a first sMRI channel weight and a first PET channel weight

[0021] Normalize the first sMRI spatial weight and the first PET spatial weight respectively to obtain a second sMRI spatial weight and a second PET spatial weight

[0022] Normalize the first sMRI channel weight and the first PET channel weight respectively to obtain a second sMRI channel weight and a second PET channel weight

[0023] Add the second sMRI spatial weight and the second sMRI channel weight together, multiply the result of the addition by the sMRI features to obtain the target sMRI features. Add the second PET spatial weight and the second PET channel weight together, multiply the result of the addition by the PET features to obtain the target PET features;

[0024] Process the target sMRI features and the target PET features using a cross-attention mechanism to output the image features.

[0025] Optionally, in the sMRI image classification method based on multi-modal graph fusion, wherein the first sMRI spatial weight and the first PET spatial weight are respectively normalized to obtain a second sMRI spatial weight and a second PET spatial weight Specifically includes:

[0026] Perform normalization calculations based on the first sMRI spatial weight and the first PET spatial weight to obtain a second sMRI spatial weight and a second PET spatial weight

[0027]

[0028] Optionally, in the sMRI image classification method based on multi-modal graph fusion, wherein the first sMRI channel weight and the first PET channel weight are respectively normalized to obtain a second sMRI channel weight and a second PET channel weight Specifically includes:

[0029] Perform normalization calculations based on the first sMRI channel weight and the first PET channel weight to obtain a second sMRI channel weight and a second PET channel weight

[0030]

[0031] Optionally, in the sMRI image classification method based on multi-modal graph fusion, wherein the step of inputting the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features specifically includes:

[0032] Input the target clinical data into the tabular Transformer embedding module for digital input processing and classification input processing respectively, input the processing result of the digital input into a linear layer for linear transformation to obtain linear features, and input the processing result of the classification input into an embedding layer for feature extraction to obtain embedding features;

[0033] Input the embedding feature and the linear feature into the embedding column for processing, and add position information to the processing result of the embedding column to obtain a position feature. Pass the position feature through a self-attention model and a feed-forward network for feature extraction to obtain a clinical table feature.

[0034] In addition, to achieve the above object, the present invention also provides an sMRI image classification system based on multi-modal chart fusion, wherein the sMRI image classification system based on multi-modal chart fusion includes:

[0035] A target data acquisition module, configured to acquire sMRI image data, PET image data, and clinical data of a target object, and perform preprocessing on the sMRI image data, the PET image data, and the clinical data to obtain target sMRI image data, target PET image data, and target clinical data;

[0036] A target model construction module, configured to construct an sMRI image classification model, train and test the sMRI image classification model to obtain a target model, and the target model includes: a data conditional enhancement framework, a table Transformer embedding module, an image-table cross-fusion module, and a fully connected layer;

[0037] A feature extraction module, configured to input the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion to obtain image features, and input the target clinical data into the table Transformer embedding module for feature extraction to obtain clinical table features;

[0038] A feature fusion and classification module, configured to input the image features and the clinical table features into the image-table cross-fusion module for feature fusion to obtain fusion features, and after passing the fusion features through the fully connected layer, obtain the classification result of the sMRI image data.

[0039] In addition, to achieve the above object, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an sMRI image classification program based on multi-modal chart fusion stored on the memory and executable on the processor. When the sMRI image classification program based on multi-modal chart fusion is executed by the processor, the steps of the sMRI image classification method based on multi-modal chart fusion described above are implemented.

[0040] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an sMRI image classification program based on multimodal chart fusion, and when the sMRI image classification program based on multimodal chart fusion is executed by a processor, the steps of the sMRI image classification method based on multimodal chart fusion as described above are implemented.

[0041] In the present invention, target sMRI image data, target PET image data, and target clinical data are acquired. An sMRI image classification model is constructed, and the sMRI image classification model is trained to obtain a target model, including: a data conditional augmentation framework, a tabular Transformer embedding module, an image-table cross-fusion module, and a fully connected layer; the target sMRI image data and the target PET image data are respectively input into the data conditional augmentation framework for feature extraction and fusion to obtain image features, and the target clinical data is input into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features; the image features and the clinical tabular features are input into the image-table cross-fusion module for feature fusion, and after the fused features pass through the fully connected layer, the classification result of the sMRI image data is obtained. The present invention effectively integrates imaging and clinical tabular data and improves the accuracy of sMRI image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of a preferred embodiment of the sMRI image classification method based on multimodal chart fusion of the present invention;

[0043] Figure 2 is an overall architecture diagram of the target model in the sMRI image classification method based on multimodal chart fusion of the present invention;

[0044] Figure 3 is a schematic diagram of the data conditional augmentation framework in the sMRI image classification method based on multimodal chart fusion of the present invention;

[0045] Figure 4 is a structural diagram of a preferred embodiment of the sMRI image classification system based on multimodal chart fusion of the present invention;

[0046] Figure 5 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present application provides an sMRI image classification method, system, and terminal based on multimodal chart fusion. To make the purpose, technical solution, and effect of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0049] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0050] The sMRI image classification method based on multimodal graph fusion according to a preferred embodiment of the present invention, as Figure 1 shown, the sMRI image classification method based on multimodal graph fusion includes the following steps:

[0051] Step S10: Obtain sMRI image data, PET image data, and clinical data of a target object, and preprocess the sMRI image data, the PET image data, and the clinical data to obtain target sMRI image data, target PET image data, and target clinical data.

[0052] Specifically, sMRI data refers to Structural Magnetic Resonance Imaging (sMRI) data, which is a medical imaging technology used to generate detailed images of the internal structures of the human body. sMRI obtains structural information of the brain or other body parts by using a strong magnetic field and radio waves and is commonly used in the study of brain structure and function. PET data refers to data obtained through Positron Emission Tomography (PET) technology. PET is a functional imaging technology that uses the distribution of radioactive tracers in the living body to image molecules and organs in the living body by detecting photon pairs generated by annihilation. PET data processing involves preprocessing a large amount of raw data, image reconstruction, and quantitative analysis to extract biological and physiological function information. Clinical data refers to safety, clinical performance, or effectiveness information generated during the clinical use of medical devices.

[0053] Obtain the sMRI image data, PET image data and clinical data of the target object, use these multi-modal data as the input of the network model, and before inputting into the model, preprocess the sMRI image data, the PET image data and the clinical data to obtain the target sMRI image data, target PET image data and target clinical data. Among them, the preprocessing process includes preprocessing techniques such as random cropping, data cleaning, and data standardization, and finally obtains the target data for inputting into the model.

[0054] It can be understood that the task of the present invention is image classification. Essentially, the image classification task includes two categories of sMRI image data and PET image data. For the sake of simplicity of expression, in this embodiment, the PET image data is used as a supplement to the classification of sMRI image data to achieve the purpose of enhancing the classification effect of sMRI image data.

[0055] Step S20: Construct an sMRI image classification model, train and test the sMRI image classification model to obtain a target model. The target model includes: a data conditional augmentation framework, a tabular Transformer embedding module, an image-table cross-fusion module, and a fully connected layer.

[0056] As Figure 2 shown, the target model is an image-table cross-fusion network (ITCM-Net). The image-table cross-fusion network includes: a data conditional augmentation framework (DCAF), a tabular Transformer embedding module (TTEM), an image-table cross-fusion module (ITCM), and a fully connected layer. The image-table cross-fusion network integrates gray matter structural magnetic resonance imaging (sMRI, Xg), positron emission tomography (PET, Xp), and clinical data in tabular form (Xt) for image classification. Image features are extracted through different branches and fused with tabular features through a cross-attention mechanism to achieve the classification of sMRI images.

[0057] Furthermore, before training and testing the sMRI image classification model, it also includes:

[0058] Obtain historical data, where the historical data includes historical sMRI image data, historical PET image data, historical clinical data, and the classification results corresponding to the historical sMRI image data;

[0059] Use the historical sMRI image data, the historical PET image data, and the historical clinical data as training samples, and use the classification results corresponding to the historical sMRI image data as labels to construct a data set;

[0060] Divide the dataset into a training set, a test set, and a validation set according to a preset ratio. The training set is used to train the sMRI image classification model, the test set is used to evaluate the sMRI image classification model in each round of training, and the validation set is used to evaluate the trained sMRI image classification model.

[0061] It can be understood that before training and testing the MRI image classification model, it is necessary to collect data to construct a dataset for training the model. The dataset is used to train a machine learning model so that it can learn patterns and regularities from the data and provide prediction capabilities. By using a large amount of diverse data to train the model, the accuracy and generalization ability of the model can be improved. In this embodiment, historical data is obtained, and the historical data includes historical sMRI image data, historical PET image data, historical clinical data, and classification results corresponding to the historical sMRI image data, and then a dataset is constructed; then the dataset is divided into a training set, a test set, and a validation set according to a preset ratio (for example, 8:1:1). The training set is used to train the sMRI image classification model, the test set is used to evaluate the sMRI image classification model in each round of training, and the validation set is used to evaluate the trained sMRI image classification model.

[0062] Step S30: Input the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion to obtain image features, and input the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features.

[0063] In this embodiment, the data conditional enhancement framework is a multi-modal framework based on 3DResNet50, and the data conditional enhancement framework includes an sMRI feature extraction branch, a PET feature extraction branch, and a feature fusion module.

[0064] The step of inputting the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion to obtain image features specifically includes:

[0065] Input the target sMRI image data and the target PET image data into the sMRI feature extraction branch and the PET feature extraction branch for upsampling respectively, and after reducing the dimensions of the upsampling results, perform feature extraction through a residual layer to obtain sMRI feature F g and PET feature F p ;

[0066] The obtained sMRI feature F g and the PET feature Fp Input into the feature fusion module for global average pooling and max pooling processing to obtain the sMRI feature F g The first average pooling result and the first max pooling result, as well as the PET feature F p The second average pooling result and the second max pooling result;

[0067] Concatenate the first max pooling result and the second average pooling result to obtain a first aggregated feature, and input the first aggregated feature into a one-dimensional convolution for calculation to obtain a first sMRI spatial weight and a first PET spatial weight

[0068] Concatenate the second max pooling result and the first average pooling result to obtain a second aggregated feature, and input the second aggregated feature into a two-dimensional convolution for calculation to obtain a first sMRI channel weight and a first PET channel weight

[0069] Normalize the first sMRI spatial weight and the first PET spatial weight respectively to obtain a second sMRI spatial weight and a second PET spatial weight

[0070] Normalize the first sMRI channel weight and the first PET channel weight respectively to obtain a second sMRI channel weight and a second PET channel weight

[0071] Add the second sMRI spatial weight and the second sMRI channel weight and multiply the sum by the sMRI feature F g to obtain the target sMRI feature. Add the second PET spatial weight and the second PET channel weight and multiply the sum by the PET feature F p to obtain the target PET feature;

[0072] Use the cross-attention mechanism to process the target sMRI feature and the target PET feature, and output the image feature.

[0073] As Figure 3As shown, it can be understood that the Data Conditioned Augmentation Framework (DCAF) aims to integrate spatial and channel features from different imaging modalities. In sMRI image processing, each imaging modality provides unique and complementary information. DCAF adopts a cross-attention mechanism to preserve the characteristics of each modality while enhancing the interaction between them, thereby improving classification accuracy. In the schematic diagram, the structure of the Data Conditioned Augmentation Framework (DCAF) is particularly shown. In the Data Conditioned Augmentation Framework (DCAF), F g 、F p and F i represent sMRI features, PET features, and fused image features respectively, while F g ' and F p ' are their processed forms (i.e., target sMRI features and target PET features). The Data Conditioned Augmentation Framework applies spatial and channel attention weights W, where γ ∈ {g, p} represents the imaging modality (g for gray matter sMRI, p for PET), and δ ∈ {s, c} represents the attention type (spatial or channel).

[0074] The present invention adopts a multi-modal framework based on 3DResNet50 to process structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) images. The input features are first upsampled and then feature extraction is performed through dimensionality reduction and residual layers. 3DResNet50 sets the channels as [16, 32, 64, 128] to capture multi-scale features, so as to comprehensively extract information from sMRI and PET data, obtaining sMRI features F g and PET features F p . The sMRI features F g and PET features F p are processed through global average pooling and max pooling, and then they are concatenated to form aggregated features.

[0075] Next, a convolutional block is applied to calculate the feature weights of each modality. The first max pooling result and the second average pooling result are concatenated to obtain a first aggregated feature, and the first aggregated feature is input into a one-dimensional convolution for calculation to obtain a first sMRI spatial weight and a first PET spatial weight which can be expressed as:

[0076]

[0077] The second max pooling result and the first average pooling result are concatenated to obtain a second aggregated feature, and the second aggregated feature is input into a two-dimensional convolution for calculation to obtain a first sMRI channel weight and a first PET channel weight which can be expressed as:

[0078]

[0079] Among them, Conv1 represents one-dimensional convolution, Conv2 represents two-dimensional convolution, Concat() represents a concatenation operation, Avg() represents an average pooling operation, and Max() represents a max pooling operation.

[0080] Furthermore, the second sMRI spatial weight is obtained through a softmax operation The second PET spatial weight The second sMRI channel weight and the second PET channel weight Thereby retaining the key information from two time points.

[0081] Specifically, according to the first sMRI spatial weight and the first PET spatial weight Normalization calculations are performed to obtain the second sMRI spatial weight and the second PET spatial weight

[0082] According to the first sMRI channel weight and the first PET channel weight Normalization calculations are performed to obtain the second sMRI channel weight and the second PET channel weight where e is the natural constant.

[0083] Furthermore, the cross-attention mechanism (CA, CrossAttention) is used to process the target sMRI feature and the target PET feature. In the cross-attention mechanism, the processed target sMRI feature is used as the query (query vector), while the processed target PET feature serves as the key (key vector) and value (value vector), which promotes the interaction between the two modalities. The finally output image feature is expressed as:

[0084] F i = CrossAttention(F g ', F p ');

[0085] where F i represents the image feature, F g ' represents the target sMRI feature, and F p ' represents the target PET feature.

[0086] Even further, as Figure 2As shown, inputting the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features specifically includes:

[0087] Input the target clinical data into the tabular Transformer embedding module for digital input processing and categorical input processing respectively, input the processing result of digital input into a linear layer for linear transformation to obtain linear features, and input the processing result of categorical input into an embedding layer for feature extraction to obtain embedding features;

[0088] Input the embedding features and the linear features into an embedding column for processing, add position information to the processing result of the embedding column to obtain position features, and extract features from the position features through a self-attention model and a feed-forward network to obtain clinical tabular features.

[0089] Step S40: Input the image features and the clinical tabular features into the image-tabular cross-fusion module for feature fusion to obtain fusion features, and after passing the fusion features through the fully connected layer, obtain the classification result of the sMRI image data.

[0090] As Figure 2 shown, in this embodiment, an image-tabular cross-fusion module based on cross-attention is designed to process imaging data and clinical tabular data simultaneously. Through this mechanism, during the fusion process, the image feature F i and the clinical tabular feature F t can pay attention to each other, ensuring that the model can fully utilize the key information in each modality.

[0091] The image-tabular cross-fusion module consists of multiple layers. Each layer contains a cross-attention mechanism, a self-attention mechanism, and a feed-forward network MLP, and layer normalization is incorporated to stabilize the training process. In the image-tabular cross-fusion module, a single token in the clinical tabular data is used to selectively pay attention to the corresponding image features in different modalities, thereby enhancing the feature alignment of specific modalities. Among them, the calculation process of the cross-attention mechanism can be described as follows:

[0092]

[0093] Among them, CrossAttention(Q, K, V) represents the output result of the cross-attention mechanism, Q, K, and V respectively represent the linear transformation of the clinical tabular feature, the image feature, and the weighted image feature value, T represents matrix transpose, represents the scaling factor, which is used to avoid the dot product value being too large and affecting the gradient stability.

[0094] This mechanism enables the image and clinical table features to enhance each other in each layer, associates the table markers with the corresponding image information, and thus achieves effective fusion. This cross-modal interaction creates a richer representation for sMRI image classification.

[0095] It can be seen that the present invention proposes an innovative image-table cross-fusion network, which adopts a multi-layer cross-attention and self-attention mechanism to promote the robust interaction between clinical data and imaging data, thereby enhancing the multi-modal learning and feature extraction capabilities. Experimental results show that the network framework of the present invention significantly improves the accuracy of sMRI image classification.

[0096] Furthermore, as Figure 4 shown, based on the above sMRI image classification method based on multi-modal chart fusion, the present invention also correspondingly provides an sMRI image classification system based on multi-modal chart fusion, wherein the sMRI image classification system based on multi-modal chart fusion includes:

[0097] A target data acquisition module 51, configured to acquire sMRI image data, PET image data, and clinical data of a target object, preprocess the sMRI image data, the PET image data, and the clinical data to obtain target sMRI image data, target PET image data, and target clinical data;

[0098] A target model construction module 52, configured to construct an sMRI image classification model, train and test the sMRI image classification model to obtain a target model, and the target model includes: a data conditional enhancement framework, a table Transformer embedding module, an image-table cross-fusion module, and a fully connected layer;

[0099] A feature extraction module 53, configured to input the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion to obtain image features, and input the target clinical data into the table Transformer embedding module for feature extraction to obtain clinical table features;

[0100] A feature fusion and classification module 54, configured to input the image features and the clinical table features into the image-table cross-fusion module for feature fusion to obtain fusion features, and after passing the fusion features through the fully connected layer, obtain the classification result of the sMRI image data.

[0101] Furthermore, as Figure 5 shown, based on the above sMRI image classification method and system based on multi-modal chart fusion, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30.Figure 5 Only some components of the terminal are shown, but it should be understood that it is not necessary to implement all the shown components, and more or fewer components can be implemented alternatively.

[0102] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in some other embodiments, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a sMRI image classification program 40 based on multimodal graph fusion is stored on the memory 20, and the sMRI image classification program 40 based on multimodal graph fusion can be executed by the processor 10, so as to implement the sMRI image classification method based on multimodal graph fusion in this application.

[0103] The processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips in some embodiments, and is used to run the program codes stored in the memory 20 or process data, such as executing the sMRI image classification method based on multimodal graph fusion, etc.

[0104] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other through a system bus.

[0105] In one embodiment, when the processor 10 executes the sMRI image classification program 40 based on multimodal graph fusion in the memory 20, the following steps are implemented:

[0106] Obtain sMRI image data, PET image data and clinical data of a target object, preprocess the sMRI image data, the PET image data and the clinical data to obtain target sMRI image data, target PET image data and target clinical data;

[0107] Construct an sMRI image classification model, train and test the sMRI image classification model to obtain a target model, where the target model includes: a data conditional augmentation framework, a tabular Transformer embedding module, an image-table cross-fusion module, and a fully connected layer;

[0108] Input the target sMRI image data and the target PET image data into different branches of the data conditional augmentation framework for feature extraction and fusion to obtain image features, and input the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features;

[0109] Input the image features and the clinical tabular features into the image-table cross-fusion module for feature fusion to obtain fused features, and after passing the fused features through the fully connected layer, obtain the classification result of the sMRI image data.

[0110] Among them, before training and testing the sMRI image classification model, it further includes:

[0111] Obtain historical data, where the historical data includes historical sMRI image data, historical PET image data, historical clinical data, and the classification results corresponding to the historical sMRI image data;

[0112] Use the historical sMRI image data, the historical PET image data, and the historical clinical data as training samples, and use the classification results corresponding to the historical sMRI image data as labels to construct a data set;

[0113] Divide the data set into a training set, a test set, and a validation set according to a preset ratio. The training set is used to train the sMRI image classification model, the test set is used to evaluate the sMRI image classification model in each round of training, and the validation set is used to evaluate the trained sMRI image classification model.

[0114] Among them, the data conditional augmentation framework is a multi-modal framework based on 3DResNet50, and the data conditional augmentation framework includes an sMRI feature extraction branch, a PET feature extraction branch, and a feature fusion module.

[0115] Among them, the step of inputting the target sMRI image data and the target PET image data into different branches of the data conditional augmentation framework for feature extraction and fusion to obtain image features specifically includes:

[0116] Input the target sMRI image data and the target PET image data into the sMRI feature extraction branch and the PET feature extraction branch respectively for upsampling. After reducing the dimensions of the upsampling results, perform feature extraction through the residual layer to obtain sMRI features and PET features;

[0117] Input the sMRI features and the PET features into the feature fusion module for global average pooling and max pooling operations to obtain the first average pooling result and the first max pooling result of the sMRI features, as well as the second average pooling result and the second max pooling result of the PET features;

[0118] Concatenate the first max pooling result and the second average pooling result to obtain a first aggregated feature. Input the first aggregated feature into a one-dimensional convolution for calculation to obtain a first sMRI spatial weight and a first PET spatial weight

[0119] Concatenate the second max pooling result and the first average pooling result to obtain a second aggregated feature. Input the second aggregated feature into a two-dimensional convolution for calculation to obtain a first sMRI channel weight and a first PET channel weight

[0120] Normalize the first sMRI spatial weight and the first PET spatial weight respectively to obtain a second sMRI spatial weight and a second PET spatial weight

[0121] Normalize the first sMRI channel weight and the first PET channel weight respectively to obtain a second sMRI channel weight and a second PET channel weight

[0122] Add the second sMRI spatial weight and the second sMRI channel weight Multiply the sum by the sMRI features to obtain the target sMRI features. Add the second PET spatial weight and the second PET channel weight Multiply the sum by the PET features to obtain the target PET features;

[0123] Process the target sMRI features and the target PET features using a cross-attention mechanism to output the image features.

[0124] Among them, the normalization of the first sMRI spatial weight and the first PET spatial weight respectively to obtain the second sMRI spatial weight and the second PET spatial weight Specifically includes:

[0125] According to the first sMRI spatial weight and the first PET spatial weight Perform normalization calculations to obtain the second sMRI spatial weight and the second PET spatial weight

[0126]

[0127] Among them, the normalization of the first sMRI channel weight and the first PET channel weight respectively to obtain the second sMRI channel weight and the second PET channel weight Specifically includes:

[0128] According to the first sMRI channel weight and the first PET channel weight Perform normalization calculations to obtain the second sMRI channel weight and the second PET channel weight

[0129]

[0130] Among them, the input of the target clinical data into the tabular Transformer embedding module for feature extraction to obtain clinical tabular features specifically includes:

[0131] Input the target clinical data into the tabular Transformer embedding module for digital input processing and classification input processing respectively, input the processing results of the digital input into a linear layer for linear transformation to obtain linear features, and input the processing results of the classification input into an embedding layer for feature extraction to obtain embedding features;

[0132] Input the embedding feature and the linear feature into an embedding column for processing, and add position information to the processing result of the embedding column to obtain a position feature. Pass the position feature through a self-attention model and a feed-forward network for feature extraction to obtain a clinical table feature.

[0133] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program for sMRI image classification based on multi-modal chart fusion. When the program for sMRI image classification based on multi-modal chart fusion is executed by a processor, the steps of the sMRI image classification method based on multi-modal chart fusion as described above are implemented.

[0134] In summary, the present invention proposes a method, a system and a terminal for sMRI image classification based on multi-modal chart fusion. The method includes: obtaining target sMRI image data, target PET image data and target clinical data. Constructing an sMRI image classification model, and training the sMRI image classification model to obtain a target model, including: a data conditional enhancement framework, a table Transformer embedding module, an image-table cross-fusion module and a fully-connected layer; inputting the target sMRI image data and the target PET image data into the data conditional enhancement framework for feature extraction and fusion respectively to obtain an image feature, and inputting the target clinical data into the table Transformer embedding module for feature extraction to obtain a clinical table feature; inputting the image feature and the clinical table feature into the image-table cross-fusion module for feature fusion, and passing the fused feature through the fully-connected layer to obtain a classification result of the sMRI image data. The present invention proposes an innovative image-table cross-fusion network, which adopts a multi-layer cross-attention and self-attention mechanism, promotes a robust interaction between clinical data and imaging data, effectively integrates imaging and clinical table data, thereby enhancing the multi-modal learning and feature extraction ability, and significantly improving the accuracy of sMRI image classification.

[0135] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.

[0136] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0137] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A sMRI image classification method based on multimodal graph fusion, characterized in that: The sMRI image classification method based on multimodal graph fusion includes: Acquiring sMRI image data, PET image data, and clinical data of a target object, and preprocessing the sMRI image data, the PET image data, and the clinical data to obtain target sMRI image data, target PET image data, and target clinical data; Constructing an sMRI image classification model, training and testing the sMRI image classification model to obtain a target model, wherein the target model includes: a data conditional enhancement framework, a table Transformer embedding module, an image table cross fusion module, and a fully connected layer; Input the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion, respectively, to obtain image features, and input the target clinical data into the table Transformer embedding module for feature extraction, to obtain clinical table features; The image features and the clinical table features are input into the image table cross fusion module for feature fusion to obtain fusion features, and the fusion features are passed through the fully connected layer to obtain the classification results of the sMRI image data.

2. The sMRI image classification method based on multimodal graph fusion according to claim 1, characterized in that: The sMRI image classification model is trained and tested, and the method further comprises: Acquiring historical data, the historical data including historical sMRI image data, historical PET image data, historical clinical data, and classification results corresponding to the historical sMRI image data; The historical sMRI image data, the historical PET image data and the historical clinical data are used as training samples, and the classification results corresponding to the historical sMRI image data are used as labels to construct a data set; The data set is divided into a training set, a test set and a validation set according to a preset ratio, the training set is used to train the sMRI image classification model, the test set is used to evaluate the sMRI image classification model in each round of training, and the validation set is used to evaluate the trained sMRI image classification model.

3. The sMRI image classification method based on multimodal graph fusion according to claim 1, characterized in that: The data conditional enhancement framework is a multimodal framework based on 3DResNet50, and the data conditional enhancement framework includes an sMRI feature extraction branch, a PET feature extraction branch and a feature fusion module.

4. The sMRI image classification method based on multimodal graph fusion according to claim 3, characterized in that: The target sMRI image data and the target PET image data are respectively input into different branches of the data conditional enhancement framework for feature extraction and fusion to obtain image features, specifically including: The target sMRI image data and the target PET image data are respectively input into the sMRI feature extraction branch and the PET feature extraction branch for upsampling, and after the upsampling result is dimensionally reduced, feature extraction is performed through the residual layer to obtain sMRI features and PET features; Inputting the sMRI features and the PET features into the feature fusion module for global average pooling and maximum pooling processing to obtain a first average pooling result and a first maximum pooling result of the sMRI features, and a second average pooling result and a second maximum pooling result of the PET features; The first maximum pooling result and the second average pooling result are concatenated to obtain a first aggregate feature, and the first aggregate feature is input into a one-dimensional convolution for calculation to obtain a first sMRI spatial weight and the first PET spatial weight The second maximum pooling result and the first average pooling result are concatenated to obtain a second aggregate feature, and the second aggregate feature is input into a two-dimensional convolution for calculation to obtain a first sMRI channel weight. and the first PET channel weight The first sMRI spatial weight and the first PET spatial weight Normalize them separately to get the second sMRI spatial weight and the second PET spatial weight Weighting of the first sMRI channel and the first PET channel weight Normalize them separately to get the second sMRI channel weight and the second PET channel weight The second sMRI spatial weight and the second sMRI channel weight Add, multiply the result of the addition by the sMRI feature to obtain the target sMRI feature, and add the second PET spatial weight and the second PET channel weight Add, multiply the result of the addition by the PET feature to obtain the target PET feature; The target sMRI features and the target PET features are processed using a cross-attention mechanism to output the image features.

5. The sMRI image classification method based on multimodal graph fusion according to claim 4, characterized in that: The first sMRI spatial weight and the first PET spatial weight Normalize them separately to get the second sMRI spatial weight and the second PET spatial weight Specifically include: According to the first sMRI spatial weight and the first PET spatial weight Perform normalization calculation to obtain the second sMRI spatial weight and the second PET spatial weight 6. The sMRI image classification method based on multimodal graph fusion according to claim 4, characterized in that: The weight of the first sMRI channel and the first PET channel weight Normalize them separately to get the second sMRI channel weight and the second PET channel weight Specifically include: According to the first sMRI channel weight and the first PET channel weight Perform normalization calculation to obtain the second sMRI channel weight and the second PET channel weight 7. The sMRI image classification method based on multimodal graph fusion according to claim 1, characterized in that: The target clinical data is input into the table Transformer embedding module for feature extraction to obtain clinical table features, specifically including: The target clinical data are respectively input into the table Transformer embedding module for digital input processing and classification input processing, and the processing result of the digital input is input into the linear layer for linear transformation to obtain linear features, and the processing result of the classification input is input into the embedding layer for feature extraction to obtain embedded features; The embedding features and the linear features are input into the embedding column for processing, and position information is added to the processing result of the embedding column to obtain position features, and the position features are extracted through a self-attention model and a feedforward network to obtain clinical table features.

8. A sMRI image classification system based on multimodal graph fusion, characterized in that: The sMRI image classification system based on multimodal graph fusion includes: a target data acquisition module, used to acquire sMRI image data, PET image data and clinical data of a target object, and preprocess the sMRI image data, the PET image data and the clinical data to obtain target sMRI image data, target PET image data and target clinical data; A target model building module is used to build an sMRI image classification model, train and test the sMRI image classification model to obtain a target model, wherein the target model includes: a data condition enhancement framework, a table Transformer embedding module, an image table cross fusion module, and a fully connected layer; A feature extraction module, used for inputting the target sMRI image data and the target PET image data into different branches of the data conditional enhancement framework for feature extraction and fusion, respectively, to obtain image features, and inputting the target clinical data into the table Transformer embedding module for feature extraction, to obtain clinical table features; The feature fusion classification module is used to input the image features and the clinical table features into the image table cross fusion module for feature fusion to obtain fusion features, and after passing the fusion features through the fully connected layer, obtain the classification result of the sMRI image data.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an sMRI image classification program based on multimodal chart fusion stored in the memory and executable on the processor. When the sMRI image classification program based on multimodal chart fusion is executed by the processor, the steps of the sMRI image classification method based on multimodal chart fusion as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an sMRI image classification program based on multimodal chart fusion, and when the sMRI image classification program based on multimodal chart fusion is executed by a processor, the steps of the sMRI image classification method based on multimodal chart fusion as described in any one of claims 1 to 7 are implemented.