Classification model, classification method and device based on multi-modal multi-scale fusion
Through the multimodal multi-scale fusion classification model, the misjudgment problem of small and medium-sized lesions detection in traditional methods is solved, and the efficient fusion of medical images and health data is achieved, and the accuracy and reliability of disease diagnosis is improved.
Patent Information
- Application Number
- CN202510685355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional deep learning methods are prone to neglect or misjudgment when detecting small-sized lesions such as squamous cell carcinoma and pulmonary embolism, and the differences between different data modes make it difficult to effectively integrate multimodal data fusion, affecting the accuracy of disease diagnosis.
A multimodal multi-scale fusion classification model is adopted, and the characteristics of medical images and health data are extracted through visual encoder and table encoder, and feature fusion is combined with a multi-scale convolutional attention module and a bidirectional feedback propagation unit to generate multi-scale fusion features, and finally the diagnostic results are output through the classification module.
It significantly improves the detection rate of micro lesions, realizes efficient coordination between images and health data, deeply explores potential correlations, and improves the accuracy and reliability of disease diagnosis.
Smart Images

Figure CN120565112A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a classification model, classification method, and device based on multimodal and multiscale fusion. Background Art
[0002] Computer-aided diagnosis (CAD) plays a vital role in modern clinical practice, significantly improving the efficiency and accuracy of disease detection. In recent years, deep learning technology has been widely used in the field of medical diagnosis, leveraging multiple data modalities, such as medical imaging and electronic health records, to provide comprehensive patient assessments.
[0003] However, traditional deep learning methods still face numerous challenges in detecting certain specific lesions, such as squamous cell carcinoma and pulmonary embolism. These lesions are typically small in size in images and easily overlooked or misjudged by existing models. Furthermore, different data modalities (such as medical imaging and electronic health records) differ significantly in representation space, temporal structure, and level of abstraction. This complicates effective multimodal data fusion in clinical decision-making, making it difficult to fully integrate the advantages of each modality for aided diagnosis. Summary of the Invention
[0004] Based on this, it is necessary to provide a multimodal and multi-scale fusion classification model, classification method, device and medium to address the above technical problems, so as to solve at least one problem existing in the above-mentioned prior art.
[0005] First, a classification method based on multimodal and multiscale fusion is provided, including:
[0006] Obtaining medical images to be tested and health data to be tested;
[0007] Extracting features from the medical image to be detected and the health data to be detected to obtain visual features and health features;
[0008] Performing multi-scale operations on the visual features and health features to obtain multi-scale feature pairs, performing feature fusion on the feature pairs at each scale, and performing bidirectional multi-scale fusion on the feature pairs after feature fusion to obtain multi-scale fusion features;
[0009] Based on the multi-scale fusion features, a classification result is obtained.
[0010] In a second aspect, a classification model based on multimodal multiscale fusion is provided, wherein the classification model is used to implement the classification method based on multimodal multiscale fusion as described in the first aspect above, and the classification model includes:
[0011] Visual encoder, table encoder, fusion module and classification module;
[0012] The visual encoder includes a feature extraction module, a feature pyramid operation module, a three-dimensional multi-scale convolutional attention module and a bidirectional feedback propagation unit. The three-dimensional multi-scale convolutional attention module includes a channel attention module, a spatial attention module and a deep convolutional fusion module connected in sequence, which is used to extract features from the medical image to be detected to obtain visual features;
[0013] The table encoder is a Kolmogorov-Arnold network, which is used to extract features from the health data to be detected to obtain health features;
[0014] The fusion module includes an inverted feature pyramid operation module, a cross attention module and a bidirectional multi-scale fusion module connected in sequence, and is used to fuse the visual features and health features to obtain multi-scale fusion features;
[0015] The classification module is used to perform a classification operation based on the multi-scale fusion features to output a classification result.
[0016] In a third aspect, a classification device based on multimodal and multiscale fusion is provided, comprising:
[0017] A data acquisition unit for detecting, used to acquire medical images and health data for detecting;
[0018] a feature extraction unit, configured to extract features from the medical image to be detected and the health data to be detected, respectively, to obtain visual features and health features;
[0019] a feature fusion unit, configured to perform multi-scale operations on the visual features and health features to obtain multi-scale feature pairs, perform feature fusion on the feature pairs at each scale, and perform bidirectional multi-scale fusion on the feature pairs after feature fusion to obtain multi-scale fusion features;
[0020] The classification unit is used to obtain a classification result based on the multi-scale fusion feature.
[0021] In a fourth aspect, a readable storage medium is provided, wherein the readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the classification method based on multimodal multi-scale fusion as described above are implemented.
[0022] The above-mentioned classification model, classification method and device based on multimodal multiscale fusion, and its implementation method include: obtaining a medical image to be detected and health data to be detected; performing feature extraction on the medical image to be detected and the health data to be detected respectively to obtain visual features and health features; performing multiscale operations on the visual features and health features to obtain multiscale feature pairs, performing feature fusion on the feature pairs of each scale, and performing bidirectional multiscale fusion on the feature pairs of each scale after feature fusion to obtain multiscale fusion features; based on the multiscale fusion features, a classification result is obtained. In an embodiment of the present application, lesion features are captured from multiple dimensions through feature extraction and multiscale operations, combined with bidirectional multiscale fusion, the perception ability of small lesions (such as squamous cell carcinoma and pulmonary embolism) is enhanced, and the lesion detection rate is significantly improved. The visual features and health features are refined, and the information barriers between modalities are broken through the construction and deep fusion of multiscale feature pairs, so as to achieve efficient collaboration between images and health data and deeply explore the potential associations between data. Through a bidirectional multi-scale fusion mechanism, features at different scales are fully integrated, preserving image details while integrating the semantic knowledge of health data. This makes multi-scale fusion features more expressive and provides richer and more accurate information for disease diagnosis. This effectively improves the accuracy and reliability of disease diagnosis and provides strong support for clinical auxiliary diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 1 is a schematic diagram of a model architecture of a classification model based on multimodal and multiscale fusion in one embodiment of the present application;
[0025] Figure 2 This is a schematic diagram of a model architecture of a visual encoder in one embodiment of the present application;
[0026] Figure 3 This is a schematic diagram of a model architecture of a fusion module in one embodiment of the present application;
[0027] Figure 4 1 is a flow chart of a classification method based on multimodal and multiscale fusion in one embodiment of the present application;
[0028] Figure 5 1 is a schematic diagram of classification results of each model in a comparative experiment in one embodiment of the present application;
[0029] Figure 61 is a structural diagram of a classification device based on multimodal multi-scale fusion in one embodiment of the present application;
[0030] Figure 7 Schematic diagram of a computer device in one embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] The classification method based on multi-modal and multi-scale fusion provided in this embodiment can be applied in Figure 1 In the model architecture of the multimodal and multi-scale fusion classification model, the classification model includes a visual encoder, a table encoder, a fusion module and a classification module.
[0033] It should be noted that this visual encoder uses the Efficient 3D Multi-scale PENet, combining a feature pyramid with an enhanced 3D multi-scale convolutional attention module (E3D-MSCA) to capture both local and global features of medical images. It includes a feature extraction module (PENet encoder), a feature pyramid operation module, a three-dimensional multi-scale convolutional attention module (E3D-MSCA), and a bidirectional feedback propagation unit (BFPU). The PENet encoder is used to encode the medical image to extract visual features at different scales. These visual features are then subjected to a feature pyramid operation and then fed into the E3D-MSCA module for feature fusion. Finally, the bidirectional feedback propagation unit fuses features from multiple scales to obtain multi-scale visual features.
[0034] like Figure 2As shown, the E3D-MSCA module may include a three-dimensional multi-scale convolutional attention module, which includes a 3D channel attention block (CAB), a 3D spatial attention block (SAB) and a 3D deep convolutional fusion block (DCFB) connected in sequence. First, a feature pyramid operation is performed on the medical image to be detected, and then the multi-scale features extracted by the feature pyramid operation are weighted by channel and spatial attention through the 3D multi-scale convolutional attention module, and finally fused through the 3D deep convolutional fusion module.
[0035] The table encoder is a Kolmogorov–Arnold network (KAN), which is used to extract features from the health data to be tested to obtain health features;
[0036] like Figure 3 As shown in the figure, the fusion module includes an inverted feature pyramid operation module, a cross attention module and a bidirectional multi-scale fusion module (BSF) connected in sequence. First, the inverted feature pyramid operation is performed on the extracted visual features and health features, and then the cross attention calculation is performed on the obtained feature pairs to achieve cross-modal fusion. Finally, the two-scale fusion module fuses the semantic and detail information of features of different scales through forward and reverse information flows to generate the final multi-scale visual features.
[0037] The classification module is a classification head, which can be composed of several fully connected layers (FC layers) or convolutional layers. It is used to perform classification operations on the input multi-scale fusion features to output classification results. The classification results may include lung disease categories, such as adenocarcinoma and squamous cell carcinoma.
[0038] In the embodiments of the present application, medical images and health data are cross-modally fused, and features of different scales are captured through an efficient 3D multi-scale convolutional attention module to enhance the recognition of small lesions; a multi-scale cross-attention module is used to break the barriers between modalities and achieve deep interactive fusion; and a dual-scale fusion module is used to integrate features of different linear scales, balance details and semantic information, significantly enhance the multimodal data integration capability, effectively improve the accuracy and reliability of disease diagnosis, and provide strong support for clinical auxiliary diagnosis.
[0039] In one embodiment, if Figure 4 As shown, a classification method based on multimodal and multi-scale fusion is provided, which includes the following steps:
[0040] In step S110, the medical image to be detected and the health data to be detected are obtained;
[0041] The medical image to be tested may be a CT (Computed Tomography) image or a PET (Positron Emission Tomography) image. The health data to be tested refers to the electronic medical record data corresponding to the medical image to be tested, which may include identity information, medical records, medical history, physical examination results, and other health-related data.
[0042] It should be noted that during the training phase of the classification model, a public dataset provided by The Cancer Imaging Archive, named Lung-PET-CT-Dx, can be used. This dataset contains CT or PET scans and tabulated clinical information of 355 patients. The dataset provides a tumor classification label for each patient. Most PET / CT data are stored in DICOM format, and all data are de-identified. The dataset contains 251 adenocarcinoma samples and 61 squamous cell carcinoma samples. Since the number of samples in the adenocarcinoma and squamous cell carcinoma categories in the original dataset is extremely unbalanced, the squamous cell carcinoma category can be oversampled. For example, a random oversampling strategy can be used to supplement the squamous cell carcinoma samples in the training set from 34 to 198, and the number of validation and test sets remains unchanged at 12 and 15, respectively.
[0043] In step S120, feature extraction is performed on the medical image to be detected and the health data to be detected to obtain visual features and health features;
[0044] Optionally, different encoders can be used to extract features from the medical images to be tested and the health data to be tested. The medical images to be tested can first be preprocessed, such as by converting them into grayscale images and performing data enhancement operations such as random flipping and sharpening to obtain enhanced grayscale images. Feature extraction of the enhanced grayscale images can then be performed using a visual encoder to obtain multi-scale visual features. The health data to be tested can first be structured to generate health table data, which can then be digitized based on this health table data. For example, identity information and medical records can be converted into corresponding numerical variables using one-hot encoding. The lung tumor stage (TNM stage) in the table can be converted into numerical variables using a hierarchical assignment method. For example, in the T stage, T0 (no evidence of primary tumor) is assigned a value of 0, and Tis (carcinoma in situ (limited to the bronchial mucosa, without interstitial invasion)) is assigned a value of 1. The digitized data can then be standardized, such as using Z-score standardization. Finally, the standardized values are combined to obtain a health feature representation.
[0045] It should be noted that Figure 3 As shown in the figure, the visual encoder may include a feature extraction module (PENet encoder), a feature pyramid operation module, a three-dimensional multi-scale convolutional attention module (E3D-MSCA), and a bidirectional feedback propagation unit (BFPU). First, the medical image to be detected is encoded using the PENet encoder to extract visual features of different scales. The visual features of different scales are subjected to feature pyramid operation and then input into the E3D-MSCA module for feature fusion. Finally, the bidirectional feedback propagation unit fuses the features from multiple scales to obtain multi-scale visual features, which can be a one-dimensional vector of length 24.
[0046] The table encoder is a Kolmogorov–Arnold network (KAN), which is used to extract features from the health data to be detected to obtain health features. The health features can be a one-dimensional vector with a length of 24.
[0047] In step S130, multi-scale operations are performed on the visual features and health features to obtain multi-scale feature pairs, feature fusion is performed on the feature pairs at each scale, and bidirectional multi-scale fusion is performed on the feature pairs after feature fusion to obtain multi-scale fusion features;
[0048] Optionally, multi-scale operations can be performed on the obtained visual features and health features. For example, based on multi-scale operations such as feature pyramid networks, void convolution, and multi-branch networks, the features are analyzed and processed from different granularities and levels of view to capture information patterns at different scales. Then, cross-modal spatial cross-attention calculations are performed on the feature pairs of each scale obtained, and finally bidirectional multi-scale fusion is performed through a bidirectional multi-scale fusion module to obtain multi-scale fusion features. It should be noted that feature pairs refer to visual-health feature pairs, which can be divided into feature pairs of multiple scales through multi-scale operations, such as 24, 12, and 6 scales.
[0049] In step S140 , a classification result is obtained based on the multi-scale fusion features.
[0050] Optionally, the multi-scale fused features are input into a classification module, which may include multiple fully connected layers or convolutional layers. The input multi-scale fused features are first preprocessed, such as through global pooling and flattening. Then, the features are mapped into unnormalized class scores through fully connected or convolutional layers. Finally, an activation function, such as Softmax, is used to generate a multi-class probability distribution, from which the class with the highest probability is selected as the final classification class.
[0051] It should be noted that for the training stage, the loss value can be calculated based on the final classification label and the true label. If the loss value is greater than the preset loss threshold, the model parameters of the classification model can be optimized, and the next round of iteration can be performed through the training data until the classification model meets the preset convergence conditions, such as the number of iterations reaches the preset number, or the loss value is less than or equal to the preset loss threshold.
[0052] Among them, the loss value can be calculated by the cross entropy loss function, then the classification loss L cls It can be calculated by the following formula:
[0053]
[0054] Where N represents the number of samples, y i represents the true label of sample i, It is the predicted probability value of the model for sample i, ranging between (0, 1).
[0055] In an embodiment of the present application, a classification method based on multimodal multiscale fusion is provided, comprising: obtaining a medical image to be detected and health data to be detected; performing feature extraction on the medical image to be detected and the health data to be detected respectively to obtain visual features and health features; performing multiscale operations on the visual features and health features to obtain multiscale feature pairs, performing feature fusion on the feature pairs of each scale, and performing bidirectional multiscale fusion on the feature pairs of each scale after feature fusion to obtain multiscale fusion features; and obtaining a classification result based on the multiscale fusion features. In an embodiment of the present application, lesion features are captured from multiple dimensions through feature extraction and multiscale operations, and combined with bidirectional multiscale fusion, the perception ability of microlesions (such as squamous cell carcinoma and pulmonary embolism) is enhanced, and the lesion detection rate is significantly improved. The visual features and health features are refined, and the information barriers between modalities are broken through the construction and deep fusion of multiscale feature pairs, and efficient collaboration between images and health data is achieved, and potential correlations between data are deeply mined. Through a bidirectional multi-scale fusion mechanism, features at different scales are fully integrated, preserving image details while integrating the semantic knowledge of health data. This makes multi-scale fusion features more expressive and provides richer and more accurate information for disease diagnosis. This effectively improves the accuracy and reliability of disease diagnosis and provides strong support for clinical auxiliary diagnosis.
[0056] In one embodiment of the present application, the feature extraction of the medical image to be detected and the health data to be detected to obtain visual features and health features includes:
[0057] Extracting features from the medical image to be detected based on a visual encoder to obtain the visual features; and
[0058] The health data to be detected is converted into table data, and feature extraction is performed on the table data based on a table encoder to obtain the health feature.
[0059] Optionally, different encoders can be used to extract features from the medical images to be detected and the health data to be detected. For the medical images to be detected, a visual encoder can be used to extract features. The visual encoder may include a feature extraction module (PENet encoder), a feature pyramid operation module, a three-dimensional multi-scale convolutional attention module (E3D-MSCA), and a bidirectional feedback propagation unit BFPU. First, the medical image to be detected can be encoded by the PENet encoder to extract visual features of different scales, and the visual features of different scales are subjected to feature pyramid operations. Then, the features are input into the E3D-MSCA module for feature fusion. Finally, the features from multiple scales are fused through the bidirectional feedback propagation unit to obtain multi-scale visual features.
[0060] The table encoder is a Kolmogorov–Arnold network (KAN), which performs structured processing on the health data to be tested to obtain corresponding table data, which is then input into the KAN for feature extraction to obtain the health features.
[0061] In one embodiment of the present application, extracting features from the medical image to be detected based on a visual encoder to obtain the visual features includes:
[0062] Preprocessing the medical image to be detected to obtain a preprocessed medical image;
[0063] performing feature extraction on the preprocessed medical image to obtain initial visual features;
[0064] Based on the initial visual features, multi-scale visual features are obtained.
[0065] Optionally, the medical image to be inspected is first preprocessed, such as converting it to a grayscale image and performing data augmentation operations such as random flipping and sharpening to obtain an enhanced grayscale image. The enhanced grayscale image can then be subjected to feature extraction using the PENet backbone network to obtain initial visual features. A feature pyramid operation is then performed on the extracted initial visual features. The multi-scale features extracted from the feature pyramid operation are then weighted by channel and spatial attention using a 3D multi-scale convolutional attention module. Finally, the features are fused using a bidirectional feedback propagation unit to obtain multi-scale visual features, which can be a one-dimensional vector of length 24.
[0066] In one embodiment of the present application, obtaining multi-scale visual features based on the initial visual features includes:
[0067] Performing a feature pyramid operation on the initial visual features to obtain visual features of different scales;
[0068] Perform channel attention and spatial attention weighting on the visual features of each scale to obtain multiple multi-scale features;
[0069] Multiple multi-scale features are fused to obtain multi-scale visual features.
[0070] like Figure 2 As shown in the figure, first, a feature pyramid operation is performed on the initial visual features, and multiple visual features of different scales are generated through a combination of bottom-up and top-down paths. The low-level features have high resolution and rich details, which are suitable for capturing small targets, while the high-level features have strong semantic information and are more conducive to detecting large targets. Then, for the visual features of each scale, the channel attention module (CAB) and the spatial attention module (SAB) are used to perform channel attention calculation and spatial attention calculation in sequence for weighting. Among them, the channel attention calculates the importance weight of each channel through global pooling and multi-layer perceptron to highlight the semantically significant channels, while the spatial attention focuses on the spatial area and emphasizes the key position information. The two work together to further enhance the expressive ability of the features and suppress background interference. Finally, the multiple multi-scale features after attention weighting are fused through the deep convolution fusion module, and the advantageous information at different scales is integrated to obtain multi-scale visual features containing rich details and semantics, providing better input for subsequent tasks such as target detection and image segmentation.
[0071] It should be noted that the channel attention module may include two parallel branches, each of which may include a pooling layer, a convolution layer, a modified activation function module, and a convolution layer connected in sequence. The input features enter the two branches respectively, and are processed by the maximum pooling through the pooling layer respectively, and then enter the respective convolution layers for convolution operations. The features after the convolution operation enter the modified activation function module, and through the activation functions such as the modified linear unit, nonlinear transformations are introduced to increase the nonlinear expression ability of the model and make the features more discriminative. Then, the convolution operation is performed again through the convolution layer to further refine and adjust the features, and the features output by the two branches are added element by element, and then activated through the S-type activation function (sigmoid function) to obtain the channel attention weight, which is applied to the features of the initial input to achieve channel attention weighting.
[0072] Among them, the spatial attention module may include a pooling layer, a convolution layer and an activation function in sequence. First, the input features are subjected to average pooling and maximum pooling respectively, and then spliced. The spliced features are input into the convolution layer for convolution operation, and then activated by an S-type activation function (sigmoid function) to obtain the spatial attention weight, which is applied to the initial input features to realize spatial attention weighting.
[0073] In addition, the deep convolution fusion module may include a sequentially connected convolution layer, a normalization layer, a rectified linear activation function layer, a multi-branch deep convolution module (including multiple parallel branches), a feature fusion layer, a channel shuffling layer, a convolution layer, a normalization layer, and an output layer. First, the features weighted by channel attention and spatial attention are input into the convolution layer for convolution operation. The features that have undergone convolution operation enter the normalization layer for normalization operation. The features after normalization operation are introduced into the rectified linear activation function, such as ReLU, to increase the model's expressiveness. The nonlinearly transformed features are then input into three parallel deep convolution branches. Each deep convolution branch independently performs convolution operation on each channel and extracts local features within the channel. Each branch then performs normalization operation and rectified linear activation function processing in turn to further activate the features. Afterwards, the features processed by the three branches are summed and fused to integrate the feature information extracted by different branches. The fused features enter the channel shuffling layer to disrupt the channel order, allowing the information between different channels to be fully mixed, enhancing the information flow and feature expression between channels. The channel-shuffled features are convolved again to further extract features, then normalized, and finally added and fused with the initial input features to obtain the final multi-scale visual features.
[0074] In an embodiment of the present application, preprocessing the medical image to be detected to obtain a preprocessed medical image includes:
[0075] Extracting a data table from a storage file corresponding to the medical image to be detected, and reading original image data based on the data table, wherein the original image data includes a pixel sequence of the data table;
[0076] Determining a target window position and a target window width corresponding to the medical image to be detected based on the pixel sequence;
[0077] generating a grayscale image corresponding to the medical image to be detected based on the target window level and the target window width;
[0078] An enhancement operation is performed on the grayscale image to obtain the enhanced grayscale image for use in visual feature extraction.
[0079] Alternatively, the medical image to be examined is typically stored in the DICOM format. DICOM (Digital Imaging and Communications in Medicine) is an international standard for medical images and related information. Therefore, a data structure in the form of a data table can be extracted from the DICOM file corresponding to the medical image to be examined. This data table is then read to obtain the raw image data. The key component of raw image data is the pixel sequence, which records information such as the grayscale value of each pixel in the image.
[0080] The pixel sequence read is the raw grayscale value matrix stored in the DICOM file. Each pixel corresponds to a specific numerical value. By filtering and mapping these grayscale values, the target window level and width are determined. Window level and width are important concepts in the display of medical images. The window width determines the range of grayscale values displayed, while the window level is the center value of this range. Different window level and width settings can highlight the details of different tissues in the image. Therefore, based on the determined target window level and width, the raw image data can be processed, and the grayscale values within the corresponding range can be mapped into a new grayscale image, resulting in a grayscale image that only highlights information in the lung region. To increase data diversity, the grayscale image can be enhanced. For example, the grayscale image can be flipped at random angles, including horizontal and vertical flips and arbitrary rotations. This simulates lung images captured from different angles, allowing the model to learn richer image features and avoid overfitting. Image sharpening algorithms can also be used to enhance image edges and details, making lung tissue boundaries clearer and textures more pronounced. After these data enhancement operations, an enhanced grayscale image is obtained, which contains more diverse image information and helps improve the results of subsequent analysis and model training.
[0081] In one embodiment of the present application, the multi-scale operation is performed on the visual features and health features to obtain multi-scale feature pairs, feature fusion is performed on the feature pairs at each scale, and bidirectional multi-scale fusion is performed on the feature pairs after feature fusion to obtain multi-scale fusion features, including:
[0082] Performing continuous inverted pyramid operations on the visual features and health features respectively to obtain multiple feature pairs of different scales;
[0083] Perform cross-attention calculation on feature pairs of each scale to obtain multiple fused feature pairs of different scales;
[0084] Multiple fusion feature pairs of different scales are bidirectionally fused to obtain multi-scale fusion features.
[0085] Optionally, after obtaining the visual features and health features, they can be input into the fusion module for multi-scale feature fusion processing. The fusion module includes an inverted feature pyramid operation module, a cross attention module and a bidirectional multi-scale fusion module (BSF) connected in sequence. First, the extracted visual features and health features are subjected to an inverted feature pyramid operation to divide them into multi-scale features, for example, into three scale features of 24, 12, and 6, and then the cross attention module is used to perform cross attention calculation on each scale feature pair obtained. One of the visual features and health features can be used as a query (Q), and the other as a key (K) and value (V). Subsequently, Q, K, and V are subjected to multi-head operations. Q and K can be used to calculate the attention score between the two modalities, these attention scores are normalized, and then multiplied by V to obtain the final output. This process can be expressed by the following formula:
[0086]
[0087] Among them, the shape of Q is [B,N,D q ], and the shape of key K and value V is Indicates the number of samples, N indicates the number of query vectors, and M indicates the number of key-value pairs. W represents the feature dimension of each key vector and value vector. q 、W k 、W υ are the weight matrices used to perform linear transformations on Q, K, and V, respectively. q is the dimension of Q′ after transformation, D k is the original dimension of K, D υ is the original dimension of V, Q′, K′, and V′ represent the query, key, and value vectors after linear transformation, respectively.
[0088] After reshaping Q′, K′, and V′, they can be expressed as:
[0089] Q h =reshape(Q',B,H,N,C),
[0090] K h =reshape(K′,B,H,M,C),
[0091] V h =reshape(V',B,H,M,C),
[0092] Among them, C represents the attention dimension of each attention head, and H represents the number of attention heads.
[0093] Then, the attention calculation can be expressed as:
[0094]
[0095] Among them, Attention represents the attention score matrix, Oh represents the final output feature, and T represents transpose.
[0096] It should be noted that the attention score is used to measure the degree of correlation between different modal features. The higher the score, the stronger the correlation between the two features.
[0097] Then, the bidirectional multi-scale fusion module can be used to perform bidirectional multi-scale fusion on the multiple fusion feature pairs of different scales obtained after the cross-attention calculation.
[0098] In one embodiment of the present application, the bidirectional multi-scale fusion of multiple fusion feature pairs at different scales to obtain multi-scale fusion features includes:
[0099] Performing linear transformation on the fusion feature pairs of different scales;
[0100] After performing element-by-element dot multiplication on the fused feature pairs after linear transformation, activation processing is performed through the preset activation function to obtain the key features;
[0101] After performing element-wise dot multiplication on the key feature and the fusion feature, they are concatenated in the channel dimension to obtain the multi-scale fusion feature.
[0102] like Figure 4 As shown, the bidirectional multi-scale fusion module (BSF) can include two linear layers. First, the linear layers can be used to linearly transform the visual features and health features of each scale, and each modality can correspond to one linear layer. Then, the features transformed by the linear layer will be element-wise multiplied. The sigmoid activation function maps the input value to a value between 0 and 1, and outputs a probability-like value, thereby highlighting the important fusion features and suppressing the unimportant parts. The result after the sigmoid function processing is then element-wise multiplied with the feature previously processed by the linear layer. Operation, and finally splicing the processed results in the channel dimension to obtain the final multi-scale fusion features.
[0103] In order to further verify the effectiveness of the multimodal and multi-scale classification module provided in this application, a comparative experiment was conducted. The experiment used a public dataset provided by The Cancer Imaging Archive, named Lung-PET-CT-Dx, which contains CT or PE scans and tabulated clinical information of 355 patients. The dataset provides tumor classification labels for each patient. Most PET / CT data are stored in DICOM format, and all data have been de-identified. The dataset contains 251 adenocarcinoma samples and 61 squamous cell carcinoma samples. Since the number of samples in the adenocarcinoma and squamous cell carcinoma categories in the original dataset is extremely unbalanced, we decided to oversample the squamous cell carcinoma category. We used a random oversampling strategy to supplement the squamous cell carcinoma samples in the training set from 34 to 198, and the number of validation and test sets remained unchanged at 12 and 15, respectively.
[0104] Each sample includes a comprehensive CT image and health data, including gender, age, weight, TMN stage, and smoking history. Data preprocessing involves converting nominal variables in the clinical information table into numerical variables, and filling missing data with the mean value of samples within the same category. Furthermore, several operations can be performed on the CT images, including random angle flipping, sharpening, and other data enhancements.
[0105] The SGD optimizer is used in the experiments. The training epochs, learning rate, batch size, and weight decay are set to 0.0001, 4, and 0.01, respectively.
[0106] Furthermore, the performance of this model is evaluated using the following comparative experiments: (1) PECon: a multimodal classification model proposed by Sanjeev et al.; (2) MedFuse: a multimodal classification model proposed by Hayat et al.; (3) Drfuse: a multimodal classification model proposed by Yao et al.; (4) MMTM: a multimodal classification model proposed by Joze et al.; (5) PEfusion: a multimodal classification model proposed by Huang et al.; (6) daft: et al.; (7) SAM+E3D-MSCA: The SAM model proposed by Kirillov, et al. is combined with the E3D-MSCA module proposed by us and only uses image data for classification; (8) PENet: The image classification model proposed by Huang et al.; (9) PENet+E3D-MSCA+drop: The single image classification model proposed by this method, which combines PENet and E3D-MSCA with dropout operation; (10) Cross-Attention: Cross-Attention is used to fuse image and table modality data; (11) clip_Fusion: Clip proposed by Radford et al. is used to align and fuse image and table modality data; (12) Late_Fusion: Late fusion is used to fuse image and table modality data. The results obtained by the above comparative experimental models are compared with the results obtained by using this method. The experimental results shown in Tables 1 to 3 below are as follows:
[0107] Table 1: Comparative experimental results
[0108] Method AUROC ACC F1 Specificity Sensitivity PPV NPV PECon 0.786 0.744 0.645 0.786 0.667 0.625 0.815 MedFuse 0.786 0.721 0.571 0.821 0.533 0.615 0.767 Drfuse 0.613 0.676 0.353 0.800 0.333 0.375 0.769 MMTM 0.802 0.698 0.581 0.750 0.600 0.562 0.778 PEfusion 0.740 0.721 0.600 0.786 0.600 0.600 0.786 daft 0.727 0.729 0.667 0.786 0.650 0.684 0.759 MMCAF-Net 0.786 0.791 0.690 0.857 0.667 0.714 0.828
[0109] Table 2: Image encoder ablation experiment results
[0110] Method AUROC ACC F1 Specificity Sensitivity PPV NPV SAM+E3D-MSCA 0.610 0.651 0.444 0.786 0.400 0.500 0.710 PENet 0.560 0.721 0.400 0.964 0.267 0.800 0.711 PENet+E3D-MSCA 0.644 0.721 0.500 0.893 0.400 0.667 0.735 PENet+E3D-MSCA+drop 0.712 0.767 0.545 0.964 0.400 0.857 0.857
[0111] Table 3 Fusion method ablation experiment results
[0112] Method AUROC ACC F1 Specificity Sensitivity PPV NPV Cross-Attention 0.752 0.744 0.686 0.714 0.800 0.600 0.870 clip_Fusion 0.717 0.698 0.629 0.679 0.733 0.550 0.826 Late_Fusion 0.729 0.674 0.462 0.821 0.400 0.545 0.719 MSCA_Fusion 0.786 0.791 0.690 0.857 0.667 0.714 0.828
[0113] As shown in Table 1 above, MMCAF-Net surpasses all comparison methods in multiple key indicators, including accuracy (ACC), F1 score (F1), specificity, sensitivity, positive predictive value (PPV) and negative predictive value (NPV). Although MMCAF-Net is 1.6% lower than MMTM in the area under the receiver operating characteristic curve (AUROC), it exceeds MMTM by about 10% in both ACC and F1 score, and shows a significant 15% improvement in PPV. This shows that MMCAF-Net achieves a lower false positive rate, demonstrating its high accuracy and reliability. Overall, it is worth noting that all models perform relatively poorly in the F1 indicator, which indicates that there may be noise or outliers in the dataset. In addition, the classification results are visualized for challenging cases. As shown in Figure 5 As shown, our model outperforms other models in distinguishing subtle and difficult cases.
[0114] As shown in Table 2, the proposed efficient 3D multi-scale PENet is compared with the SegmentAnything Model (SAM) that incorporates multi-scale technology and the original PENet. The results show that the proposed method outperforms the other two methods on all evaluation metrics, achieving improvements of 10% and 15% in AUROC and 10% and 14% in F1 score, respectively. In addition, compared to the standalone PENet, the proposed method shows significant improvements in AUROC and NPV, indicating that the improved architecture enhances the reliability of negative class recognition.
[0115] In addition, we conducted ablation experiments on different fusion methods. As shown in Table 3, we compared the proposed MSCA with Cross-Attention, Late Fusion, and Clip Aligned Fusion. The results show that MSCA outperforms the other three fusion methods in most metrics, with accuracy improvements of 5%, 10%, and 12%, respectively. This demonstrates that our method outperforms the other methods in overall performance.
[0116] In the embodiment of the present application, through feature extraction and multi-scale operation, the lesion characteristics are captured from multiple dimensions, combined with bidirectional multi-scale fusion, the perception ability of small lesions (such as squamous cell carcinoma, pulmonary embolism) is enhanced, and the lesion detection rate is significantly improved. The visual features and health features are refined, and the information barriers between modalities are broken through the construction and deep fusion of multi-scale features, and the efficient collaboration of images and health data is achieved, and the potential correlation between data is deeply mined. Through the bidirectional multi-scale fusion mechanism, features of different scales are fully integrated, which not only retains the image detail information, but also integrates the semantic knowledge of health data, making the multi-scale fusion features more representative and providing richer and more accurate information for disease diagnosis. Effectively improve the accuracy and reliability of disease diagnosis, and provide strong support for clinical auxiliary diagnosis.
[0117] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0118] In one embodiment, a classification device based on multimodal multiscale fusion is provided, and the classification device based on multimodal multiscale fusion corresponds one-to-one with the classification method based on multimodal multiscale fusion in the above embodiment. Figure 6 As shown, the classification device based on multimodal multiscale fusion includes a data acquisition unit 10, a feature extraction unit 20, a feature fusion unit 30 and a classification unit 40. The functional modules are described in detail as follows:
[0119] The data acquisition unit 10 is used to acquire the medical image and health data to be detected;
[0120] A feature extraction unit 20 is used to extract features from the medical image to be detected and the health data to be detected, respectively, to obtain visual features and health features;
[0121] A feature fusion unit 30 is configured to perform multi-scale operations on the visual features and health features to obtain multi-scale feature pairs, perform feature fusion on the feature pairs at each scale, and perform bidirectional multi-scale fusion on the feature pairs after feature fusion to obtain multi-scale fusion features;
[0122] The classification unit 40 is configured to obtain a classification result based on the multi-scale fusion features.
[0123] In one embodiment of the present application, the feature fusion unit 30 is further configured to:
[0124] Performing continuous inverted pyramid operations on the visual features and health features respectively to obtain multiple feature pairs of different scales;
[0125] Perform cross-attention calculation on feature pairs of each scale to obtain multiple fused feature pairs of different scales;
[0126] Multiple fusion feature pairs of different scales are bidirectionally fused to obtain multi-scale fusion features.
[0127] In one embodiment of the present application, the feature fusion unit 30 is further configured to:
[0128] Performing linear transformation on the fusion feature pairs of different scales;
[0129] After performing element-by-element dot multiplication on the fused feature pairs after linear transformation, activation processing is performed through the preset activation function to obtain the key features;
[0130] After performing element-wise dot multiplication on the key feature and the fusion feature, they are concatenated in the channel dimension to obtain the multi-scale fusion feature.
[0131] In one embodiment of the present application, the feature extraction unit 20 is further configured to:
[0132] Extracting features from the medical image to be detected based on a visual encoder to obtain the visual features; and
[0133] The health data to be detected is converted into table data, and feature extraction is performed on the table data based on a table encoder to obtain the health feature.
[0134] In one embodiment of the present application, the feature extraction unit 20 is further configured to: pre-process the medical image to be detected to obtain a pre-processed medical image;
[0135] performing feature extraction on the preprocessed medical image to obtain initial visual features;
[0136] Based on the initial visual features, multi-scale visual features are obtained.
[0137] In one embodiment of the present application, the feature extraction unit 20 is further configured to:
[0138] Performing a feature pyramid operation on the initial visual features to obtain visual features of different scales;
[0139] Perform channel attention and spatial attention weighting on the visual features of each scale to obtain multiple multi-scale features;
[0140] Multiple multi-scale features are fused to obtain multi-scale visual features.
[0141] In one embodiment of the present application, the feature extraction unit 20 is further configured to:
[0142] Extracting a data table from a storage file corresponding to the medical image to be detected, and reading original image data based on the data table, wherein the original image data includes a pixel sequence of the data table;
[0143] Determining a target window position and a target window width corresponding to the medical image to be detected based on the pixel sequence;
[0144] generating a grayscale image corresponding to the medical image to be detected based on the target window level and the target window width;
[0145] An enhancement operation is performed on the grayscale image to obtain the enhanced grayscale image for use in visual feature extraction.
[0146] In the embodiment of the present application, through feature extraction and multi-scale operation, the lesion characteristics are captured from multiple dimensions, combined with bidirectional multi-scale fusion, the perception ability of small lesions (such as squamous cell carcinoma, pulmonary embolism) is enhanced, and the lesion detection rate is significantly improved. The visual features and health features are refined, and the information barriers between modalities are broken through the construction and deep fusion of multi-scale features, and the efficient collaboration of images and health data is achieved, and the potential correlation between data is deeply mined. Through the bidirectional multi-scale fusion mechanism, features of different scales are fully integrated, which not only retains the image detail information, but also integrates the semantic knowledge of health data, making the multi-scale fusion features more representative and providing richer and more accurate information for disease diagnosis. Effectively improve the accuracy and reliability of disease diagnosis, and provide strong support for clinical auxiliary diagnosis.
[0147] For the specific definition of the classification device based on multimodal multi-scale fusion, please refer to the definition of the classification method based on multimodal multi-scale fusion above, and will not be repeated here. Each module in the above-mentioned classification device based on multimodal multi-scale fusion can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0148] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer-readable instructions are executed by the processor, a classification method based on multimodal and multi-scale fusion is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0149] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the above-mentioned classification method based on multimodal multi-scale fusion are implemented.
[0150] In an embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the classification method based on multimodal multi-scale fusion as described above are implemented.
[0151] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include processes in the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0152] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0153] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A classification method based on multimodal and multiscale fusion, characterized in that: The method comprises: Obtaining medical images to be tested and health data to be tested; Extracting features of the medical image to be detected and the health data to be detected to obtain visual features and health features; Performing multi-scale operations on the visual features and health features to obtain multi-scale feature pairs, performing feature fusion on the feature pairs at each scale, and performing bidirectional multi-scale fusion on the feature pairs after feature fusion to obtain multi-scale fusion features; Based on the multi-scale fusion features, a classification result is obtained.
2. The classification method based on multimodal and multiscale fusion according to claim 1, characterized in that: The multi-scale operation is performed on the visual features and the health features to obtain multi-scale feature pairs, feature fusion is performed on the feature pairs of each scale, and bidirectional multi-scale fusion is performed on the feature pairs of each scale after feature fusion to obtain multi-scale fusion features, including: Performing continuous inverted pyramid operations on the visual features and health features respectively to obtain multiple feature pairs of different scales; Perform cross-attention calculation on feature pairs of each scale to obtain multiple fused feature pairs of different scales; Multiple fusion feature pairs of different scales are bidirectionally fused to obtain multi-scale fusion features.
3. The classification method based on multimodal and multiscale fusion according to claim 2, characterized in that: The bidirectional multi-scale fusion of multiple fusion feature pairs of different scales to obtain multi-scale fusion features includes: Performing linear transformation on the fusion feature pairs of different scales; After performing element-by-element dot multiplication on the fused feature pairs after linear transformation, activation processing is performed through the preset activation function to obtain the key features; After performing element-wise dot multiplication on the key feature and the fusion feature, they are concatenated in the channel dimension to obtain the multi-scale fusion feature.
4. The classification method based on multimodal and multiscale fusion according to claim 1, wherein: The extracting features of the medical image to be detected and the health data to be detected respectively to obtain visual features and health features includes: Extracting features from the medical image to be detected based on a visual encoder to obtain the visual features; and The health data to be detected is converted into table data, and feature extraction is performed on the table data based on a table encoder to obtain the health feature.
5. The classification method based on multimodal and multiscale fusion according to claim 4, characterized in that: The extracting features of the medical image to be detected based on a visual encoder to obtain the visual features includes: Preprocessing the medical image to be detected to obtain a preprocessed medical image; performing feature extraction on the preprocessed medical image to obtain initial visual features; Based on the initial visual features, multi-scale visual features are obtained.
6. The classification method based on multimodal and multiscale fusion according to claim 5, characterized in that: The obtaining of multi-scale visual features based on the initial visual features includes: Performing a feature pyramid operation on the initial visual features to obtain visual features of different scales; Perform channel attention and spatial attention weighting on the visual features of each scale to obtain multiple multi-scale features; Multiple multi-scale features are fused to obtain multi-scale visual features.
7. The classification method based on multimodal and multiscale fusion according to claim 5, characterized in that: The preprocessing of the medical image to be detected to obtain a preprocessed medical image includes: Extracting a data table from a storage file corresponding to the medical image to be detected, and reading original image data based on the data table, wherein the original image data includes a pixel sequence of the data table; Determining a target window position and a target window width corresponding to the medical image to be detected based on the pixel sequence; generating a grayscale image corresponding to the medical image to be detected based on the target window level and the target window width; An enhancement operation is performed on the grayscale image to obtain the enhanced grayscale image for use in visual feature extraction.
8. A classification model based on multimodal and multiscale fusion, characterized in that: The classification model is used to implement the classification method based on multimodal and multiscale fusion according to any one of claims 1 to 7, and the classification model includes: Visual encoder, table encoder, fusion module and classification module; The visual encoder includes a feature extraction module, a feature pyramid operation module, a three-dimensional multi-scale convolutional attention module and a bidirectional feedback propagation unit. The three-dimensional multi-scale convolutional attention module includes a channel attention module, a spatial attention module and a deep convolutional fusion module connected in sequence, which is used to extract features from the medical image to be detected to obtain visual features; The table encoder is a Kolmogorov-Arnold network, which is used to extract features from the health data to be detected to obtain health features; The fusion module includes an inverted feature pyramid operation module, a cross attention module and a bidirectional multi-scale fusion module connected in sequence, and is used to fuse the visual features and health features to obtain multi-scale fusion features; The classification module is used to perform a classification operation based on the multi-scale fusion features to output a classification result.
9. A classification device based on multimodal and multiscale fusion, characterized in that: The device comprises: A data acquisition unit for detecting, used to acquire medical images and health data for detecting; a feature extraction unit, configured to extract features from the medical image to be detected and the health data to be detected, respectively, to obtain visual features and health features; a feature fusion unit, configured to perform multi-scale operations on the visual features and health features to obtain multi-scale feature pairs, perform feature fusion on the feature pairs at each scale, and perform bidirectional multi-scale fusion on the feature pairs after feature fusion to obtain multi-scale fusion features; The classification unit is used to obtain a classification result based on the multi-scale fusion feature.
10. A readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the classification method based on multimodal multi-scale fusion as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Remote sensing image change detection method based on language guidance
CN119169449A
Multi-modal multi-scale transformation fusion method and system based on exchange
CN119538188A
Heterogeneous cross-modal fusion-based pulmonary embolism auxiliary diagnosis method and system
CN119601215A
Multi-modal data fusion method and device, equipment and storage medium
CN120030496A