A high-order tensor multi-modal fusion method and model based on dual-mode spectral interactive learning

By employing a high-order tensor multimodal fusion method based on dual-mode spectral interactive learning, the problem of insufficient information interaction in multimodal medical data fusion is solved, enabling more accurate multimodal diagnosis and improving the robustness and diagnostic capability of the model.

CN117315427BActive Publication Date: 2025-12-05XINJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311380426.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-12-05
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

Existing research has neglected the complementarity between multimodal data in multimodal medical data fusion, and cannot effectively explain the dynamic modeling within and between modes, resulting in insufficient information interaction and affecting diagnostic accuracy.

Method used

A high-order tensor multimodal fusion method based on dual-mode spectral interactive learning is adopted. By combining high-order tensor outer product and cross-modal interactive learning with orthogonality loss, reconstruction loss and adversarial loss, non-cascaded feature-level fusion of Raman spectroscopy and infrared spectroscopy data is achieved to extract effective information.

Benefits of technology

It improves the robustness and accuracy of the model, better explains the complementarity of multimodal data, and enhances its diagnostic capabilities in the biomedical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315427B_ABST
    Figure CN117315427B_ABST
Patent Text Reader

Abstract

The application specifically relates to a high-order tensor multi-modal fusion method and model based on double-mode spectrum interactive learning.A high-order tensor multi-modal fusion method based on double-mode spectrum interactive learning comprises the following steps: S10, inputting Raman spectrum and infrared spectrum data; S20, performing non-cascaded multi-mode spectrum fusion representation on the Raman spectrum and the infrared spectrum data through high-order tensor outer product; and S30, calculating orthogonality loss, reconstruction loss and adversarial loss, and compensating for heterogeneity difference by performing unique representation learning on the Raman spectrum feature and the infrared spectrum feature.The high-order tensor multi-modal fusion method and model based on double-mode spectrum interactive learning effectively realize multi-modal data information fusion by obtaining high-order interactive fusion features of double-mode spectrum information through BHTF, and realize more accurate cross-modal representation by learning the heterogeneity between multi-modal data through CMIL cross-modal learning, thereby improving the robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of medical data processing, specifically relating to a high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning. Background Technology

[0002] In recent years, molecular vibrational spectroscopy of biofluids has become a valuable source of medical data for disease diagnosis. Its advantages include non-invasiveness and high resolution, leading to numerous studies combining artificial intelligence with spectroscopy for non-invasive disease diagnosis, such as lung cancer, kidney cancer, esophageal cancer, cervical cancer, breast cancer, glioma, and other malignant tumors. It has also been shown to be useful in diagnosing infectious diseases and systemic lupus erythematosus. Infrared and Raman spectroscopy, both molecular vibrational spectra, can provide rich biomedical information in a non-destructive manner. However, the aforementioned studies primarily rely on single-type data, which cannot provide sufficiently accurate information to support accurate diagnosis. Many research methods have demonstrated that multimode studies outperform single-mode studies. The differences between infrared and Raman spectroscopy in vibrational modes, wavelength ranges, and detection levels enable the fusion of these two spectra to be complementary and improve the acquisition of information from samples. This complementarity allows these two technologies to be applied in the biomedical field.

[0003] The continuous advancement and development of smart healthcare has led to the generation of massive amounts of medical data from sensors, medical-related equipment, and communication technologies. Multimodal medical data is widely used in disease diagnosis, treatment, and clinical research. In predicting complex diseases, the effective fusion of medical data to handle this vast amount of multi-source data has become an increasingly important research area.

[0004] In view of this, the present invention proposes a high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning, which extracts effective information from different modal data and performs non-cascaded feature-level fusion to achieve effective fusion. Summary of the Invention

[0005] In the field of intelligent medical diagnosis, medical data is becoming increasingly multimodal, and vibrational spectroscopy-assisted detection technology is rapidly developing due to its advantages such as non-invasiveness and high speed. However, multimodal data fusion can extract richer information than single-modal fusion, and the effective fusion of multispectral data faces challenges. Existing research often adopts simple cascade processing methods, neglecting the complementarity between multimodal data fusions and failing to adequately explain the dynamic modeling within and between modes. Therefore, exploring the complementarity of information between multispectral data and the optimal fusion strategy is crucial for multimodal tasks.

[0006] Based on this, the present invention provides a high-order tensor multimodal fusion method based on dual-mode spectral interaction learning, which constructs important high-order interactions of non-cascaded fusion features of multimodal data in a high-dimensional space; it adopts a cross-modal interaction learning approach to learn between heterogeneous modalities, achieves more accurate cross-modal representation, improves the robustness of the model, and thus achieves effectiveness and universality in the biomedical multimodal field.

[0007] To achieve the above objectives, the technical solution adopted is as follows:

[0008] A high-order tensor multimodal fusion method based on dual-mode spectral interactive learning includes the following steps:

[0009] S10: Input Raman and infrared spectral data;

[0010] S20: By using high-order tensor outer product, the non-cascaded multimode spectral characterization of the Raman spectrum and infrared spectral data is fused to achieve the learning of the interaction between the two modes;

[0011] S30: Calculate the orthogonality loss, reconstruction loss, and adversarial loss, and compensate for the heterogeneity differences between different spectra by performing unique characterization learning on the Raman spectral features and infrared spectral features.

[0012] Furthermore, the Raman and infrared spectral data are subjected to baseline correction, smoothing, and normalization before input.

[0013] Furthermore, the Raman and infrared spectral data mentioned are Raman and infrared spectral data of serum samples.

[0014] Furthermore, the high-order tensor multimodal fusion method also includes step S15, feature importance alignment, which calculates the feature importance of the Raman and infrared spectral data and performs feature selection.

[0015] Furthermore, the specific processing of step S15 is as follows: first, the feature importance ranking matrix is ​​obtained by scoring and ranking the feature importance of the original data while maintaining the original data dimensions; then, the data dimensions are pruned according to the contribution of each feature to the target variable; finally, the dimension with the best performance in different datasets is used as the baseline.

[0016] Furthermore, the specific processing of step S20 is as follows: first, a unified pre-aligned encoder is used to perform nonlinear mapping to obtain the pre-fusion representation of the dual-mode spectrum; then, a high-order tensor fusion method is adopted to upgrade the different spectral representations to three-dimensional space in the form of rows and columns according to the high-order matrix multiplication rules, and batch matrix multiplication is used for efficient fusion; finally, the non-cascaded fusion features of high-order interaction are mapped back to the low-dimensional space.

[0017] Furthermore, the specific processing steps of step S20 are as follows:

[0018] ① Through encoder E with a multi-layer fully connected architecture c Learning nonlinear characterizations of different spectral modes z m The formula is as follows:

[0019] z m =E c (h m ;w),m∈{r,i}

[0020] In the formula, E c This represents the encoder function for dual-mode spectral prefusion, where w represents E. c The parameters {r,i} represent the Raman and infrared spectral data, respectively.

[0021] ② The bimodal eigenvectors are increased in dimensionality in the second and third dimensions to adapt to batch matrix multiplication in a higher-dimensional space, forming a 3-D fusion tensor of higher-order bimodal spectral interactions. This allows for full learning of intermodal dynamics modeling, as shown in the following formula:

[0022] z m' =[z m ,1],m∈{r,i},

[0023] In the formula, z represents the outer product between vectors. f ∈R (r+1)*(i+1) It is a 3-D tensor cube of all possible combinations of two-mode spectra.

[0024] ③ Combine the non-cascaded tensor z f The data is fed into the decoder and classifier for task classification, and the formula is as follows:

[0025] z′ f =D c (z f ;w),

[0026] In the formula, w represents D c The parameter, z′ f This represents the fused feature representation after passing through the fully connected layer decoder.

[0027] Furthermore, in step S30, a shared representation encoder for dual-mode representation is constructed. Obtain features and specific representation encoder Obtain features Then through the encoder equation Where n represents the shared encoder and the specific encoder, respectively, the orthogonality loss, reconstruction loss, and adversarial loss are calculated;

[0028] The orthogonality loss L diff The calculation formula is:

[0029] The reconstruction loss L rec The calculation formula is:

[0030] The aforementioned adversarial loss L adv The calculation formula is: L adv =L f +L t

[0031]

[0032]

[0033] Among them, L f For pseudo-adversarial loss, L t For true adversarial losses, To obtain the target through the discriminator The predicted value.

[0034] This invention also provides a high-order tensor multimodal fusion model based on dual-mode spectral interactive learning, which is a lightweight and efficient general-purpose model for processing unstructured and structured multimodal data. It has a powerful ability to fuse multimodal data and high accuracy.

[0035] To achieve the above objectives, the technical solution adopted is as follows:

[0036] A high-order tensor multimodal fusion model based on dual-mode spectral interactive learning is obtained using the aforementioned high-order tensor multimodal fusion method.

[0037] Furthermore, it includes: an input module, a feature importance alignment module, a dual-modal high-order tensor fusion module, and a cross-modal interactive learning module.

[0038] The input module is used to input Raman and infrared spectral data;

[0039] The feature importance alignment module is used to calculate the feature importance of the Raman and infrared spectral data and perform feature selection.

[0040] The dual-mode high-order tensor fusion module uses high-order tensor outer product to enable non-cascaded multimode spectral fusion characterization of Raman and infrared spectral data, thereby achieving the learning of the interaction between the two modes;

[0041] The cross-modal interactive learning module consists of orthogonality loss, reconstruction loss, and adversarial loss. It compensates for the heterogeneity differences between different spectra by performing unique characterization learning on the Raman spectral features and infrared spectral features.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] This invention presents a high-order tensor multimodal fusion method and model based on dual-mode spectral interaction learning. Based on Raman and infrared spectroscopy, it extracts effective information from different modal data and performs non-cascaded feature-level fusion to achieve effective fusion, resulting in a dual-mode interaction learning high-order tensor multimodal fusion (BHTMF) model. It also considers the interactions within, between, and between bimodal modes of multimodal data. Inspired by multimodal tasks in the field of sentiment analysis in natural language processing, dual-mode fusion features obtain important high-order interactions between modes in a high-dimensional space through tensor fusion. Multimodal heterogeneity is learned through different loss modules to learn the relationships between different modes, reducing intermodal differences. Extensive experiments were conducted on two multispectral datasets to verify the effectiveness and universality of the BHTMF model in the biomedical multimodal field. The main advantages are as follows:

[0044] 1. The BHTMF model proposed in this invention is developed in a modular manner in the form of main and auxiliary tasks, which fully considers the interactions within the mode, between modes and between bimodals, as well as the cross-modal expressive ability of the model.

[0045] 2. In the technical solution of the present invention, a dual-mode high-order tensor fusion module (BHTF) is designed. BHTF uses 3-D tensors to efficiently construct important high-order interactions of non-cascaded fusion features in a high-dimensional space. It can process unstructured and structured multimodal data in a lightweight and efficient manner, and more intuitively explain the input and output of multimodal data in the fusion process.

[0046] 3. In the technical solution of the present invention, a cross-modal interactive learning module (CMIL) is designed. CMIL improves the performance of the main task by interactively learning more accurate cross-modal representations between heterogeneous modalities through the joint loss task of orthogonality loss, reconstruction loss and adversarial loss. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of a high-order tensor multimodal fusion (BHTMF) model.

[0048] Figure 2 Characteristics of early and late fusion;

[0049] Figure 3 An intermediate fusion method for fusing higher-order tensors. Detailed Implementation

[0050] To further illustrate the high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning proposed in this invention, and to achieve the intended purpose of the invention, the following detailed description, in conjunction with preferred embodiments, details the specific implementation, structure, features, and effects of the high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning proposed in this invention. In the following description, different "an embodiment" or "an embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0051] Before elaborating on the high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning of the present invention, it is necessary to further explain the relevant background mentioned in the present invention in order to achieve better results.

[0052] In recent years, many scholars have dedicated themselves to multimodal medical tasks. In the multimodal fusion strategy, strategies can be categorized into early, mid, and late-stage fusion based on the input method of the fusion layer. Compared to traditional machine learning methods, deep learning methods have significant advantages in multimodal learning. Shallow layers in deep models learn simple abstractions of data, while deeper layers combine these abstractions into more abstract representations, providing a superior approach to data fusion compared to shallow learning. Modal-related features only become apparent at higher levels of abstraction; for example, fully connected neural networks (FCNNs) improve the predictions of the final classifier by discovering simple dependencies between potential unentangled factors. More importantly, multimodal deep learning can model nonlinear intramodal and cross-modal relationships, leading to increased research and applications of deep learning models across various fields.

[0053] In multimodal fusion strategies in the medical field, early fusion, through simple feature concatenation, cannot fully learn the intermodal edge representations, while late fusion, by combining the decisions of different classifiers on individual unimodal sub-models to obtain the final decision, cannot fully learn the joint representations between modalities. For example, Yin et al., 2016, using the multimodal dataset DEAP, concatenated multiple brain electrophysiological signals but did not achieve intermodal information interaction, proposing an ensemble classifier based on stacked autoencoders with multiple fusion layers to identify emotions. Muzammel et al., 2021, through early fusion, simply concatenated audio unimodal representations with visual or text features to form spliced ​​fusion features, proposing a multimodal fusion system for depression detection. Zhu et al., 2022, by combining different classifiers to fuse EEG and eye-tracking data at the decision layer, did not fully consider the joint representations between modalities, proposing a content-based multi-evidence fusion for depression detection. Bi et al., 2022 constructed fusion features from brain regions and genes using graph learning, addressing the complex interactions that are difficult to interpret, and proposed a feature aggregation graph convolutional network for Alzheimer's disease diagnosis. Chen et al., 2023 obtained fusion features by cascading Fourier transform infrared spectroscopy, Raman spectroscopy, and their first derivative spectra, thus expanding spectral information at the data level, which also falls under the category of early fusion. Based on AlexNet, they developed a rapid diagnostic method for renal cell carcinoma, achieving an accuracy of 93%. Leng et al., 2023 obtained early fusion features by concatenating raw Raman and infrared spectral features and their PLS-selected features. Based on GS-SVM, MFCNN, and CNN-LSTM models, they diagnosed lung cancer, esophageal cancer, and glioma, achieving an accuracy of 79.91%. Importantly, biomedical applications are interested in the ability to perform mid-term fusion, reducing the heterogeneity gap between modalities by applying different networks and network types to each modality, thereby achieving effective fusion of imaging, molecular, and clinical modalities. Therefore, this invention uses a mid-term fusion approach to construct multimodal joint features. Before using each modality to learn the joint representation or make direct predictions, it learns the edge representation of each modality to discover the correlation within the modality.

[0054] However, research on multimodal biomedical data fusion is still in its early stages, although some studies have demonstrated its significance for medical research. However, fusion algorithms and strategies require further improvement. In the field of NLP, for example, Zadehe et al. (2017) proposed a tensor fusion network model capable of end-to-end learning of intramodal and intermodal dynamics. Liu et al. (2018) proposed a low-rank multimodal fusion method using low-rank tensors, significantly reducing computational complexity. The core method is tensor fusion; tensors are powerful because they capture important high-order interactions across time, feature dimensions, and multiple modalities. Hazarika et al. (2020) proposed modality-invariant and modality-specific representations at the modality representation level, introducing a novel multimodal sentiment analysis model. Wu et al. (2022) proposed a novel bimodal multi-head attention fusion network for sentiment prediction by considering intramodal, intermodal, and bimodal interactions and fusing different groups of textual, visual, and acoustic signals. In summary, numerous multimodal fusion ideas have emerged and can be widely applied across various fields. While relatively sophisticated fusion algorithms have been developed in the field of natural language processing, there is a continued need to develop new fusion algorithms and strategies suitable for medical data fusion. This is crucial for maximizing the utilization of multimodal data and enhancing its feasibility in clinical applications.

[0055] Having understood the relevant materials mentioned in this invention, the following will provide a more detailed description of a high-order tensor multimodal fusion method and model based on dual-mode spectral interactive learning, in conjunction with specific embodiments:

[0056] Example 1.

[0057] To address the issue of insufficient information exchange between different data fusion methods in multimodal spectroscopy, this invention proposes a dual-modal interactive learning high-order tensor multimodal fusion model. In this model, Raman spectroscopy and infrared spectroscopy are used as inputs. Furthermore, the model consists of a primary task and an auxiliary task, representing BHTF and CMIL, respectively. Figure 1 As shown, the input to the BHTMF model is preprocessed and feature importance aligned Raman and infrared spectra; the main task BHTF is used to extract single-peak spectral features and model inter-modal fusion features, which can finally be used for disease prediction; the auxiliary task CMIL is used to extract shared and specific representations of single-peak spectra, and performs cross-modal interactive learning through a multi-loss module to assist the main task in learning performance.

[0058] Specifically, this model obtains feature representations from multimodal data through nonlinear projection using different encoding modules. In the main task, a high-order tensor outer product is used for the first time to efficiently learn the interactions between bimodalities in batches. This results in a non-cascaded multimodal spectral fusion representation of Raman and infrared spectra. In the auxiliary task, a cross-modal interaction learning module, mainly composed of orthogonality loss, reconstruction loss, and adversarial loss, compensates for the heterogeneity differences between different spectra by uniquely learning representations of Raman and infrared spectral features. Furthermore, it intuitively explains how tensor fusion can fully integrate multimodal information in high-dimensional space, significantly improving computational efficiency and using less memory. In summary, the main task obtains important intermodal interaction features for disease prediction, while the auxiliary task assists in disease prediction and recognition rates by interactively learning the relationships between bimodal spectra, thus improving the model's cross-modal expressive ability. The algorithm is described in Table 1 below:

[0059] Table 1: Algorithm Description

[0060]

[0061]

[0062] 1. Dual-mode high-order tensor fusion module (BHTF module)

[0063] like Figure 2-3 As shown, traditional early and late-stage fusion methods simply fuse features from different modalities through cascading and late-stage weighted decision-making. This approach ignores complementary, redundant, or collaborative information between multimodal features. The BHTF module offers a non-cascading multimodal feature fusion method. Specifically, to address the differences in input dimensions between different modalities, a unified pre-aligned encoder is used for nonlinear mapping to obtain the pre-fusion representation of the dual-mode spectra. Then, BHTF employs a high-order tensor fusion approach, upscaling the different spectral representations to a three-dimensional space according to high-order matrix multiplication rules, using batch matrix multiplication for efficient fusion. The 3-D tensor outer product formed in the high-dimensional space fully considers the differences between samples from a holistic perspective, thus accurately capturing patterns and features in the data. Finally, the non-cascading fusion features with high-order interactions are mapped back to a low-dimensional space for subsequent learning.

[0064] First, technically speaking, the two inputs to the dual-mode fusion module are: Raman spectrum h r ∈R r×1 Infrared spectrum h i ∈R i ×1 , (r represents h) r (and so on) the BHTF module first passes through the encoder E with a multi-layer fully connected (FC) architecture. cLearning nonlinear characterizations of different spectral modes z m It can be represented as follows:

[0065] z m =E c (h m ;w),m∈{r,i} (1)

[0066] E c This represents the encoder function for dual-mode spectral prefusion, where w represents E. c The parameters {r,i} represent two types of spectral data, Raman spectrum and infrared spectrum. By using formula (1), the information of different modes can be nonlinearly mapped to the feature subspace. Through the self-optimizing parameter w, the pre-fusion encoder can learn the feature representation of each mode from within the mode before being sent to the fusion module, thus reducing the mode gap.

[0067] Secondly, in the fusion module, each local part is filled using (1) to preserve the interaction of any subset within the mode. The bimodal eigenvectors filled with (1) are then increased in dimensionality in the second and third dimensions to adapt to batch matrix multiplication in high-dimensional space, forming a 3-D fusion tensor of high-order bimodal spectral interactions, thus fully learning the intermodal dynamics modeling. As shown in Equation 2-3:

[0068] z m' =[z m ,1],m∈{r,i} (2)

[0069]

[0070] here Let z denote the outer product between vectors, and z f ∈R (r+1)*(i+1) It is a 3D tensor cube representing all possible combinations of two-mode spectra. The outer product of two vectors is a special case of the tensor product, essentially a bilinear mapping, which is simply matrix multiplication. A bilinear mapping generates an element in a third vector space from elements in two vector spaces. The fused tensor obtained through high-dimensional interaction operations can fully represent the interaction between two-mode spectra. Figure 2 It demonstrates its formation process.

[0071] Finally, the non-cascaded fusion tensor z f The data is fed into the decoder and classifier for disease prediction, where w represents D. c The parameter, z′ f This represents the fused feature representation after passing through the fully connected layer decoder, which is then connected to the softmax classifier to obtain... Disease prediction. As shown in formulas 4 and 5:

[0072] z′ f =D c (z f ;w) (4)

[0073]

[0074] 2. Cross-modal interactive learning module

[0075] To ensure that the bimodal feature representations are fully learned, a CMIL (Common Mode Injection Logic) module was designed for cross-modal interaction. This module is a joint loss network model consisting of orthogonality loss, reconstruction loss, and adversarial loss. Specifically, the network includes two unique encoders. First, a shared representation encoder for the bimodal modes is constructed. get and specific representation encoder get By using encoder equations Here, n represents the shared encoder and the specific encoder, respectively. Subsequently, these two unique representations are used for cross-modal interactive learning to compensate for the heterogeneity differences between the two-mode spectra, thereby improving the generalization ability and prediction performance of the overall model.

[0076] (1)L diff -Orthogonality Loss.

[0077] This loss is designed to ensure that shared representations and specific representations capture different aspects of the input. This is achieved by imposing a soft orthogonality constraint between the two representations, which is calculated as follows:

[0078]

[0079] In the formula, It is the square of the Frobenius norm. By ensuring good separation of representations in the shared and specific subspaces, intermodal differences between two-mode spectra are learned. Specifically, the following loss function is used to enhance the orthogonality between the representations in the shared and specific subspaces for each mode, which can be computed as:

[0080]

[0081] (2)L rec -Reconstruction Loss.

[0082] When orthogonality loss is applied, special cases of intermodal representation learning still exist, such as the encoder function potentially obtaining orthogonal vectors that approximate the modalities, leading to insufficient learning. To avoid this, a reconstruction loss is designed. This loss ensures that the representations of the main task BHTF module and the CMIL module can work synergistically and complementaryly to provide a comprehensive multimodal representation of the patient, complementing the modal information of the main and auxiliary tasks during representation learning. Specifically, this is achieved by using a reconstruction decoder equation D composed of several FC layers. rec Learning the main task fusion feature z′ f and modality-specific representation The joint statement, in which The input represents f. m The reconstruction of can be calculated as follows:

[0083]

[0084] The core task of this loss is to recover the original representation of cross-modal specific representations and fused features. The reconstruction loss is f... m and Mean squared error loss between them It is the square of the L2 specification, which can be calculated as:

[0085]

[0086] (3)L adv -Adversarial Loss.

[0087] Because the different statistical properties between heterogeneous modalities make it difficult to provide a comprehensive exploration for multimodal fusion, an adversarial training method was designed for bimodal modality representation to effectively narrow the modal gap. This method transforms the general modal distribution in the bimodal spectrum into a distribution representing a better modality. Specifically, the adversarial network establishes an adversarial game between a shared encoder and two discriminators.

[0088] First, adversarial training is used to aid the cross-modal representation distribution of bimodal spectral modes in a common subspace. Given the widespread use of bimodal spectroscopy in disease prediction, the spectra with superior performance are designated as true target labels, while the remaining modes are used as false labels. Then, two discriminators are defined to guide the shared encoder as adversaries. Modality distribution transformation of the generated shared representations; simultaneously The generator attempts to deceive the two discriminators into making the distribution of the false target's representation more similar to that of the real target. Finally, the distribution of the two-mode spectral representations is effectively similarized in the common subspace. Therefore, the adversarial loss function L... adv It consists of two parts: pseudo-adversarial loss L fTrue adversarial loss L t As shown below:

[0089] L adv =L f +L t (10)

[0090] L f and L t The definition is as follows, where The representative obtains the pair through the discriminator. The predicted value. By minimizing L adv To compensate for the heterogeneity differences in dual-mode spectral data:

[0091]

[0092]

[0093] 3. Feature importance alignment

[0094] Due to the high dimensionality of spectral data, inconsistencies in dimensionality may exist between different modalities. Processing high-dimensional data can lead to insufficient information fusion, high computational costs, and inefficiency. In typical multimodal fusion, using all original modal information may ultimately negatively impact the performance of deep learning algorithms. Therefore, feature selection becomes crucial in improving model learning capabilities while reducing the original dimensionality. This paper proposes a feature selection method based on Random Forest Feature Importance Alignment (RFIA). By calculating feature importance, we can determine which features have a significant impact on the model's predictions. Impurity-based feature importances are a method for calculating feature importance, commonly used in decision tree and random forest models. Impurity-based feature importance provides a metric for measuring the contribution of features to the model's predictive performance; features with high impurity may contribute more to node splitting.

[0095] Specifically, firstly, a feature importance ranking matrix is ​​obtained by scoring and ranking the feature importance of the original data, maintaining the original data dimensions. The feature importance score reflects the contribution of each feature to the model performance. Then, the data dimensions are pruned based on the contribution of each feature to the target variable. This method not only selects and retains the most important features and removes redundant information, maximizing the complementarity of highly correlated features between different modalities, but also solves the problem of dimensional alignment in heterogeneous multimodal data, improving model performance and reducing the number of parameters. Finally, the dimension with the best performance in different datasets is used as the baseline for subsequent experiments. This ensures that the feature selection method can achieve good results in practical applications.

[0096] n m =w m *G m -w left *G left -w right *G right (13)

[0097]

[0098] M fi =sort(FI m (15) (descending=True)

[0099] Where n m w represents the impurity of a given node. m ,w left ,w right G represents the ratio of the number of training samples in node m and its left and right child nodes to the total number of training samples. m G left G right Let be the impurities of node m and its left and right child nodes, respectively. The importance FI of each feature is obtained from the node impurities. i Finally, the features are sorted to obtain the feature importance ranking matrix M. fi .

[0100] 4. Learning

[0101] The model's overall learning is achieved through minimization:

[0102] L all =L log +αL diff +βL rec +γL adv (16)

[0103]

[0104] In the formula, α, β, γ are the loss weights, and L all This determines the contribution of each regularization component to the overall loss. For classification tasks, the standard cross-entropy loss L is used. log Let N represent the spectrum N. b For details on the remaining loss modules, see section 3, Feature Importance Alignment. During the training phase, the total loss L is used. all Optimize BHTMF. PyTorch is a high-level neural network framework written in Python for implementing CAMR on Windows systems with NVIDIA GeForce RTX 3090 graphics processors.

[0105] Example 2.

[0106] The model from Example 1 was used in the experiment, and some of the processing steps were as follows:

[0107] (1) Dataset and Preprocessing

[0108] The samples used in this embodiment were from the Systemic Lupus Erythematosus (SLE) dataset from the People's Hospital of Xinjiang Uygur Autonomous Region and the Cancer dataset (Cancer) from the Affiliated Cancer Hospital of Xinjiang Medical University, which includes non-small cell lung cancer (NSCLC), renal cell carcinoma (RCC), and esophageal cancer. Informed consent was waived for all studies. All serum samples were obtained from fresh blood. The samples contained no anticoagulants and were collected in the morning after a 12-hour fast to avoid interference from diet or other factors on serum composition. 3 mL of fresh blood was collected from each sample. The blood samples were then centrifuged at 4000 rpm at 4°C to extract the uppermost clear liquid, obtaining serum. The serum was separated into centrifuge tubes and stored at -80°C for analysis. The SLE dataset was tested three times at different locations on each serum sample, and the Cancer dataset was tested five times. The average value of each sample was used for further experiments.

[0109] Because serum spectra acquired by the spectrometer are susceptible to interference from measurement conditions, detection environment, and hardware facilities, the analytical results can be significantly affected. Therefore, preprocessing operations were performed on both datasets. For the SLE dataset, serum Raman spectra underwent baseline correction using iterative adaptive weighted penalized least squares (airPLS) to subtract fluorescence background. The polynomial degree was set to 3, the number of iterations to 100, and the gradient of the polynomial loss to 0.001. Mean smoothing was then applied to smooth the spectra with a smoothing window width of 9. Serum infrared spectra underwent baseline correction using only airPLS. Data preprocessing was performed in Matlab R2022a. For the Cancer dataset, the convex rubber band correction method in OPUS spectral processing software was used to perform unified baseline correction on Raman and infrared spectra. The number of iterations was set to 5, and the number of baseline points was 64. Atmospheric compensation was also applied to the baseline-corrected spectra using CO2 compensation. Specific sample information is shown in Table 2.

[0110] For the original spectral data, the data is normalized by rescaling the data range to [0,1] using a min-max method. The following is the definition of min-max normalization:

[0111]

[0112] x max and x min These are the maximum and minimum values, respectively. Within the range [a, b], minimum-maximum normalization converts the numerical values ​​to x′.

[0113] Table 2: Raman and infrared spectral sample information, acquisition information, original spectral dimensions and ethical number.

[0114]

[0115]

[0116] (2) Baselines

[0117] To verify the generalization ability of the BHTMF model, a large number of comparative experiments were conducted using classic deep learning models. These baselines included classic convolutional neural network architectures and classic neural network architectures for multimodal fusion.

[0118] LeNet: LeNet is a classic convolutional neural network model in deep learning, proposed by Yann LeCun et al. in 1998. It uses convolutional layers and pooling layers to capture the local features and spatial structure of the input image, and is one of the important milestones in deep learning.

[0119] AlexNet: AlexNet is another significant milestone in the field of deep learning, proposed by Alex Krizhevsky et al. in 2012. It was the first convolutional neural network to achieve a significant breakthrough on the large-scale image dataset ImageNet. AlexNet contains multiple convolutional layers, pooling layers, and fully connected layers, and uses Dropout and ReLU activation functions. Its success marked the rise of deep learning in the field of computer vision.

[0120] VGG-19: VGG-19 is a very deep convolutional neural network model proposed by Karen Simonyan and Andrew Zisserman in 2014. It uses multiple convolutional and pooling layers to extract features from images layer by layer, learning more complex and abstract image features, which makes the network perform well in image recognition and classification tasks.

[0121] ResNet: ResNet is a deep residual network proposed by Kaiming He et al. in 2015. ResNet addresses the training difficulty of deep networks by introducing residual blocks, allowing the network to reach deeper layers. Each residual block contains skip connections, enabling the model to optimize better during training. ResNet's feature extraction layers allow the network to learn feature representations at different scales, making it advantageous for handling tasks with various image sizes and complexities.

[0122] DenseNet: DenseNet is a densely connected convolutional neural network proposed by Gao Huang et al. in 2017. It introduces dense connections between convolutional layers, allowing all feature maps from the previous layer to be directly passed to the next. This dense connection promotes feature reuse and information transfer, making the network more compact and efficient. DenseNet has achieved excellent performance in many image recognition tasks.

[0123] EF-LSTM: EF-LSTM is a multimodal fusion model based on early fusion proposed by Jennifer Williams et al. in 2018. By concatenating the original features of multimodal data and then inputting them into an LSTM, it captures the long-term correlation between long multimodal sequences and has achieved good results in sentiment analysis tasks.

[0124] LF-DNN: LF-DNN is a multimodal fusion model based on late fusion proposed by Wenmeng Yu et al. in 2020. It learns unimodal features by using DNN and then concatenates their representations as input to the prediction layer.

[0125] TFN : Tensor Fusion Network (TFN) is a model proposed by Amir Zadeh et al. in 2017. This model uses tensor fusion to form unimodal, bimodal, and trimodal interactions from multimodal data, and the resulting multimodal tensors are used for sentiment analysis tasks.

[0126] LMF: LMF is a low-rank multimodal fusion (LMF) method proposed by Zhun Liu et al. in 2018. It is an improvement based on tensor fusion network. This method uses low-rank tensors for multimodal fusion, which greatly reduces the computational complexity and improves the efficiency of model execution.

[0127] (3) Model evaluation indicators

[0128] The model evaluation metrics used the following five standard measures: accuracy, precision, recall, F1 score, and AUC.

[0129]

[0130]

[0131]

[0132]

[0133] Among them, TP is a true positive, TN is a true negative, FP is a false positive, and FN is a false negative.

[0134] The ROC curve is plotted with the model's true positive rate on the ordinate and the false positive rate on the abscissa. AUC represents the area under the ROC curve; a larger AUC value indicates better model performance and a better ability to distinguish between positive and negative samples.

[0135] (4) Settings

[0136] In this embodiment, a stratified sampling five-fold cross-validation method was used to divide the dataset into training and test sets in an 8:2 ratio. Stratified sampling ensures that the proportion of samples from different categories in each subset remains consistent with the proportion in the original dataset. The experimental data used the original full-dimensional data and the feature importance alignment dimension data mentioned in Section 3. Furthermore, the optimal dimension in the feature selection gradient experiment was used as the benchmark for subsequent experiments.

[0137] A consistent initial learning rate of 0.00001 was used with the Adam optimizer in all experiments. Specifically, for the SLE dataset, there were 70 hidden units, 0.1 dropout, and 100 epochs. For the Cancer dataset, there were 80 hidden units, 0.1 dropout, and 200 epochs. In both multi-class classification tasks, 3-class accuracy (Acc-3), 4-class accuracy (Acc-4), and class accuracy (Acc-HC, Acc-LN, Acc-SLE, Acc-LC, Acc-KC, Acc-EC) are reported, representing the control group, lupus nephritis, systemic lupus erythematosus, non-small cell lung cancer, renal cell carcinoma, and esophageal cancer, respectively. Precision, recall, F1 score, and AUC are also reported. For all mentioned metrics, higher values ​​indicate better performance.

[0138] (5) Results and Analysis

[0139] ① Model Comparison Analysis

[0140] To demonstrate the effectiveness of the BHTF, CMIL, and RFIA modules, extensive experiments were conducted on the Systemic Lupus Erythematosus (SLE) and Cancer datasets. First, the RFIA module was used to perform feature selection on the bimodal spectral feature dimensions, cropping the feature range to 10-100 dimensions. Then, using the original feature dimensions and the selected feature dimensions as experimental datasets, full-dimensional experiments and feature gradient experiments were performed. The main task module containing only BHTF was tested first, followed by the main and auxiliary task modules combining BHTF and CMIL, to fully verify the model's generalization ability. For each feature selection dimension in the feature gradient experiments, the number of dimensions was equal to the number of hidden layer units in the model, while other hyperparameters remained unchanged. Finally, the dataset with the optimal dimensions from the feature gradient experiments was used as the dataset for subsequent ablation experiments. The baseline model mentioned in Part IV was used as a comparison model to fully verify the advantages of the BHTMF model.

[0141] Table 3-4 presents the experimental results for single-mode and dual-mode spectra on the SLE and Cancer datasets. It can be noted that BHTMF outperforms single-mode experiments in both dual-mode and dual-mode experiments, demonstrating the effectiveness of dual-mode spectral data fusion and showing significant performance.

[0142] Table 5-6 presents the experimental results of feature importance alignment gradients for bimodal spectral fusion on the SLE and Cancer datasets. The results show that RFIA can efficiently handle high-dimensional data fusion, significantly reduce model complexity, and substantially improve model performance. It achieves the best performance in systemic lupus erythematosus classification with 70 feature selection dimensions, showing a 3.35% improvement over all dimensions with the BHTF model, and a 4.54% improvement with the BHTF combined with CMIL model. For cancer classification with 80 feature selection dimensions, it achieves the best performance, showing a 10.07% improvement over all dimensions with the BHTF model, and a 10.1% improvement with the BHTF combined with CMIL model.

[0143] Tables 7-8 present the comparative experimental results of bimodal spectral fusion models on the SLE and Cancer datasets. While deep convolutional models such as DenseNet, ResNet, and AlexNet exhibit good performance, they are highly complex and have a large number of parameters. BHTMF has 2.75M and 3.54M parameters on the two datasets, respectively, which is less than the number of parameters in classic deep convolutional models. The performance difference between LSTM based on early fusion and DNN based on late fusion is significant, possibly because it does not fully understand the interaction information between modalities. For the tensor fusion network series TFN and LMF, although the parallel decomposition of LMF can significantly reduce model complexity, it does not fully consider the overall intermodal information interaction in terms of fusion efficiency, and its performance is inferior to that of the tensor fusion network TFN. BHTMF demonstrates the best performance in all comparative experiments.

[0144] Table 3-4 shows the classification performance of the BHTMF model for single-mode and dual-mode spectral features on the SLE and cancer datasets. R represents Raman spectra, I represents Infrared spectra, and R&I represents Raman spectra with Infrared spectra.

[0145] Table 3

[0146]

[0147] Table 4

[0148]

[0149] Table 5-6: Feature importance aligned gradient experiments on SLE and cancer datasets, using dual-mode spectral fusion features. BHTF represents using only the main task module, and BHTF&CMIL represents the combination of main and auxiliary tasks.

[0150] Table 5

[0151]

[0152] Table 6

[0153]

[0154]

[0155] Table 7-8: Comparison analysis of baseline models on SLE and cancer datasets. Using dual-mode spectral fusion features, BHTMF showed the best performance in all cases.

[0156] Table 7

[0157]

[0158] Table 8

[0159]

[0160]

[0161] ② Ablation Research

[0162] To demonstrate the effectiveness of multispectral fusion, extensive experiments were conducted in two aspects to verify the effects of different combinations of single-mode and dual-mode spectral features. One type of experiment was single-mode spectral classification, and the other was dual-mode spectral classification. Validation was performed on both the SLE and Cancer datasets. In the ablation experiments, we report the main metrics: three-class accuracy (Acc-3), four-class accuracy (Acc-4), class accuracy, F1 score (F1), and AUC. For TFN, LMF, EF-LSTM, and LF-DNN, which were modified to use dual-mode input models, their respective tensor fusion, early fusion, and late fusion methods were used in dual-mode spectral experiments. In single-mode spectral experiments, the model was modified to use single-mode input models to ensure accuracy. For LeNet, AlexNet, VGG-19, ResNet, and DenseNet, feature concatenation was used as the model input in dual-mode spectral experiments, while single-mode spectral features were used as the model input in single-mode spectral experiments. Furthermore, we conducted full-dimensional experiments and feature importance alignment experiments, with data settings of Unaligned and Aligned, respectively. This was done to avoid causing high-dimensional input to the model, taking into account the high-dimensional nature of the spectrum, and to ensure more complete and direct interaction of dual-mode spectral feature information.

[0163] 1) Single-modal classification experiment

[0164] In this type of experiment, Raman and infrared spectra were used as single-peak features input into the model. Tables 9-10 show that in the three-class classification experiment on the SLE dataset, LF-DNN exhibits excellent performance on both single-peak Raman and infrared spectra. This may be because simple fully connected neural networks are more suitable for information mining from single-mode spectra. The optimal accuracy is 93.79% for single-mode Raman and 76.21% for single-mode infrared. Tables 11-12 show that in the four-class classification experiment on the Cancer dataset, BHTMF shows good performance compared to other models with aligned single-peak Raman and infrared spectra, with accuracies of 76.63% and 93.51%, respectively. However, without aligned single-peak features, LF-DNN and ResNet show the best performance, with accuracies of 77.29% and 93.16%, respectively. We infer that this is due to the high dimensionality of cancer data and the large number of classification categories. Furthermore, single-peak experiments revealed that Raman spectroscopy data showed superior diagnostic advantages in the SLE dataset, while infrared spectroscopy was superior in the Cancer dataset. This also provides strong support for the adversarial loss module in the CMIL module during dual-mode experiments. Aligned models significantly reduce model complexity and memory overhead, while also improving model performance in some experiments.

[0165] Table 9-10: Comparison analysis of baseline models on the SLE dataset, using single-mode Raman spectroscopy and single-mode infrared spectroscopy respectively, with BHTMF as our model.

[0166] Table 9

[0167]

[0168]

[0169] Table 10

[0170]

[0171] Table 11-12: Comparison analysis of baseline models on the cancer dataset, using single-mode Raman spectroscopy and single-mode infrared spectroscopy respectively, with BHTMF as our model.

[0172] Table 11

[0173]

[0174]

[0175] Table 12

[0176]

[0177] 2) Bimodal classification experiment

[0178] In this type of experiment, Raman and infrared spectroscopy were combined as dual-mode features input into the model. Tables 13-14 show that in the three-class classification experiment on the SLE dataset and the four-class classification experiment on the Cancer dataset, BHTMF with aligned dual-mode fusion features exhibited excellent performance, achieving accuracies of 97.22% and 94.82%, respectively. However, with unaligned dual-mode fusion features, BHTMF showed the best performance on the SLE dataset, achieving an accuracy of 92.68%, while DenseNet showed the best performance on the Cancer dataset, achieving an accuracy of 91.56%. In summary, in multimodal fusion work, the results indicate that if features with high contribution and complementary information are fully extracted from the original features, the interaction between the two peaks may be more conducive to information fusion. RFIA addresses this issue, and the aligned feature dataset can better improve model performance. This also demonstrates that fully considering the interactions within, between, and between the two peaks, as well as improving the model's cross-modal expressive ability, can lead to better performance. This also verifies that the BHTMF model can effectively integrate multimodal data by combining the BHTF and CMIL main and auxiliary task modules, demonstrating the advantages of BHTMF in multimodal tasks.

[0179] Table 13-14: Comparison analysis of baseline models on SLE and cancer datasets, using dual-mode spectral fusion features, with BHTMF as our model.

[0180] Table 13

[0181]

[0182]

[0183] Table 14

[0184]

[0185] This invention, considering multimodal fusion, utilizes two types of spectral data, focusing on the information interaction between multimodalities and the selection of fusion strategies. It proposes a high-order tensor multimodal fusion disease method and model based on bimodal interaction learning (BHTMF) to assist in medical data analysis and diagnosis. This method effectively achieves multimodal data information fusion by obtaining high-order interaction fusion features of bimodal spectral information through BHTMF, and improves diagnostic model performance by learning the heterogeneity between multimodal data through CMIL cross-modal learning. BHTMF was evaluated on the Systemic Lupus Erythematosus (SLE) and Cancer datasets, and the model's principles and key aspects of the multimodal fusion classification task are intuitively explained. Results show that the model is effective and outperforms many existing classic deep learning models.

[0186] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A high-order tensor multi-modal fusion method based on dual-mode spectral interaction learning, characterized in that, The method comprises the following steps: S10: inputting Raman spectrum and infrared spectrum data; S20: learning the interaction between the two modalities by high-order tensor outer product to make the non-cascaded multi-modal spectrum fusion representation of the Raman spectrum and infrared spectrum data; The specific processing of step S20 is: first, using a unified pre-alignment encoder for non-linear mapping to obtain the fusion representation of the two modal spectrum; then, using high-order tensor fusion, the different spectrum representations are respectively upgraded to three-dimensional space in the form of rows and columns according to the high-order matrix multiplication rule, and batch matrix multiplication is used for efficient fusion; finally, the non-cascaded fusion features of high-order interaction are mapped back to a low-dimensional space; the specific processing steps are: ① by an encoder E having a multi-layer fully connected architecture c learning a non-linear representation z of different spectral modalities m which is formulated as follows: z m = E c (h m ; w), m e {r, i} In the formula, E c represents the encoder function of the dual-mode spectral pre-fusion, w represents the parameters of E c , and {r, i} respectively represent the Raman spectrum and infrared spectrum data. ②The two modal feature vectors are respectively upgraded in the second and third dimensions to adapt to batch matrix multiplication in a high-dimensional space, forming a 3-D fusion tensor of the high-order interaction of the two modal spectrum, and fully learning the inter-modal dynamic modeling, and the formula is as follows: z m' = [z m ,1], m e {r,i}, where denotes the outer product between vectors, z f ∈R (r+1)*(i+1) is a 3-D tensor cube of all possible combinations of bimodal spectra iii. The non-cascaded fusion tensor z f is fed into the decoder and classifier for task classification, whose formula is as follows: z' f = D c (z f ; w), In the formula, w represents D c the parameters of z′ f represent the fusion feature representation after the full connection layer decoder. S30: calculating the orthogonality loss, reconstruction loss and adversarial loss to compensate for the heterogeneity difference between different spectra by unique representation learning of the Raman spectrum features and infrared spectrum features; In the step S30, a shared representation encoder of the dual-mode modalities is constructed obtained particular representation encoder obtained encoder equation m∈{r,i},n∈{c,u}, wherein n respectively represents a shared encoder and a particular encoder, and then a orthogonality loss, a reconstruction loss and an adversarial loss are calculated The orthogonality loss L diff The calculation formula is: The reconstruction loss L rec The calculation formula is: The calculation formula of the confrontation loss L adv The calculation formula of the confrontation loss L adv = L f + L t where L f is the false adversarial loss, L t is the true adversarial loss, is the predicted value of y by the discriminator.

2. The high-order tensor multi-modal fusion method according to claim 1, wherein the Raman spectrum and infrared spectrum data are subjected to baseline correction, smoothing processing and normalization processing before being inputted.

3. The high-order tensor multi-modal fusion method according to claim 1, wherein the Raman spectrum and infrared spectrum data are Raman spectrum and infrared spectrum data of serum samples.

4. The high-order tensor multi-modal fusion method according to claim 1, wherein the high-order tensor multi-modal fusion method further comprises a step S15 of feature importance alignment, which calculates the feature importance of the Raman spectrum and infrared spectrum data and performs feature selection.

5. The high-order tensor multi-modal fusion method according to claim 3, wherein the specific processing of step S15 is: first, scoring and sorting the feature importance degree of the original data to obtain a feature importance sorting matrix that maintains the original data dimension; then, according to the contribution degree of each feature to the target variable, the data dimension is trimmed; finally, the optimal performance dimension in different data sets is taken as the baseline. The high-order tensor multi-modal fusion device is used to implement a high-order tensor multi-modal fusion model based on the interaction learning of two modal spectrums. The high-order tensor multi-modal fusion model is obtained by using the high-order tensor multi-modal fusion method according to any one of claims 1-5. The device comprises an input module, a feature importance alignment module, a two-modal high-order tensor fusion module and a cross-modal interaction learning module. The input module is used to input Raman spectrum and infrared spectrum data.

6. A high-order tensor multi-modal fusion device based on dual-mode spectral interaction learning, characterized in that, The feature importance alignment module is used to calculate the feature importance of the Raman spectrum and infrared spectrum data and perform feature selection. The two-modal high-order tensor fusion module learns the interaction between the two modalities by high-order tensor outer product to make the non-cascaded multi-modal spectrum fusion representation of the Raman spectrum and infrared spectrum data.

7. The high-order tensor multi-modal fusion apparatus of claim 6, ​ ​ ​ ​ ​ The cross-modal interaction learning module is composed of orthogonality loss, reconstruction loss and adversarial loss, and the heterogeneity difference between different spectra is compensated by unique feature learning of the Raman spectrum features and infrared spectrum features.

Citation Information

Patent Citations

  • Target detection method of multispectral double-flow network

    CN116486233A

  • Multi-mode sentiment analysis method and device based on dual-mode multi-granularity interaction and medium

    CN116912642A