Data processing method and device based on multi-modal contrast fusion and graph neural network, and medium
By mapping multimodal biomedical data to a unified space and performing cross-modal semantic alignment, and combining Transformer and graph convolutional networks, the problem of difficulty in mining complementary information between modalities in multimodal data processing is solved, achieving efficient data fusion and improved classification performance.
Patent Information
- Application Number
- CN202510894605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-28
AI Technical Summary
Existing multimodal biomedical data processing methods struggle to effectively integrate heterogeneous data, lack intermodal nonlinear complementary information mining, and are difficult to construct consistent representations across modalities. Consequently, the fused features fail to fully preserve the semantic correlations and sample association patterns of each modality.
By mapping three modalities of data—mRNA expression, DNA methylation, and miRNA expression—to a unified latent space, cross-modal semantic alignment is performed using the NT-Xent contrastive learning mechanism. Furthermore, the Transformer multi-head self-attention mechanism is combined to model high-order interactions between modalities, constructing a sample similarity graph structure. Finally, a graph convolutional network is used to extract high-order adjacency features of nodes for classification.
It achieves efficient semantic fusion and complementary information mining of multimodal data, improves the deep analysis capability and classification performance of medical data, and significantly enhances the robustness and generalization ability of the model, especially in scenarios with scarce labels.
Smart Images

Figure CN121034429A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to medical data processing systems, and more particularly to a data processing method, device, and medium based on multimodal contrast fusion and graph neural networks. Background Technology
[0002] In recent years, the explosive growth of multimodal biomedical data has brought new opportunities and challenges to biomedical information processing. Compared with single-modal data, multi-source heterogeneous data can reflect the characteristics of biological samples from different dimensions, providing a rich information foundation for uncovering potential biomedical laws. However, how to efficiently integrate these heterogeneous data and overcome significant differences in data dimensions, scales, and semantics has become a core challenge in the field of multimodal data processing.
[0003] Early multimodal data fusion methods primarily employed techniques such as ensemble learning, nonnegative matrix factorization, and multi-kernel learning. For example, ensemble learning improves model performance by combining feature representations from different modalities, nonnegative matrix factorization achieves feature dimensionality reduction and fusion through low-rank approximation, and multi-kernel learning uses kernel functions to map different modalities to a shared space. While these methods have achieved the integration of multimodal features to some extent, they are mostly limited to the independent representation of features within a modality, lacking effective mining of nonlinear complementary information between modalities. Consequently, the fused features fail to fully preserve the semantic correlation between the various modalities.
[0004] To address these challenges, Graph Neural Networks (GNNs) have gradually become an important tool for multimodal data fusion, especially Graph Convolutional Networks (GCNs). GCNs, by constructing graph structures between samples to model potential similarities, can aggregate information from adjacent nodes, demonstrating certain advantages in biomedical data processing. However, existing GCN-based methods still have significant shortcomings: on the one hand, they still focus on the independent processing of features within a modality, failing to establish high-order interaction models between modalities, thus limiting the in-depth mining of complementary information across multiple modalities; on the other hand, the inherent distributional differences in multimodal biomedical data, such as the fundamental differences in feature dimensions and scales between imaging data and molecular data, make unified representation difficult to achieve, and existing methods struggle to construct consistent cross-modal representations while preserving modality-specific features. Furthermore, most models only focus on the feature representation of individual samples, ignoring the group structural relationships between samples, resulting in insufficient ability to capture potential sample association patterns in biomedical data. Therefore, the question is how to perform more efficient multimodal data fusion based on the characteristics of medical data, to achieve intermodal interaction modeling, unified representation of heterogeneous data, and utilization of sample structure relationships, thereby realizing the deep integration and semantic mining of multimodal biomedical data. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a data processing method, device and medium based on multimodal contrast fusion and graph neural network. By mapping three modal data of mRNA expression, DNA methylation and miRNA expression to a unified latent space and performing cross-modal semantic alignment, it effectively integrates multi-dimensional biomedical features such as transcriptome, epigenetics and non-coding RNA, retains the specificity of each modality and mines complementary information between modalities.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] According to one aspect of the present invention, a data processing method based on multimodal contrast fusion and graph neural networks is provided, the specific steps of which include:
[0008] S1. Collect multimodal biomedical dataset samples including mRNA expression data, DNA methylation data, and miRNA expression data;
[0009] S2. Each modality feature vector is mapped to a latent space of a unified dimension through an independent feedforward neural network encoder to obtain the embedding matrices of the three modalities. An NT-Xent-based contrastive learning mechanism is used to perform unsupervised contrastive alignment on the embedding matrices of different modalities. By maximizing the similarity between homologous modal pairs and minimizing the similarity between heterologous modal representations, the aligned embedding matrices of the three modalities are generated.
[0010] S3. Stack the aligned modal embeddings into a sequence, input them into a multi-layer Transformer encoder, model the high-order interactions between modalities through a multi-head self-attention mechanism, and obtain a unified fusion representation through modal dimension average pooling; normalize the fusion representation, calculate the cosine similarity matrix between samples, retain several similar neighbors of each sample to generate a sparse adjacency matrix, and construct a sample similarity graph structure.
[0011] S4. Based on the graph structure and fusion representation, high-order adjacency features of nodes are extracted through a three-layer graph convolutional network, and sample classification is completed by combining two fully connected layers. The probability of a sample under different classification results is output, where the classification loss function is a joint optimization of cross-entropy loss and contrast loss.
[0012] Furthermore, the mRNA expression data dimension is N×G1, where G is the number of detected genes and N is the number of samples, and each element represents the standardized expression level of a gene in the corresponding sample; the DNA methylation data dimension is N×G2, where G2 is the total number of CpG sites detected by the sequencing platform and N is the number of samples, and each element represents the methylation level of a sample at a specific site using a β value; the miRNA expression data dimension is N×G3, where G3 is the number of detected miRNAs and N is the number of samples, and each element is a standardized expression value representing the expression level of the corresponding miRNA in the sample.
[0013] Furthermore, in S2, the feedforward neural network encoder includes a normalization layer and a ReLU activation function, which maps the original features of each modality from the original dimension to a unified dimension.
[0014] Furthermore, in S2, based on the NT-Xent contrastive learning mechanism, the embedding matrices of any two modalities from the three modalities of mRNA, DNA, and miRNA of any sample are selected to form a modality pair. After normalizing the embedding matrices of the modality pair, the cosine similarity matrix S is calculated, and the expression is:
[0015]
[0016] in, and Let be the embedding matrix of mode a of sample i and mode b of sample j, and τ be a parameter that controls the sharpness of the similarity distribution.
[0017] Furthermore, the NT-Xent contrastive learning mechanism utilizes a multimodal contrastive loss function. To bring the representations of the same sample closer together and increase the distance between the modal representations of different samples, the expression is:
[0018]
[0019] Where M is the number of modes, with a value of 3. Let be the cross-entropy loss term between modes a and b.
[0020] Furthermore, in step S3, the aligned modes are embedded. Stacked as a sequence The input is a Transformer encoder, and the output is obtained by average pooling to obtain the fused representation z. i :
[0021]
[0022] Where M is the number of dimensions, Stacked as a sequence.
[0023] Furthermore, in S3, for each row of the adjacency matrix, its Top-K similar neighbors are retained, and the remaining elements are set to zero, thereby generating a sparse adjacency matrix, expressed as:
[0024]
[0025] Among them, A ij For elements in the sparse adjacency matrix, For the set of indexes of the K samples that rank K most similar to sample i, the diagonal elements A ii Set to zero, that is
[0026] Furthermore, the input to the graph convolutional network is the fused representation and the sparse adjacency matrix. The graph convolution operation involves extracting high-order adjacency features in one pass through a three-layer graph convolutional network, mapping them to the class space through two fully connected layers, and generating predicted probabilities after Softmax activation. The objective of the loss function of the graph convolutional network is to minimize the cross-entropy loss.
[0027]
[0028] Where N is the number of samples and C is the number of categories. Let be the predicted probability of sample i belonging to category c. One-hot encoding for the actual label.
[0029] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.
[0030] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] (1) Efficient semantic fusion and complementary information mining of multimodal data: By mapping three modal data of mRNA expression, DNA methylation and miRNA expression to a unified latent space through independent encoders, and combining the NT-Xent contrastive learning mechanism to achieve cross-modal semantic alignment, it can effectively integrate multi-dimensional biomedical features such as transcriptome, epigenetics and non-coding RNA, retain the specificity of each modality while mining complementary information between modalities, and provide a more comprehensive feature representation for in-depth analysis of medical data.
[0033] (2) Dynamic interactive modeling and cross-modal feature fusion optimization: The Transformer multi-head self-attention mechanism is used to perform high-order interactive modeling on the aligned modal embedding. Compared with traditional static splicing or weighted fusion, it can dynamically capture the nonlinear correlation between mRNA, DNA methylation and miRNA data, generate a unified fusion representation containing cross-modal semantic commonality and complementary characteristics, and improve the fusion efficiency and representation ability of multi-source heterogeneous medical data.
[0034] (3) Sample population structure modeling enhances the performance of medical data discrimination: By constructing a sample similarity graph structure and combining it with a graph neural network for node classification, the potential similarity between patients is transformed into graph structure information. When processing multimodal data such as mRNA, DNA methylation and miRNA, the model not only focuses on individual sample characteristics, but also uses the population structure relationship to optimize the classification results, which significantly improves the robustness and generalization ability of medical data processing. Attached Figure Description
[0035] Figure 1 This is a data flow graph based on a data processing method using multimodal contrast fusion and graph neural networks. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0037] like Figure 1 The diagram illustrates the data flow of a data processing method based on multimodal contrast fusion and graph neural networks. The specific steps of the method include:
[0038] S1. Collect samples from a multimodal biomedical dataset that includes mRNA expression data, DNA methylation data, and miRNA expression data;
[0039] S2. Each modality feature vector is mapped to a latent space of a unified dimension through an independent feedforward neural network encoder to obtain the embedding matrices of the three modalities. An NT-Xent-based contrastive learning mechanism is used to perform unsupervised contrastive alignment on the embedding matrices of different modalities. By maximizing the similarity between homologous modal pairs and minimizing the similarity between heterologous modal representations, the aligned embedding matrices of the three modalities are generated.
[0040] S3. Stack the aligned modal embeddings into a sequence, input them into a multi-layer Transformer encoder, model the high-order interactions between modalities through a multi-head self-attention mechanism, and obtain a unified fusion representation through modal dimension average pooling; normalize the fusion representation, calculate the cosine similarity matrix between samples, retain several similar neighbors of each sample to generate a sparse adjacency matrix, and construct a sample similarity graph structure.
[0041] S4. Based on graph structure and fusion representation, a three-layer graph convolutional network is used to extract high-order adjacency features of nodes. Combined with two fully connected layers, sample classification is completed, and the probability of the sample under different classification results is output. The classification loss function is a joint optimization of cross-entropy loss and contrastive loss.
[0042] The mRNA expression data is N×G1, where G is the number of detected genes and N is the number of samples. Each element represents the standardized expression level of a gene in the corresponding sample, and has been standardized (e.g., TPM, FPKM, or log2(count+1)). This data reflects the transcriptional activity status of genes in each sample and is used to explore disease-related gene expression characteristics.
[0043] DNA methylation data is presented in an N×G2 dimension, where G2 represents the total number of CpG sites detected by the sequencing platform, N is the number of samples, and each element is represented by a β value indicating the methylation level of a specific site in a given sample. This data is used to assess the epigenetic status of gene regulatory regions and is of great significance for transcriptional regulatory mechanisms and disease subtype classification.
[0044] The miRNA expression data dimension is N×G3, where G3 is the number of detected miRNAs, N is the number of samples, and each element is a standardized expression value, representing the expression level of the corresponding miRNA in the sample. miRNAs play a crucial role in the development and progression of various diseases by regulating mRNA expression.
[0045] A multimodal biomedical dataset containing mRNA expression data, DNA methylation data, and miRNA expression data from N patient samples. in, This represents the input feature vector of the i-th sample in the m-th modality, where M is the total number of available modalities, which is 3, and d m Let be the feature dimension of the m-th modality. The corresponding label y i ∈{1,2,…,C} represents the categories of the sample, and C represents the total number of categories.
[0046] In S2, the feedforward neural network encoder includes a normalization layer and a ReLU activation function, mapping the original features of each modality from their original dimensions to a unified dimension. For the m-th modality, its encoding function... Represented as:
[0047]
[0048] The input features are derived from their original dimension d. m Mapping to a unified latent space dimension d is a prerequisite for achieving semantic alignment between different modalities. For the i-th sample, its modality m embedding is represented as:
[0049]
[0050] This encoder is composed of a feedforward neural network, which combines normalization and nonlinear activation to improve modeling ability and representation robustness.
[0051] In the NT-Xent contrastive learning mechanism, any two modalities from the three modalities of mRNA, DNA, and miRNA in any sample are selected to form a modality pair. After normalizing the embedding matrices of the modality pair, the cosine similarity matrix S is calculated, expressed as:
[0052]
[0053] in, and Let be the embedding matrix of mode a of sample i and mode b of sample j, and τ be a parameter that controls the sharpness of the similarity distribution.
[0054] The NT-Xent contrastive learning mechanism uses a multimodal contrastive loss function. By bringing the representations of the same sample closer together across different modalities and widening the distance between the modal representations of different samples, the model's ability to model modal consistency and individual discriminability is improved. The expression is as follows:
[0055]
[0056] Where M is the number of modes, with a value of 3. Let be the cross-entropy loss term between modes a and b.
[0057] In S3, the aligned modalities are embedded. Stacked as a sequence The input is a Transformer encoder, and the output is obtained by average pooling to obtain the fused representation z. i :
[0058]
[0059] Where M is the number of dimensions, Stacked as a sequence.
[0060] In S3, for each row of the adjacency matrix, its Top-K similar neighbors are retained, and the remaining elements are set to zero, thus generating a sparse adjacency matrix, expressed as:
[0061]
[0062] Among them, A ij For elements in the sparse adjacency matrix, For the set of indexes of the K samples that rank K most similar to sample i, the diagonal elements A ii Set to zero, that is
[0063] The graph convolutional network takes a fused representation and a sparse adjacency matrix as input. The graph convolution operation involves extracting high-order adjacency features in one pass through a three-layer graph convolutional network, mapping them to the class space through two fully connected layers, and generating predicted probabilities after softmax activation. The graph convolutional network consists of three graph convolutional modules and two fully connected layers. The former is used to extract high-order adjacency structure features, and the latter is used for node-level classification prediction.
[0064] The graph convolution operation takes the following form:
[0065]
[0066] in, W is the normalized adjacency matrix. (l) Let be the learnable parameters of the l-th layer, and σ(·) be the ReLU activation function. The final representation is the predicted probability distribution of each node after mapping through the fully connected layer:
[0067]
[0068] Where C represents the total number of categories.
[0069] The objective of the loss function in graph convolutional networks is to minimize the cross-entropy loss.
[0070]
[0071] Where N is the number of samples and C is the number of categories. Let be the predicted probability of sample i belonging to category c. One-hot encoding of the ground truth labels. Cross-entropy loss optimizes node representations under graph structure guidance, thereby achieving efficient and robust patient-level classification.
[0072] Throughout the entire data processing network, end-to-end optimization is performed on all parameters Θ, including the contrastive coding layer, graph convolutional layer, and fully connected layer. The total loss function is:
[0073]
[0074] This loss function drives the graph neural network to learn node representations that are both discriminative and structurally consistent based on multimodal fusion features, thereby achieving accurate sample-level, i.e., patient-level classification.
[0075] To verify the performance of this embodiment, experiments were conducted on four public datasets (BRCA, LGG, KIPAN, and ROSMAP), and the effects of using the multimodal contrastive feature fusion mechanism versus the feature concatenation method alone were analyzed and compared. All experiments were conducted according to the predetermined partitioning scheme of each dataset. The experimental results are shown in Table 1. As can be seen from the comparison results in Table 1, the multimodal contrastive feature fusion mechanism significantly outperforms the simple feature concatenation method on all datasets, showing varying degrees of improvement in multiple evaluation metrics such as accuracy, weighted F1 score, and AUC (Area Under the Curve). Taking the BRCA dataset as an example, this mechanism improved accuracy and F1 score by 2.52% and 2.41%, respectively, and also improved AUC by 0.50%. The improvement is even more significant on the ROSMAP dataset, with accuracy and F1 score increasing by 4.00% and 4.14% respectively, and AUC improving by 5.57%, indicating that the mechanism has a stronger discriminative ability for highly heterogeneous neurodegenerative disease data.
[0076] Experimental results across various datasets demonstrate that the proposed multimodal contrastive fusion mechanism effectively enhances the mining of intermodal complementarity and the learning of representational consistency while maintaining representational capabilities, thereby improving multimodal classification performance and verifying the advantages and broad applicability of this invention in multi-omics disease modeling.
[0077] Table 1. Experimental comparison results between this embodiment and the method using only feature stitching.
[0078]
[0079] This embodiment effectively aligns homologous representations of different modalities without labels by introducing a contrastive loss based on NT-Xent, enhancing modality consistency and demonstrating strong unsupervised modality alignment capabilities, making it suitable for scenarios with scarce labels. Furthermore, it employs a Transformer attention mechanism to achieve sample-level modality weight modeling, which, compared to static concatenation or weighted fusion, more fully exploits the nonlinear interactions and complementary characteristics between modalities, and offers a more flexible dynamic fusion mechanism. It introduces graph structure information to model the relationships between samples, transforming the similarity between patients into a graph structure and using GCN for structure-aware classification. This allows the model to not only focus on individual features but also integrate group structural information, improving overall discriminative performance. Finally, it introduces a reconstruction loss through a modality decoder to assist in latent representation learning, effectively enhancing the model's ability to preserve original features and generalize, thus improving robustness.
[0080] In biomedical research, mRNA expression, DNA methylation, and miRNA expression reflect different levels of biological characteristics, such as gene transcription, epigenetic modification, and non-coding RNA regulation. Cross-modal fusion can overcome the limitations of single-modal data representation. This approach maps three types of data to a unified space using independent encoders and achieves cross-modal semantic alignment by combining NT-Xent contrastive learning. This enables the capture of complex regulatory networks in biological samples from multiple dimensions, including the transcriptome and epigenome. For example, in tumor subtype classification, this fusion mechanism can simultaneously integrate abnormal gene expression and changes in methylation modification, providing a more comprehensive feature representation for revealing the multi-stage molecular mechanisms of cancer development. Compared to single-modal analysis, it is easier to discover potential biomarker combinations. However, there are complex nonlinear regulatory relationships between gene expression and epigenetic modification (such as the inhibitory effect of methylation on mRNA expression and the silencing effect of miRNA on genes), and traditional feature splicing methods struggle to capture such high-order interactions. This approach uses a Transformer multi-head self-attention mechanism to dynamically weight intermodal interactions, adaptively learning the regulatory weights between mRNA, methylation, and miRNA data. Taking neurodegenerative disease research as an example, this mechanism can accurately capture the spatiotemporal correlations between β-amyloid protein-related gene expression and epigenetic modifications, generating fusion representations containing cross-modal semantic dependencies. This provides algorithmic support for analyzing molecular interaction networks in disease progression and significantly improves the accuracy of complex disease classification. Furthermore, biomedical samples (such as patient tissues) are highly heterogeneous, and clinical data often suffers from small sample sizes and high dimensionality, making it difficult for traditional models to utilize the potential similarities between samples. This approach, by constructing a sample similarity graph structure and combining it with GCN modeling, can transform the correlations of clinical characteristics and molecular phenotypes of patient groups into graph topological information. For example, in rare disease research, this method can construct disease association networks through the molecular feature similarities of a small number of samples, assisting in the discovery of patient subgroups with similar pathological mechanisms. This provides a new approach for disease classification and prognosis prediction in small sample scenarios, effectively alleviating the problem of insufficient model generalization ability caused by sample scarcity in biomedical research.
[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0082] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0083] Multiple components in the device are connected to an I / O interface, including: input units such as a keyboard, mouse, etc.; output units such as various types of displays, speakers, etc.; storage units such as disks, optical disks, etc.; and communication units such as network interface cards, modems, wireless transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. The processing unit performs the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention may be implemented as a computer software program tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or the communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the method of the present invention described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the method of the present invention by any other suitable means (e.g., by means of firmware).
[0084] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0085] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0086] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0087] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A data processing method based on multimodal contrast fusion and graph neural networks, characterized in that, The specific steps include: S1. Collect samples from a multimodal biomedical dataset that includes mRNA expression data, DNA methylation data, and miRNA expression data; S2. Each modality feature vector is mapped to a latent space of a unified dimension through an independent feedforward neural network encoder to obtain the embedding matrices of the three modalities. An NT-Xent-based contrastive learning mechanism is used to perform unsupervised contrastive alignment on the embedding matrices of different modalities. By maximizing the similarity between homologous modal pairs and minimizing the similarity between heterologous modal representations, the aligned embedding matrices of the three modalities are generated. S3. Stack the aligned modal embeddings into a sequence, input them into a multi-layer Transformer encoder, model the high-order interactions between modalities through a multi-head self-attention mechanism, and obtain a unified fusion representation through modal dimension average pooling; normalize the fusion representation, calculate the cosine similarity matrix between samples, retain several similar neighbors of each sample to generate a sparse adjacency matrix, and construct a sample similarity graph structure. S4. Based on the graph structure and fusion representation, high-order adjacency features of nodes are extracted through a three-layer graph convolutional network, and sample classification is completed by combining two fully connected layers. The probability of a sample under different classification results is output, where the classification loss function is a joint optimization of cross-entropy loss and contrast loss.
2. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, The mRNA expression data dimension is N×G1, where G is the number of detected genes and N is the number of samples. Each element represents the standardized expression level of a gene in the corresponding sample. The DNA methylation data dimension is N×G2, where G2 is the total number of CpG sites detected by the sequencing platform and N is the number of samples. Each element represents the methylation level of a sample at a specific site using a β value. The miRNA expression data dimension is N×G3, where G3 is the number of detected miRNAs and N is the number of samples. Each element is a standardized expression value, representing the expression level of the corresponding miRNA in the sample.
3. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, In S2, the feedforward neural network encoder includes a normalization layer and a ReLU activation function, which maps the original features of each modality from the original dimension to a unified dimension.
4. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, In S2, based on the NT-Xent contrastive learning mechanism, the embedding matrices of any two modalities from the three modalities of mRNA, DNA, and miRNA of any sample are selected to form a modality pair. After normalizing the embedding matrices of the modality pair, the cosine similarity matrix S is calculated, and the expression is: in, and Let be the embedding matrix of mode a of sample i and mode b of sample j, and τ be a parameter that controls the sharpness of the similarity distribution.
5. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, The NT-Xent contrastive learning mechanism uses a multimodal contrastive loss function. To bring the representations of the same sample closer together and increase the distance between the modal representations of different samples, the expression is: Where M is the number of modes, with a value of 3. Let be the cross-entropy loss term between modes a and b.
6. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, In step S3, the aligned modes are embedded. Stacked as a sequence The input is a Transformer encoder, and the output is obtained by average pooling to obtain the fused representation z. i : Where M is the number of dimensions, Stacked as a sequence.
7. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, In step S3, for each row of the adjacency matrix, its Top-K similar neighbors are retained, and the remaining elements are set to zero, thereby generating a sparse adjacency matrix, expressed as: Among them, A ij For elements in the sparse adjacency matrix, For the set of indexes of the K samples that rank K most similar to sample i, the diagonal elements A ii Set to zero, that is 8. The data processing method based on multimodal contrast fusion and graph neural networks according to claim 1, characterized in that, The graph convolutional network takes a fused representation and a sparse adjacency matrix as input. The graph convolution operation involves extracting high-order adjacency features in one pass through a three-layer graph convolutional network, mapping them to the class space through two fully connected layers, and generating predicted probabilities after Softmax activation. The loss function of the graph convolutional network aims to minimize the cross-entropy loss. Where N is the number of samples and C is the number of categories. Let be the predicted probability of sample i belonging to category c. One-hot encoding for the real label.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.