Gene multi-omics data fusion method in modal deletion scene
By processing genomic multi-omics data using the modality missing rate algorithm and hypergraph neural network, the modality missing problem was solved, accurate data fusion and efficient prediction were achieved, and the stability and accuracy of the analysis were improved.
Patent Information
- Application Number
- CN202510832558.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
In multi-omics data analysis, the problem of missing modalities makes data integration and analysis difficult, and existing methods cannot effectively judge the similarity between data.
The modality missing rate algorithm is used to process genomic multi-omics data, and the common latent representation is obtained through non-negative matrix decomposition. Hyperedge groups and hypergraph adjacency matrices are constructed. Hypergraph neural networks are used for information propagation. The prototype network and feature encoder are combined to extract features for multimodal data fusion and prediction.
It improves the stability and accuracy of data analysis in modality-missing scenarios, comprehensively captures the local and global characteristics of samples, and improves the accuracy and reliability of analysis results.
Smart Images

Figure CN120705480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method for fusion of multi-omics gene data in a modality-missing scenario. Background Art
[0002] The rapid development of high-throughput biomedical technologies has made the collection of multiple types of omics data more sophisticated. Researchers can obtain genome-wide data on different molecular processes for the same set of samples, thereby providing multi-omics data support for disease research. Although each omics technology can only capture part of the biological complexity, integrating multiple omics data can reveal the underlying biological processes more comprehensively. Studies on human diseases have shown that the integration of multi-omics data significantly improves the accuracy of predicting patient clinical outcomes compared to single-type omics data.
[0003] However, in real-world clinical scenarios, missing modalities are a common problem in multi-omics data analysis, posing significant challenges to data integration and analysis. Currently, common approaches to addressing missing modalities rely on auxiliary information from similar data, compensating for data loss by inferring feature representations of the missing modalities in the latent space. However, determining similarity between data remains a significant challenge.
[0004] Therefore, a method for fusion of multi-omics data in a modality-missing scenario is provided to solve the above problems. Summary of the Invention
[0005] In order to solve the above problems, the present invention provides a method for fusion of genetic multi-omics data in a modality-missing scenario, which solves the problem of modality-missing in multi-omics data analysis through an accurate and comprehensive data similarity determination method.
[0006] To achieve the above objectives, the present invention provides a method for fusion of multi-omics data in a modality-missing scenario, which specifically includes the following steps:
[0007] S1: Using the modal missing rate algorithm, we process the missing data of multi-omics based on the set missing rate to obtain data that meets the missing ratio.
[0008] S2: Preprocess the data obtained in S1 that meets the missing ratio to obtain preprocessed data; preprocessing includes feature screening and normalization;
[0009] S3: Use non-negative matrix factorization algorithm to obtain the common latent representation of the data based on the preprocessed data in S2;
[0010] S4: Construct similarity relations between samples based on the cosine similarity between common latent representations, construct hyperedge groups based on the similarity relations between samples, and connect the hyperedge groups to obtain the hypergraph adjacency matrix;
[0011] S5: Using the prototype network layer and feature encoder, extract the omics data features based on the preprocessed data in S2 as the hypergraph neural network node;
[0012] S6: Based on the hypergraph adjacency matrix and omics data features, the information propagation mechanism of the hypergraph neural network is used to aggregate similar sample information, interpolate missing modalities in the latent space, and fuse existing modal data and auxiliary information to obtain interpolated and enhanced features;
[0013] S7: Input the interpolated and enhanced features in S6 into the multimodal learning module to perform multimodal data fusion and output the prediction results.
[0014] Preferably, the objective function of the non-negative matrix factorization algorithm in S3 is expressed as:
[0015]
[0016] Where M represents the number of modes, X m Represents the preprocessed data of the mth mode, U m represents the basis matrix of the latent subspace of the mth modality, P represents the public latent representation of the sample, λ and β are the trade-off parameters, represents the i-th sample data of the m-th mode after preprocessing, represents the jth sample data of the mth mode after preprocessing, P i represents the public potential representation of the i-th sample, P j represents the public latent representation of the j-th sample.
[0017] Preferably, the hyperedge group in S4 is constructed by the k-hop neighbor method, which finds the vertices related to the central vertex according to the reachable positions in the graph structure; s The domain of a vertex v is defined as:
[0018]
[0019] Among them, the value range of k is [2,n v ];n v Representation graph G s The number of vertices in ; the hyperedge group E with k-hop hop The formula is defined as:
[0020]
[0021] Where G = (V, E) represents the graph structure of the similarity relationship between samples; A represents the graph G s The adjacency matrix of .
[0022] Preferably, S5 specifically includes the following steps:
[0023] S51: Input the preprocessed data into the prototype network layer and extract the prototype activation state as prototype level information;
[0024] S52: inputting the preprocessed data into a feature encoder to extract feature level information;
[0025] S53: Fusion of prototype-level information and feature-level information to generate omics data feature Z m , as a hypergraph node;
[0026] S54: The objective function designed to construct the prototype network layer is:
[0027]
[0028] Where N represents the total number of samples; M represents the number of modalities; represents the i-th sample data of the m-th mode after preprocessing; Indicates that the mth mode belongs to label y i All prototypes of ;p j express The j-th prototype in ; s represents the data Partial gene sequence in L cluster To minimize the clustering loss, l separation To minimize separation cost losses.
[0029] Preferably, S6 specifically includes the following steps:
[0030] S61: Hypergraph adjacency matrix H and omics data features Z m Input the hypergraph neural network, and after the spectral convolution operation in the hypergraph neural network, the features of aggregating similar sample information are obtained.
[0031]
[0032] Among them, D v and D e represents the degree matrix of vertices and hyperedges in the hypergraph; W represents the diagonal matrix of hyperedge weights; the wave filter Θ is the parameter that needs to be learned during training; σ represents the nonlinear activation function, HGNN is a hypergraph neural network, The feature of the i-th aggregated similar sample information;
[0033] S62: Using attention fusion from and Z m Extract the appropriate amount of information from:
[0034]
[0035] Among them, α and β represent the importance of self information and similar sample information respectively. β=1-αAdd constraint α+β=1; W o and W s Represent the weight matrix respectively; represents the importance of the i-th self-information of mode m, Indicates the importance of the i-th similar sample information of modality m.
[0036] Preferably, S7 specifically includes the following steps:
[0037] S71: Representation Z of each mode m Connect to form input z 0 :
[0038] z 0 =[Z 1 ; Z 2 ;…,Z m ];
[0039] S72: z 0 Input l multimodal fusion modules, expressed as:
[0040]
[0041] Where l=1,…,L represents the number of stacked multimodal fusion modules; MHSA represents the multi-head self-attention mechanism; FFN represents the feedforward network; LayerNorm represents the layer normalization operation;
[0042] S73: The construction formula of the predictor is expressed as:
[0043]
[0044] Among them, W final Represents the weight matrix; the Sigmoid function is used to convert the output into probability, Indicates that z 0 The final result obtained by inputting L multimodal fusion modules;
[0045] S74: Use cross entropy loss as the prediction loss function, expressed as:
[0046]
[0047] Among them, B represents the batch size, y i ∈{0,1} |C| represents the true label, Represents the true label y i The transpose of .
[0048] Preferably, the non-negative matrix factorization algorithm uses an alternating optimization strategy to m Iteratively update P, which specifically includes the following steps:
[0049] S31: Fixed matrix P, minimize the objective function of the non-negative matrix factorization algorithm to update the matrix U m ;
[0050] S32: Fixed matrix U m ,Minimize the objective function of the non-negative matrix factorization algorithm to update the matrix P;
[0051] S33: Repeat S31-S32 until a preset convergence condition is met.
[0052] Preferably, the feature encoder in S52 uses the maximum category probability and the true category probability to quantify the classification confidence of different modalities, specifically including the following steps:
[0053] S521: Build a classifier f for each modality m m : Classifier f m The observed value x m Convert to predictive distribution The classifier is trained using the maximum likelihood estimation framework to minimize the Kullback-Leibler divergence between the predicted distribution and the true distribution:
[0054]
[0055] Among them, y k represents the kth element; Denotes the classifier f m For the predicted probability of the kth category, the maximum category probability The inferred maximum class probability is regarded as the confidence of the classifier in the prediction result;
[0056] S522: For a given prediction distribution And the corresponding label y, the true category probability TCP m The formula is:
[0057]
[0058] Where (·) represents the inner product;
[0059] S523: Construct a confidence neural network g for each modality m m :x m →TCP m To approximate TCP m , use l2 loss to train the confidence network gm :
[0060]
[0061] in,
[0062] Preferably, the total loss function is expressed as:
[0063]
[0064] Here, λ1 and λ2 represent the hyperparameters that control the loss.
[0065] Preferably, S2 specifically includes:
[0066] The data obtained by S1 that meet the missing ratio are used to screen out key expression gene sequences through feature selection methods, and the key expression gene sequences are normalized to obtain preprocessed data.
[0067] Therefore, the present invention adopts the above-mentioned gene multi-omics data fusion method in the modality missing scenario, which can effectively deal with the modality missing problem. Through the data of similar samples, the characteristic representation of the missing modality is inferred, thereby improving the stability of the model; further, it provides an accurate method for determining the similarity between samples, which can comprehensively capture the local and global characteristics of the samples and improve the accuracy and reliability of the analysis results.
[0068] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 Schematic diagram of a process for fusion of multi-omics data in a modality-missing scenario in the present invention;
[0070] Figure 2 Schematic diagram of the network structure in an embodiment of the present invention. DETAILED DESCRIPTION
[0071] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0072] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0073] The words “include” or “comprising” and similar words used in the present invention mean that the elements before the word include the elements listed after the word, and do not exclude the possibility of also including other elements. The orientation or position relationship indicated by the terms “inside”, “outside”, “upper”, “lower”, etc. is based on the orientation or position relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation of the present invention. When the absolute position of the described object changes, the relative position relationship may also change accordingly. In the present invention, unless otherwise clearly stipulated and limited, the terms such as “attachment” should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral whole; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0074] Example
[0075] A method for fusion of multi-omics data in a modality-missing scenario, such as Figure 1-2 As shown, the specific steps include:
[0076] S1: Using the modal missing rate algorithm, we process the missing data of multi-omics based on the set missing rate to obtain data that meets the missing ratio.
[0077] Specifically, the multi-omics data are missing, and the modal missing rate algorithm is used to introduce modal missingness into the multi-omics data according to the set missing rate to generate data that meets the missing ratio.
[0078] S2: Preprocess the data obtained in S1 that meets the missing ratio to obtain preprocessed data; preprocessing includes feature screening and normalization;
[0079] Specifically, the data obtained in S1 that meet the missing ratio are screened by a feature selection method to screen out key expression gene sequences, and the key expression gene sequences are normalized to obtain preprocessed data.
[0080] S3: Use non-negative matrix factorization algorithm to obtain the common latent representation of the data based on the preprocessed data in S2;
[0081] The objective function of the non-negative matrix factorization algorithm is expressed as:
[0082]
[0083] Where M represents the number of modes, X mRepresents the preprocessed data of the mth mode, U m represents the basis matrix of the latent subspace of the mth modality, P represents the public latent representation of the sample, λ and β are the trade-off parameters, represents the i-th sample data of the m-th mode after preprocessing, represents the jth sample data of the mth mode after preprocessing, P i represents the public potential representation of the i-th sample, P j represents the public latent representation of the j-th sample.
[0084] The non-negative matrix factorization algorithm uses an alternating optimization strategy to m and P are iteratively updated to reduce the computational complexity and improve the stability of the decomposition results. The alternating optimization strategy includes the following steps:
[0085] S31: Fixed matrix P, minimize the objective function of the non-negative matrix factorization algorithm to update the matrix U m ;
[0086] S32: Fixed matrix U m ,Minimize the objective function of the non-negative matrix factorization algorithm to update the matrix P;
[0087] S33: Repeat S31-S32 until a preset convergence condition is met.
[0088] S4: Construct similarity relations between samples based on the cosine similarity between common latent representations, construct hyperedge groups based on the similarity relations between samples, and connect the hyperedge groups to obtain the hypergraph adjacency matrix;
[0089] Hyperedge groups are obtained by k-hop neighbor method (E hop )Build, E hop Aims to find the vertices related to the central vertex based on the reachable positions in the graph structure; Graph G s The domain of a vertex v is defined as:
[0090]
[0091] Among them, the value range of k is [2,n v ];n v Representation graph G s The number of vertices in ; the hyperedge group E with k-hop hop The formula can be defined as:
[0092]
[0093] Where G = (V, E) represents the graph structure of the similarity relationship between samples; A represents the graph G s The adjacency matrix of .
[0094] S5: Using the prototype network layer and feature encoder, extract the omics data features based on the preprocessed data in S2 as the hypergraph neural network node;
[0095] S5 specifically includes the following steps:
[0096] S51: Input the preprocessed data into the prototype network layer and extract the prototype activation state as prototype level information;
[0097] S52: simultaneously inputting the preprocessed data into a feature encoder to extract feature level information;
[0098] The feature encoder uses the maximum category probability and the true category probability to quantify the classification confidence of different modalities. Specifically, it includes the following steps:
[0099] S521: Build a classifier f for each modality m m : Classifier f m The observed value x m Convert to predictive distribution The classifier is trained using the maximum likelihood estimation framework to minimize the Kullback-Leibler divergence between the predicted distribution and the true distribution:
[0100]
[0101] Among them, y k represents the kth element; Denotes the classifier f m For the predicted probability of the kth category, the maximum category probability The inferred maximum class probability is regarded as the confidence of the classifier in the prediction result;
[0102] S522: For a given prediction distribution And the corresponding label y, the true category probability TCP m The formula is:
[0103]
[0104] Where (·) represents the inner product;
[0105] S523: Since TCP requires label information and cannot be used directly in the testing phase, a confidence neural network g is constructed for each modality m m :x m →TCP m To approximate TCP m , use l2 loss to train the confidence network g m :
[0106]
[0107] in,
[0108] S53: Fusion of prototype-level information and feature-level information to generate omics data feature Z m , as a hypergraph node;
[0109] S54: The objective function designed to construct the prototype network layer is:
[0110]
[0111] Where N represents the total number of samples; M represents the number of modalities; represents the i-th sample data of the m-th mode after preprocessing; Indicates that the mth mode belongs to label y i All prototypes of ;p j express The j-th prototype in ; s represents the data Part of the gene sequence in; by minimizing the clustering loss L cluster , making each gene sequence closer to at least one prototype of the category to which it belongs; by minimizing the separation cost L separation , moving the gene sequence away from prototypes that do not belong to its category.
[0112] S6: Based on the hypergraph adjacency matrix and omics data features, the information propagation mechanism of the hypergraph neural network (HGNN) is used to aggregate similar sample information, interpolate missing modal data in the latent space, and fuse existing modal data and auxiliary information to obtain interpolated and enhanced features;
[0113] S6 specifically includes the following steps:
[0114] S61: Hypergraph adjacency matrix H and omics data features Z m Input the hypergraph neural network, and after the spectral convolution operation in the hypergraph neural network, the features of aggregating similar sample information are obtained.
[0115]
[0116] Among them, D v and D e represents the degree matrix of vertices and hyperedges in the hypergraph; W represents the diagonal matrix of hyperedge weights; the wave filter Θ is the parameter that needs to be learned during training; σ represents the nonlinear activation function, HGNN is a hypergraph neural network, The feature of the i-th aggregated similar sample information;
[0117] S62: Using attention fusion from and Z m Extract the appropriate amount of information from:
[0118]
[0119] α m =Sigmoid(Z m W o ),
[0120] Among them, α and β represent the importance of self information and similar sample information respectively. β=1-αAdd constraint α+β=1; W o and W s Represent the weight matrix respectively; represents the importance of the i-th self-information of mode m, Indicates the importance of the i-th similar sample information of modality m.
[0121] S7: Input the interpolated and enhanced features in S6 into the multimodal learning module to perform multimodal data fusion and output the prediction results.
[0122] S7 specifically includes the following steps:
[0123] S71: Representation Z of each mode m Connect to form input z 0 :
[0124] z 0 =[Z 1 ; Z 2 ;…,Z m ];
[0125] S72: z 0 Input l multimodal fusion modules, expressed as:
[0126]
[0127] Where l=1,…,L represents the number of stacked multimodal fusion modules; MHSA represents the multi-head self-attention mechanism; FFN represents the feedforward network; LayerNorm represents the layer normalization operation;
[0128] S73: The construction formula of the predictor is expressed as:
[0129]
[0130] Among them, W final Represents the weight matrix; the Sigmoid function is used to convert the output into probability, Indicates that z 0The final result obtained by inputting L multimodal fusion modules;
[0131] S74: Use cross entropy loss as the prediction loss function to train the model. The specific loss function is:
[0132]
[0133] Among them, B represents the batch size, y i ∈{0,1} |C| represents the true label, Represents the true label y i The transpose of .
[0134] The loss function of the entire model is:
[0135]
[0136] Here, λ1 and λ2 represent the hyperparameters that control the loss.
[0137] Example 1
[0138] In order to verify the effectiveness of the gene multi-omics data fusion method provided in this embodiment in a modality missing scenario, a comparative experiment was conducted on various existing models for modality missing scenarios, such as the DMRNet model, the ShaSpec model, and the m3care model. The ROSMAP dataset from the AMP-AD knowledge portal was used. The ROSMAP dataset contains two groups of data: Alzheimer's disease (AD) and normal control (NC). The statistical results of various indicators of the DMRNet model and the model of this embodiment are shown in Table 1 below, where the percentage represents the missing rate, such as 20% means 20% of the modality is missing.
[0139] Table 1 Statistical results of various indicators of the model in this embodiment
[0140] AD ACC / AUC / F1 0% 0.7147(±0.0317) / 0.7827(±0.0227) / 0.7558(±0.0340) 0%_Our 0.8120(±0.0227) / 0.8795(±0.0261) / 0.8369(±0.0267) 10% 0.7160(±0.0373) / 0.7997(±0.0239) / 0.7479(±0.0370) 10%_Our 0.7827(±0.0224) / 0.8384(±0.0158) / 0.8099(±0.0334) 20% 0.6907(±0.0484) / 0.7574(±0.0399) / 0.7381(±0.0391) 20%_Our 0.7320(±0.0283) / 0.7855(±0.0306) / 0.7730(±0.0286) 30% 0.6880(±0.0299) / 0.7516(±0.0531) / 0.7351(±0.0300) 30%_Our 0.7320(±0.0340) / 0.7683(±0.0289) / 0.7702(±0.0362)
[0141] The statistical results of various indicators of the ShaSpec model and the model of this embodiment verified on the ROSMAP dataset are shown in Table 2 below:
[0142] Table 2 Statistical results of various indicators of the ShaSpec model and the model of this embodiment verified on the ROSMAP dataset
[0143]
[0144]
[0145] The statistical results of various indicators of the m3care model and the model of this embodiment verified in the ROSMAP dataset are shown in Table 3 below:
[0146] Table 3 Statistical results of various indicators of the m3care model and the model of this embodiment verified on the ROSMAP dataset
[0147] AD ACC / AUC / F1 0% 0.7040(±0.0309) / 0.7823(±0.0381) / 0.7039(±0.0607) 0%_Our 0.8120(±0.0227) / 0.8795(±0.0261) / 0.8369(±0.0267) 10% 0.7453(±0.0448) / 0.8210(±0.0456) / 0.7706(±0.0430) 10%_Our 0.7827(±0.0224) / 0.8384(±0.0158) / 0.8099(±0.0334) 20% 0.6960(±0.0980) / 0.7383(±0.0976) / 0.6980(±0.2341) 20%_Our 0.7320(±0.0283) / 0.7855(±0.0306) / 0.7730(±0.0286) 30% 0.6813(±0.1118) / 0.7403(±0.0583) / 0.6912(±0.2329) 30%_Our 0.7320(±0.0340) / 0.7683(±0.0289) / 0.7702(±0.0362)
[0148] It can be seen from Tables 1-3 that the method provided in this embodiment is superior to existing models in terms of accuracy (ACC), area under the curve (AUC) and F1 score (F1Score) under various modality missing conditions. For example, in a scenario with a modality missing rate of 30%, the accuracy of this embodiment reaches 0.7320, which is 0.044 higher than the DMRNet model, 0.0307 higher than the ShaSpec model, and 0.0507 higher than the m3care model. The F1 score of this embodiment is 0.0351, 0.0268 and 0.079 higher than the DMRNet model, ShaSpec model and m3care model, respectively. In general, the method proposed in this embodiment is superior to these existing methods and can effectively complete missing modalities and efficiently predict targets in modality missing scenarios.
[0149] Therefore, the present invention adopts the aforementioned multi-omics data fusion method for the missing modality scenario. Using data from similar samples, it infers the characteristic representation of the missing modality, thereby improving the stability of the model. Furthermore, it provides an accurate method for determining similarity between samples, which can comprehensively capture the local and global characteristics of the samples, improving the accuracy and reliability of the analysis results. By combining deep learning technology with knowledge from the field of multi-omics, it achieves the goals of data noise reduction, missing modality completion, and efficient prediction.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for fusion of multi-omics data in a modality-missing scenario, characterized by: The specific steps include: S1: Using the modal missing rate algorithm, we process the missing data of multi-omics based on the set missing rate to obtain data that meets the missing ratio. S2: Preprocess the data that meets the missing ratio obtained in S1 to obtain preprocessed data; Preprocessing Including feature screening and normalization processing; S3: Use non-negative matrix factorization algorithm to obtain the common latent representation of the data based on the preprocessed data in S2; S4: Construct similarity relations between samples based on the cosine similarity between common latent representations, construct hyperedge groups based on the similarity relations between samples, and connect the hyperedge groups to obtain the hypergraph adjacency matrix; S5: Using the prototype network layer and feature encoder, extract the omics data features based on the preprocessed data in S2 as the hypergraph neural network node; S6: Based on the hypergraph adjacency matrix and omics data features, the information propagation mechanism of the hypergraph neural network is used to aggregate similar sample information, interpolate missing modalities in the latent space, and fuse existing modal data and auxiliary information to obtain interpolated and enhanced features; S7: Input the interpolated and enhanced features in S6 into the multimodal learning module to perform multimodal data fusion and output the prediction results.
2. The method for fusion of multi-omics data in a modality-missing scenario according to claim 1, characterized in that: The objective function of the non-negative matrix factorization algorithm in S3 is expressed as: Where M represents the number of modes, x m Represents the preprocessed data of the mth mode, U m represents the basis matrix of the latent subspace of the mth modality, P represents the public latent representation of the sample, λ and β are the trade-off parameters, represents the i-th sample data of the m-th mode after preprocessing, represents the jth sample data of the mth mode after preprocessing, P i represents the public potential representation of the i-th sample, P j represents the public latent representation of the j-th sample.
3. The method for fusion of multi-omics data in a modality-missing scenario according to claim 1, characterized in that: The hyperedge group in S4 is constructed by the k-hop neighbor method, which finds the vertices related to the central vertex according to the reachable positions in the graph structure; s The domain of a vertex v is defined as: Among them, the value range of k is [2,n v ];n v Representation graph G s The number of vertices in ; the hyperedge group E with k-hop hop The formula is defined as: Where G = (V, E) represents the graph structure of the similarity relationship between samples; A represents the graph G s The adjacency matrix of .
4. The method for fusion of multi-omics data in a modality-missing scenario according to claim 3, characterized in that: S5 specifically includes the following steps: S51: Input the preprocessed data into the prototype network layer and extract the prototype activation state as prototype level information; S52: inputting the preprocessed data into a feature encoder to extract feature level information; S53: Fusion of prototype-level information and feature-level information to generate omics data feature Z m , as a hypergraph node; S54: The objective function designed to construct the prototype network layer is: Where N represents the total number of samples; M represents the number of modalities; represents the i-th sample data of the m-th mode after preprocessing; Indicates that the mth mode belongs to label y i All prototypes of ;p j express The j-th prototype in ; s represents the data Partial gene sequence in L cluster To minimize the clustering loss, L separation To minimize separation cost losses.
5. The method for fusion of multi-omics data in a modality-missing scenario according to claim 4, characterized in that: S6 specifically includes the following steps: S61: Hypergraph adjacency matrix H and omics data features Z m Input the hypergraph neural network, and after the spectral convolution operation in the hypergraph neural network, the features of aggregating similar sample information are obtained. Among them, D v and D e represents the degree matrix of vertices and hyperedges in the hypergraph; W represents the diagonal matrix of hyperedge weights; the wave filter Θ is the parameter to be learned during training; σ represents the nonlinear activation function, HGNN is a hypergraph neural network, The feature of the i-th aggregated similar sample information; S62: Using attention fusion from and Z m Extract the appropriate amount of information from: α m =Sigmoid(Z m W o ), Among them, α and β represent the importance of self information and similar sample information respectively. β=1-αAdd constraint α+β=1; W o and W s Represent the weight matrix respectively; represents the importance of the i-th self-information of mode m, Indicates the importance of the i-th similar sample information of modality m.
6. The method for fusion of multi-omics data in a modality-missing scenario according to claim 2, characterized in that: S7 specifically includes the following steps: S71: Representation Z of each mode m Connect to form input z 0 : With 0 =[Z 1 ;WITH 2 ;…,WITH m ]; S72: z 0 Input l multimodal fusion modules, expressed as: Where l=1,…,L represents the number of stacked multimodal fusion modules; MHSA represents the multi-head self-attention mechanism; FFN represents the feedforward network; LayerNorm represents the layer normalization operation; S73: The construction formula of the predictor is expressed as: Among them, W final Represents the weight matrix; the Sigmoid function is used to convert the output into probability, Indicates that z 0 The final result obtained by inputting L multimodal fusion modules; S74: Use cross entropy loss as the prediction loss function, expressed as: Among them, B represents the batch size, y i ∈{0,1} |C| represents the true label, Represents the true label y i The transpose of .
7. The method for fusion of multi-omics data in a modality-missing scenario according to claim 6, characterized in that: The non-negative matrix factorization algorithm uses an alternating optimization strategy to m Iteratively update P, which specifically includes the following steps: S31: Fixed matrix P, minimize the objective function of the non-negative matrix factorization algorithm to update the matrix U m ; S32: Fixed matrix U m ,Minimize the objective function of the non-negative matrix factorization algorithm to update the matrix P; S33: Repeat S31-S32 until a preset convergence condition is met.
8. The method for fusion of multi-omics data in a modality-missing scenario according to claim 1, characterized in that: The feature encoder in S52 uses the maximum category probability and the true category probability to quantify the classification confidence of different modalities. Specifically, it includes the following steps: S521: Build a classifier for each modality m Classifier f m The observed value x m Convert to predictive distribution The classifier is trained using the maximum likelihood estimation framework to minimize the Kullback-Leibler divergence between the predicted distribution and the true distribution: Among them, y k represents the kth element; Represents the classifier f m For the predicted probability of the kth category, the maximum category probability The inferred maximum class probability is regarded as the confidence of the classifier in the prediction result; S522: For a given prediction distribution And the corresponding label y, the true category probability TCP m The formula is: Where (·) represents the inner product; S523: Construct a confidence neural network g for each modality m m :x m →TCP m To approximate TCP m ,use Loss training confidence network g m : in, 9. The method for fusion of multi-omics data in a modality-missing scenario according to claim 8, characterized in that: The total loss function is expressed as: Here, λ1 and λ2 represent the hyperparameters that control the loss.
10. The method for fusion of multi-omics data in a modality-missing scenario according to claim 1, characterized in that: S2 specifically includes: The data obtained by S1 that meet the missing ratio are used to screen out key expression gene sequences through feature selection methods, and the key expression gene sequences are normalized to obtain preprocessed data.