Incomplete multi-omics cancer subtype data clustering method

By combining variational autoencoders and multi-omics diffusion models with a two-level contrastive learning framework, the problem of cancer subtype identification in high-dimensional, heterogeneous, and incomplete multi-omics data was solved, achieving higher clustering accuracy and robustness.

CN121034428AActive Publication Date: 2025-11-28FUJIAN NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511558495.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-11-28
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively process high-dimensional, heterogeneous, and incomplete multi-omics cancer subtype data, resulting in insufficient accuracy and robustness in cancer subtype identification.

Method used

A variational autoencoder framework is used to extract shared and specific features. Combined with a multi-omics diffusion model and a two-level contrastive learning framework, a comprehensive representation is generated. Clustering is optimized by the K-means algorithm, and multi-omics information is integrated to improve clustering performance.

Benefits of technology

It significantly improves the accuracy and robustness of cancer subtype identification in incomplete multi-omics data, better handles complex data scenarios, and enhances the understanding and analysis of biomolecular exchanges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034428A_ABST
    Figure CN121034428A_ABST
Patent Text Reader

Abstract

The invention provides an incomplete multi-omics cancer subtype data clustering method, and belongs to the field of machine learning, and the method comprises the steps: generating a shared potential representation and an omics specific potential representation for each sample through a variational auto-encoder, capturing omics specific biological features, and carrying out the regularization of representation distribution through reconstruction loss and KL divergence; a diffusion model framework is adopted to fuse shared potential representations of multiple omics and specific potential representations of the omics, comprehensive representations are generated, incomplete data are effectively processed, and comprehensive extraction of information of the multiple omics is enhanced; the comprehensive representation is optimized through two-stage contrast learning, and the consistency of the representation is enhanced; according to the method, clustering input is jointly optimized through joint optimization of clustering loss, reconstruction loss and semantic comparison loss to generate high-quality potential representation, and based on the clustering input after joint optimization, an improved K-means clustering algorithm is used for analyzing multi-omics data to obtain a final cancer subtype clustering result. According to the invention, the accuracy and robustness of cancer subtype recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of machine learning, and particularly relates to an incomplete multi-omics cancer subtype data clustering method. BACKGROUND

[0002] Cancer is a highly heterogeneous disease, and the response of patients to clinical treatment varies significantly. Studies have shown that molecular characteristics have a profound impact on disease prognosis in tumors that are difficult to distinguish histopathologically. Therefore, to achieve personalized treatment, cancer subtype classification has become a research hotspot, and patients are accurately stratified through molecular or phenotypic characteristics.

[0003] With the advancement of biotechnology, it is increasingly convenient to obtain multi-omics data. Early methods rely on single-omics data and apply traditional clustering algorithms to predict cancer subtypes. However, single-omics only partially represents molecular characteristics, and often faces the problem of incomplete data, such as missing mRNA, miRNA, DNA methylation, and other omics in some samples, limiting the comprehensiveness and accuracy of analysis. Integrating multi-omics data can provide a more comprehensive perspective, reveal the potential patterns of cancer subtypes, and deepen the understanding of the complex interactions between biological molecules at multiple levels, especially in the context of incomplete data. Multi-omics integration is crucial in subtype identification, but traditional clustering methods struggle to cope with the complexity of high-dimensional, heterogeneous, and incomplete data, requiring innovative methods to improve the accuracy and robustness of clustering. Variational autoencoder effectively captures omics-specific features by generating probabilistic latent representations, and uses a diffusion model framework to integrate the latent distribution of incomplete omics, generating a robust comprehensive representation, combining contrastive learning to optimize multi-omics consistency and sample discriminativeness, providing an efficient solution for cancer subtype clustering of incomplete multi-omics data. SUMMARY

[0004] The purpose of the present application is to provide an incomplete multi-omics cancer subtype data clustering method that can effectively handle incomplete multi-omics data and improve the accuracy and robustness of cancer subtype identification.

[0005] To achieve the above-mentioned purpose, the technical solution of the present application is as follows: an incomplete multi-omics cancer subtype data clustering method, specifically comprising:

[0006] S1, omics encoding based on a variational autoencoder framework: for the non-missing omics data of each sample in the incomplete multi-omics cancer subtype data, a shared encoder is used to extract shared features across omics and generate shared latent representations, and omics-specific encoders are used to extract specific features of each omics and generate omics-specific latent representations;

[0007] S2. Information aggregation based on multi-omics diffusion model framework: The generated shared latent representation and omics-specific latent representation are used as the initial latent representation. During the forward diffusion process, noise is added to capture cross-omics shared biological features and omics-specific variations of the data. During the reverse denoising process, the hidden links between omics are integrated through a bidirectional cross-attention mechanism to generate a joint distribution. Finally, starting from the noise distribution, the data is repeatedly sampled and a comprehensive representation is generated through the reverse denoising process.

[0008] S3. Using the generated comprehensive representation as input, information aggregation and optimization are performed through a two-level contrastive learning framework: the two-level contrastive loss is composed of high-level contrastive loss and semantic-level contrastive loss, and two-level contrastive learning is performed to generate optimized representation;

[0009] S4. Joint Optimization and Clustering: Based on the extension of the K-means algorithm, cancer subtype data are clustered by using the optimized representation as the initial clustering input through an iterative optimization algorithm. At the same time, the clustering input is jointly optimized by clustering loss, reconstruction loss and semantic contrast loss to improve clustering performance and finally output the clustering results.

[0010] Preferably, the step of extracting shared features across omics and generating shared latent representations through a shared encoder, and extracting specific features of each omics and generating omics-specific latent representations through an omics-specific encoder, is as follows:

[0011] S1.1, using non-missing omics data from each sample in incomplete multi-omics cancer subtype data. As input; where: It is the first The first sample omics data, and To use the data after min-max normalization, Indicates sample The non-missing omics set, , For the set of real numbers, It is the feature dimension;

[0012] S1.2. Map the non-missing omics of each sample to the latent space through a shared encoder and an omics-specific encoder to generate Gaussian distribution parameters:

[0013]

[0014]

[0015] in: For the first Omics-specific encoders; , They represent data respectively mean and standard deviation of the approximate posterior distribution generated as input to the omics-specific encoder ; as the shared encoder , represent the mean and standard deviation of the approximate posterior distribution output as input to the shared encoder ; as the shared encoder , are parameters of the omics-specific encoder and the shared encoder ;

[0016] The approximate posterior distribution is represented as:

[0017]

[0018] wherein represents the approximate posterior distribution defined by the output of the omics-specific encoder , represents the approximate posterior distribution defined by the output of the shared encoder ; represents a Gaussian distribution; represents a diagonal matrix; and represent parameters obtained by averaging the shared Gaussian distribution parameters over all non-missing omics of the sample ;

[0019] ,

[0020] , are the omics-specific latent representation and the shared latent representation of the th omics of the sample

[0021]

[0022] wherein represents an element-wise multiplication, is a random vector following a standard normal distribution , and I is an identity matrix.

[0023] Preferably, the lower bound of the evidence is optimized by variational inference:

[0024]

[0025] wherein: represents the th a likelihood of the data given the omics-specific latent representation and the shared latent representation, is a KL divergence weight, denotes a KL divergence, and denote prior distributions set for capturing the omics-specific latent representation distribution and the shared latent representation distribution, respectively, denotes a sample The first omcs-specific reconstruction loss:

[0026] reconstructing the data from the omics-specific latent representation and the shared latent representation by the omics-specific decoder

[0027]

[0028] wherein is the omics specific decoder, is a parameter of the decoder ;

[0029] The reconstruction distribution defined by the decoder output is denoted as:

[0030]

[0031] wherein is a fixed hyperparameter;

[0032] The loss function is denoted as:

[0033]

[0034] wherein: denotes an expectation under the approximate posterior distribution denotes an expectation under the approximate posterior distribution ; denotes the omics-specific log-likelihood, denotes the shared log-likelihood, the log-likelihood terms and and the reconstruction distribution The optimization objective ELBO of the variational autoencoder interrelates.

[0035] Preferably, the step S2 is implemented in the following manner:

[0036] S2.1, concatenating the shared latent representation and the omics-specific latent representation of each sample generated in the step S1 to obtain an initial latent representation ​​​and the initial latent representation as input data has a dimension of , has a dimension of , has a dimension of for each sample shared latent representation has a dimension of for each sample omics-specific latent representation

[0037] S2.2, the forward diffusion process step-by-step adds Gaussian noise:

[0038] , ;

[0039] wherein: are the latent representations at time step t in the forward diffusion process or the backward denoising process, respectively, the latent representation is the initial latent representation in the forward diffusion process, and the latent representation is the predicted initial latent representation in the backward denoising process, is the total number of time steps in the forward diffusion process or the backward denoising process; is the probability distribution of the latent representation given the latent representation at time step t in the forward diffusion process, obeys a Gaussian distribution; denotes the noise scheduling coefficient at time step t in the forward diffusion process;

[0040] S2.3, the backward denoising process uses a U-Net architecture and adopts a bidirectional cross-attention mechanism to fuse features in the middle layer, learn the denoising distribution, and optimize the parameters by minimizing the variational lower bound:

[0041]

[0042] wherein: is the probability distribution of the latent representation given the latent representation at time step t in the backward denoising process, learned by the U-Net architecture, and the U-Net architecture parameters are ; is the mean function with the latent representation and time step t as input by the U-Net architecture with parameters ; is the mean function with the latent representation and time step t as input by the U-Net architecture with parameters The covariance function is taken as input at time step t;

[0043] S2.4, From the perspective of noise distribution Initially, through a reverse denoising process Step sampling to generate a comprehensive representation Noise distribution Follows a standard normal distribution The sampling update formula in the reverse denoising process is:

[0044] , ,

[0045] in: Noise scheduling coefficient The cumulative product term, It is the index of time step t. It is the standard deviation parameter of time step t in the reverse denoising process;

[0046] Repeat the sampling process S times and take the average value to obtain the comprehensive representation. And project back Dimension;

[0047]

[0048] in, Let represent the initial latent representation of the s'-th sampling prediction.

[0049] Preferably, by minimizing the variational lower bound Parameter optimization is performed, specifically expressed as:

[0050]

[0051] in: This represents the expectation of the probability distribution of the forward diffusion process. The initial latent representation of the model prediction The probability distribution, through optimization This makes the model predict A data distribution that approximates the true value, i.e., the initial latent representation. ; Indicates the KL divergence. This indicates that the forward diffusion process is at the final time step. Time potential representation probability distribution Represents a predefined prior distribution; This indicates that during the forward diffusion process, given the initial potential representation... and potential representation Under what conditions is the latent representation the probability distribution of.

[0052] Preferably, step S3 is implemented as follows:

[0053] S3.1, data augmentation:

[0054] For the comprehensive representation , the row vector corresponding to each sample , a random disturbance is added to generate an augmented representation :

[0055] ,

[0056] wherein represents a random disturbance, , is a preset standard deviation;

[0057] For the comprehensive representation , the row vector corresponding to each sample , a column index is randomly selected to perform feature permutation and generate an augmented representation , ;

[0058] S3.2, extract high-level features of the augmented representation corresponding to each sample by a multi-layer perception MLP and high-level features of the augmented representation ; and calculate a high-level contrastive loss :

[0059]

[0060] wherein: is a temperature parameter controlling the similarity distribution, is a similarity matrix of high-level positive sample pairs, , is a similarity matrix of high-level negative sample pairs, , represents a cosine similarity:

[0061] ,

[0062] wherein: is the transpose of , and is an L2 norm;

[0063] S3.3, extract semantic representation of the high-level features by a clustering head mapping ​and advanced features semantic representation And calculate semantic-level contrastive loss. :

[0064]

[0065] in This is a semantic-level similarity matrix of positive sample pairs. , This is a semantic-level similarity matrix of negative sample pairs. :

[0066] ,

[0067] in for transpose;

[0068] S3.4, Compare losses based on higher levels Loss compared to semantic level Obtain contrast loss :

[0069]

[0070] in To control the semantic loss weights;

[0071] Calculate contrast loss gradient And perform iterative updates to the parameters:

[0072]

[0073]

[0074] in This is a set of parameters, including the weight matrix of the multilayer perceptron and the weight matrix of the clustering head; The learning rate;

[0075] Complete parameter set After iterative updates, the synthesized representation is adjusted through backpropagation. To optimize representation .

[0076] Preferably, the enhanced representation corresponding to each sample is extracted using a multilayer perceptron (MLP). Advanced features and enhancement representation Advanced features The details are as follows:

[0077]

[0078]

[0079] and Let represent the weight matrices of the first and second layer transformations of the multilayer perceptron, respectively. , , For the first hidden layer dimension, For the second hidden layer dimension, This is the ReLU activation function.

[0080] Preferably, the extraction of high-level features through cluster head mapping is... semantic representation and advanced features semantic representation The details are as follows:

[0081]

[0082]

[0083] in: This is the weight matrix in the clustering head used for feature mapping and clustering decisions. , The number of clusters, This is the Softmax activation function.

[0084] Preferably, step S4 is implemented as follows:

[0085] S4.1. Expand the optimized representation using the K-means algorithm. Select An initial cluster center To optimize the initial cluster distribution, where c'=1,..., c' represents the cluster index;

[0086] S4.2 Iterative optimization:

[0087] S4.2.1, in optimized form As initial clustering input, for optimizing the representation The row vector corresponding to each sample in the data Calculate row vectors With all cluster centers The Euclidean distance between them, and the row vectors Assign to the cluster with the smallest distance: , This represents the cluster label of the i-th sample. It is a row vector to the cluster center square of the distance;

[0088] S4.2.2, updating the cluster center : wherein, represents the sum operation on all row vectors belonging to the cluster , and is the number of samples in the cluster ;

[0089] S4.2.3, jointly optimizing the clustering input through a clustering loss , a reconstruction loss and a semantic contrast loss , and obtaining the clustering result again according to the clustering input after joint optimization; the total loss is represented as:

[0090]

[0091] wherein , represents the weight of the reconstruction loss and the weight of the semantic contrast loss ;

[0092] S4.3 outputs the clustering result after reaching the iteration termination condition, including the final clustering label of each sample and the final cluster center of each cluster .

[0093] Preferably, the clustering loss , the reconstruction loss and the semantic contrast loss are specifically calculated as follows:

[0094] Clustering loss:

[0095]

[0096] Reconstruction loss:

[0097]

[0098] wherein, is the original i-th sample data, is the result of reconstructing the row vector using the decoder of step S1, ;

[0099] Semantic contrast loss:

[0100] ​

[0101] wherein, and are two semantic representations of the row vector obtained by the data augmentation method of step S3.1, the advanced feature extraction method of step S3.2 and the semantic representation extraction method of step S3.3 respectively

[0102] Compared with the prior art, the present application has the following beneficial effects:

[0103] 1. Robust processing of incomplete data: By dynamically fusing incomplete omics cycle distribution through the multi-omics diffusion model framework, the present application can effectively cope with the challenge of omics missing in multi-omics data, and thus significantly improve the robustness and reliability of data compared with traditional supplementary methods.

[0104] 2. Efficient probability feature extraction: The variational autoencoder can capture omics-specific biological features in omics data more accurately than traditional feature extraction methods by generating probabilistic potential representation combined with reconstruction and KL divergence regularization.

[0105] 3. Multi-omics information fusion optimization: The multi-omics diffusion model framework generates efficient comprehensive representation by integrating multi-omics information, and combines contrastive learning optimization to surpass traditional multi-omics methods, which can more effectively extract shared information and is suitable for complex data scenarios.

[0106] 4. Improve clustering accuracy: Based on the comprehensive representation of contrastive learning optimization, combined with K-means algorithm, the present application can accurately identify cancer subtypes, and thus significantly improve the accuracy and biological significance of clustering results based on traditional methods.

[0107] 5. Improve complex data processing capability: By combining variational autoencoder, multi-omics diffusion model framework and double-level contrastive learning, the present application can better handle high-dimensional, complete and incomplete multi-omics data, enhance the understanding and analysis capability of complex molecular exchange, and is suitable for large-scale omics analysis and clinical application. BRIEF DESCRIPTION OF DRAWINGS

[0108] Figure 1 is the flowchart of the present application. DETAILED DESCRIPTION

[0109] The technical solutions of the present application will be specifically described below in conjunction with the accompanying drawings. Figure 1 The technical solutions of the present application will be specifically described below in conjunction with the accompanying drawings.

[0110] The present application proposes an incomplete multi-omics cancer subtype data clustering method, which specifically includes:

[0111] S1, Omics encoding based on variational autoencoder framework: For the non-missing omics data of each sample in the incomplete multi-omics cancer subtype data, shared features across omics are extracted and shared latent representations are generated through a shared encoder, and specific features of each omics are extracted and omics-specific latent representations are generated through an omics-specific encoder;

[0112] The variational autoencoder generates probabilistic latent representations for each omics by combining the shared encoder and the omics-specific encoder, instead of traditional deterministic encoders; Through variational inference and reparameterization techniques, robust latent representations are generated, enhancing the adaptability of the model to noise, incomplete data and data heterogeneity;

[0113] S2, Information aggregation based on multi-omics diffusion model framework: The generated shared latent representations and omics-specific latent representations are used as initial latent representations, and in the forward diffusion process, the shared biological features across omics and the omics-specific variations are captured by adding noise, and in the backward denoising process, the hidden links between omics are integrated through a bidirectional cross-attention mechanism to generate a joint distribution. Finally, starting from the noise distribution, the comprehensive representation is generated by repeatedly sampling and generating through the backward denoising process;

[0114] The bidirectional cross-attention integrates the links between omics, where , are the query matrix, the key matrix, and the value matrix, respectively, , are the initial latent representation and the updated representation after processing by the bidirectional cross-attention mechanism, respectively; , and , are the initial latent representation and the updated representation after processing by the bidirectional cross-attention mechanism, respectively;

[0115] S3, Information aggregation optimization based on the generated comprehensive representation as input through a two-level contrastive learning framework: The two-level contrastive loss is composed of a high-level contrastive loss and a semantic-level contrastive loss, and the two-level contrastive learning is performed to generate an optimized representation;

[0116] S4, Joint optimization and clustering: Based on the K-means algorithm extension, the cancer subtype data clustering is performed by using the optimized representation as the initial clustering input through the iterative optimization algorithm, and the clustering input is jointly optimized through the clustering loss, the reconstruction loss and the semantic contrastive loss, to improve the clustering performance and finally output the clustering result.

[0117] In this embodiment, the shared features across omics are extracted and shared latent representations are generated through a shared encoder, and specific features of each omics are extracted and omics-specific latent representations are generated through an omics-specific encoder, as follows:

[0118] S1.1, the non-missing omics data of each sample in the incomplete multi-omics cancer subtype data is used as input; wherein: is the th sample, the th omics, and the th feature of the th sample. omics data, and the data normalized using min-max normalization, representing the sample of the non-missing omics set, , is a real set, is the feature dimension;

[0119] S1.2, mapping the non-missing omics of each sample to a latent space through a shared encoder and omics-specific encoders, generating Gaussian distribution parameters:

[0120]

[0121]

[0122] wherein: is the th omics-specific encoder; , respectively represent the mean and standard deviation of the approximate posterior distribution generated with data as input of the omics-specific encoder ; is the shared encoder; , respectively represent the mean and standard deviation of the approximate posterior distribution outputted with data as input of the shared encoder ; , are respectively the parameters of the omics-specific encoder and the shared encoder ;

[0123] The approximate posterior distribution is represented as:

[0124]

[0125] wherein: represents the approximate posterior distribution defined by the output of the omics-specific encoder , represents the approximate posterior distribution defined by the output of the shared encoder ; represents a Gaussian distribution; represents a diagonal matrix; and respectively represent the parameters obtained by averaging the shared Gaussian distribution parameters for all non-missing omics of the sample , thus ensuring that the shared representation integrates cross-omics shared information:

[0126] ,

[0127] 、 respectively as the sample th omics-specific latent representation and shared latent representation, for subsequent multi-omics factor analysis framework and contrastive learning, generated using reparameterization trick:

[0128]

[0129] where, denotes element-wise multiplication, is a random vector following standard normal distribution , and I is an identity matrix.

[0130] In this embodiment, the lower bound of evidence is optimized by variational inference:

[0131]

[0132] where, denotes the sample th omics-specific variational autoencoder loss, is the KL divergence weight, denotes the KL divergence, and denote the prior distribution set for capturing the omics-specific latent representation distribution and shared latent representation distribution respectively, denotes the sample th omics-specific reconstruction loss:

[0133] reconstruct the data from the omics-specific latent representation and shared latent representation by the omics-specific decoder

[0134]

[0135] where, is the omics specific decoder, is the parameter of the decoder ;

[0136] The reconstruction distribution defined by the decoder output is denoted as:

[0137]

[0138] where, is a fixed hyperparameter;

[0139] The loss function is then represented as:

[0140]

[0141] where: denotes the expectation under the approximate posterior distribution denotes the expectation under the approximate posterior distribution denotes the omics-specific log-likelihood, denotes the shared log-likelihood, the log-likelihood term and and the reconstruction distribution are related through the optimization objective ELBO of the variational autoencoder.

[0142] In this embodiment, step S2 is implemented as follows:

[0143] S2.1, concatenate the shared latent representation and the omics-specific latent representation of each sample generated in step S1 to obtain an initial latent representation and take the initial latent representation as the input data, the dimension of the initial latent representation , is the feature dimension of the shared latent representation and the omics-specific latent representation after concatenation for each sample, , is the feature dimension of the shared latent representation for each sample, is the feature dimension of the omics-specific latent representation for each sample;

[0144] S2.2, a forward diffusion process is used to gradually add Gaussian noise:

[0145] , ;

[0146] wherein: are the latent representations at time step t in the forward diffusion process or the backward denoising process, respectively, the latent representation is the initial latent representation in the forward diffusion process, and the latent representation is the predicted initial latent representation in the backward denoising process, is the total number of time steps in the forward diffusion process or the backward denoising process (in this embodiment, ); is the latent representation in the forward diffusion process given the latent representation ​​a probability distribution of the latent representation obeys a Gaussian distribution; denotes the noise schedule coefficient at time step t in the forward diffusion process;

[0147] S2.3, the reverse denoising process uses a U-Net architecture and adopts a bidirectional cross-attention mechanism to fuse features in the middle layer, learn the denoising distribution, and optimize parameters by minimizing the variational lower bound:

[0148]

[0149] wherein: is a probability distribution of the latent representation given the latent representation at time step t in the reverse denoising process, learned by the U-Net architecture, and the U-Net architecture parameters are ; is a mean function with the U-Net architecture parameterized by , taking the latent representation and the time step t as input; is a covariance function with the U-Net architecture parameterized by , taking the latent representation and the time step t as input;

[0150] S2.4, starting from the noise distribution , the comprehensive representation is generated by sampling through the reverse denoising process , and the noise distribution obeys a standard normal distribution ; the sampling update formula in the reverse denoising process is:

[0151] , ,

[0152] wherein: is the cumulative product term of the noise schedule coefficient , is the index of the time step t, is the standard deviation parameter at time step t in the reverse denoising process;

[0153] Repeat the sampling S times, take the average value to obtain the comprehensive representation , and project it back to dimension;

[0154]

[0155] wherein, denotes the initial latent representation predicted by the s'th sampling.

[0156] In this embodiment, the parameter optimization is performed by minimizing the variational lower bound , which is specifically expressed as:

[0157]

[0158] wherein: represents the expectation of the probability distribution of the forward diffusion process, represents the probability distribution of the initial latent representation predicted by the model, by optimizing so that the model-predicted is close to the real data distribution, i.e., the initial latent representation ; represents the KL divergence, represents the probability distribution of the latent representation predicted by the forward diffusion process at the final time step , represents a predefined prior distribution; represents the probability distribution of the latent representation in the forward diffusion process given the known initial latent representation and the latent representation ; this embodiment uses the Adam optimizer with a learning rate of 1e-4 and a batch size of 32.

[0159] In this embodiment, step S3 is specifically implemented as follows:

[0160] S3.1, data augmentation:

[0161] For each row vector corresponding to a sample in the comprehensive representation , a random perturbation is added to generate an augmented representation :

[0162] ,

[0163] wherein represents a random perturbation, , , is a preset standard deviation (in this embodiment );

[0164] For each row vector corresponding to a sample in the comprehensive representation , a column index is randomly selected for feature permutation to generate an augmented representation , ;

[0165] S3.2 Extract the augmented representation corresponding to each sample using a multilayer perceptron (MLP). Advanced features and enhancement representation Advanced features And calculate the high-level contrast loss. :

[0166]

[0167] in: To control the similarity distribution of temperature parameters (in this embodiment, ), This is a high-level positive sample pair similarity matrix. , For high-level negative sample pair similarity matrix, , Cosine similarity:

[0168] ,

[0169] in: for transpose, It is an L2 norm;

[0170] S3.3 Extracting high-level features through cluster head mapping semantic representation and advanced features semantic representation And calculate semantic-level contrastive loss. :

[0171]

[0172] in This is a semantic-level similarity matrix of positive sample pairs. , This is a semantic-level similarity matrix of negative sample pairs. :

[0173] ,

[0174] in for Transpose of;

[0175] S3.4, Compare losses based on higher levels Loss compared to semantic level Obtain contrast loss :

[0176]

[0177] in To control the semantic loss weights (in this embodiment, );

[0178] Calculate contrast loss gradient And perform iterative updates to the parameters:

[0179]

[0180]

[0181] in The parameter set includes the weight matrix of the multilayer perceptron and the weight matrix of the clustering head (the weight matrix of the first layer transformation of the MLP). The weight matrix of the second layer transformation and And the weight matrix in the clustering head used to implement feature mapping and clustering decisions. ); The learning rate (in this embodiment, Train for 100 epochs until convergence;

[0182] Complete parameter set After iterative updates, the synthesized representation is adjusted through backpropagation. To optimize representation .

[0183] In this embodiment, the enhanced representation corresponding to each sample is extracted using a multilayer perceptron (MLP). Advanced features and enhancement representation Advanced features The details are as follows:

[0184]

[0185]

[0186] and Let represent the weight matrices of the first and second layer transformations of the multilayer perceptron, respectively. , , For the first hidden layer dimension, For the second hidden layer dimension, This is the ReLU activation function.

[0187] In this embodiment, the extraction of high-level features through cluster head mapping is described. semantic representation and high-level features semantic representation , in particular as follows:

[0188]

[0189]

[0190] wherein: is a weight matrix in the clustering head for implementing feature mapping and clustering decision, , is the number of clusters, is a Softmax activation function.

[0191] In this embodiment, step S4 is implemented as follows:

[0192] S4.1, using the K-means algorithm extension (i.e., the K-means++ algorithm) to select initial cluster centers from the optimized representation to optimize the initial clustering distribution, wherein c'=1,..., , c' represents the cluster index;

[0193] S4.2, iterative optimization:

[0194] S4.2.1, taking the optimized representation as the initial clustering input, for each sample in the optimized representation corresponding to the row vector , calculate the Euclidean distance between the row vector and all cluster centers , and assign the row vector to the cluster with the smallest distance: , represents the clustering label of the i-th sample, is the square of the distance between the row vector and the cluster center ;

[0195] S4.2.2, update the cluster center : wherein, represents the summation operation on all row vectors belonging to the cluster , is the number of samples in the cluster ;

[0196] S4.2.3, update the clustering loss , the reconstruction loss and semantic contrast loss Jointly optimize the clustering input, and obtain the clustering result according to the jointly optimized clustering input again; total loss is expressed as:

[0197]

[0198] wherein , represents the reconstruction loss weight and semantic contrast loss weight;

[0199] S4.3 outputs the clustering result after reaching the iteration termination condition, including the final clustering label of each sample and the final cluster center of each cluster .

[0200] In this embodiment, the clustering loss , the reconstruction loss and the semantic contrast loss are specifically calculated as follows:

[0201] Clustering loss:

[0202]

[0203] Reconstruction loss:

[0204]

[0205] wherein, is the original ith sample data, is the result of reconstructing the row vector using the decoder of step S1, ;

[0206] Semantic contrast loss:

[0207]

[0208] wherein, and are two semantic representations of the row vector obtained by the data enhancement method of step S3.1, the advanced feature extraction method of step S3.2 and the semantic representation extraction method of step S3.3, respectively.

[0209] The above is the preferred embodiment of the present application, any changes made according to the technical solutions of the present application, as long as the generated function does not exceed the scope of the technical solutions of the present application, all belong to the protection scope of the present application.

Claims

1. An incomplete multi-omic cancer subtype data clustering method, characterized in that, Specifically comprising: S1, omics encoding based on a variational autoencoder framework: for the non-missing omics data of each sample in the incomplete multi-omics cancer subtype data, shared features across omics are extracted through a shared encoder to generate shared latent representations, and omics-specific features are extracted through omics-specific encoders to generate omics-specific latent representations; S2, information aggregation based on a multi-omics diffusion model framework: taking the generated shared latent representations and omics-specific latent representations as initial latent representations, capturing the cross-omics shared biological features and omics-specific variations through adding noise in the forward diffusion process, and integrating the hidden links between omics through a bidirectional cross-attention mechanism in the backward denoising process to generate a joint distribution, and finally repeating sampling from the noise distribution through the backward denoising process and generating a comprehensive representation; S3, information aggregation optimization by a two-level contrastive learning framework with the generated comprehensive representation as input: a two-level contrastive loss composed of a high-level contrastive loss and a semantic-level contrastive loss, two-level contrastive learning and generating an optimized representation; S4, joint optimization and clustering: based on the K-means algorithm extension, taking the optimized representation as the initial clustering input to perform cancer subtype data clustering through an iterative optimization algorithm, and simultaneously optimizing the clustering input through a clustering loss, a reconstruction loss and a semantic contrastive loss to improve the clustering performance and finally output the clustering results.

2. The incomplete multi-omic cancer subtype data clustering method of claim 1, wherein, The shared features across omics are extracted through a shared encoder to generate shared latent representations, and omics-specific features are extracted through omics-specific encoders to generate omics-specific latent representations, specifically as follows: S1.1, non-missing omics data of each sample in incomplete multi-omics cancer subtype data as input; wherein: is the jth omics data of the ith sample, and is the jth omics data of the ith sample, and is the jth omics data of the ith sample, and is the jth omics data of the ith sample, and denotes the non-missing omics set of sample denotes the non-missing omics set of sample , is a real set, is the feature dimension; S1.2, map the non-missing omics of each sample to the latent space through the shared encoder and the omics-specific encoder to generate Gaussian distribution parameters: wherein: is a first omics-specific encoder; , respectively denote the mean and standard deviation of the approximate posterior distribution generated with data as input to the omics-specific encoder ; is a shared encoder; , respectively denote the mean and standard deviation of the approximate posterior distribution output with data as input to the shared encoder ; , are respectively parameters of the omics-specific encoder and the shared encoder ; The approximate posterior distribution is represented as: wherein, represents a shared encoder by omics outputs an approximated posterior distribution defined by represents a shared encoder by omics outputs an approximated posterior distribution defined by represents a Gaussian distribution; represents a diagonal matrix; and represent parameters obtained by averaging over all non-missing omics of samples represent parameters obtained by averaging over all non-missing omics of samples , , Samples No. Omics-specific and shared latent representations for each omics are generated using reparameterization techniques: wherein, denotes element-wise multiplication, is a random vector subject to a standard normal distribution , I is the identity matrix.

3. The incomplete multi-omic cancer subtype data clustering method of claim 2, wherein, Optimize the evidence lower bound through variational inference: wherein: represents a sample the thomics-specific variational autoencoder loss, is a KL divergence weight, represents a KL divergence, and represent prior distributions set for capturing the omics-specific latent representation distribution and the shared latent representation distribution, respectively, represents a sample the thomics-specific reconstruction loss: Reconstructing data from omics specific latent representations and shared latent representations by omics specific decoders : wherein, as omics a particular decoder, as a decoder parameters; By the decoder Output defined reconstruction distribution Is represented as: wherein, is a fixed hyperparameter; The loss function is represented as: where: denotes the expectation under the approximate posterior distribution denotes the expectation under the approximate posterior distribution denotes the omics-specific log-likelihood, denotes the shared log-likelihood, the log-likelihood term and and the reconstruction distribution are interconnected by the optimization objective ELBO of the variational autoencoder.​​ 4. The incomplete multi-omic cancer subtype data clustering method of claim 2, wherein, The specific implementation of step S2 is as follows: S2.1, concatenating the shared latent representation of each sample generated in step S1 and the omics-specific latent representation to obtain an initial latent representation and inputting the initial latent representation as input data has a dimension of , is a feature dimension of the shared latent representation of each sample , is a feature dimension of the omics-specific latent representation of each sample is a feature dimension of the shared latent representation of each sample is a feature dimension of the omics-specific latent representation of each sample is a feature dimension of the shared latent representation of each sample S2.2, gradually add Gaussian noise in the forward diffusion process: , ; wherein: are latent representations at time step t in a forward diffusion process or a reverse denoising process, respectively, is an initial latent representation is a predicted initial latent representation in a reverse denoising process, is a total number of time steps in the forward diffusion process or the reverse denoising process; is a probability distribution of a latent representation given a latent representation at time step t in the forward diffusion process, obeys a Gaussian distribution; denotes a noise schedule coefficient at time step t in the forward diffusion process. S2.3, the backward denoising process uses a U-Net architecture and adopts a bidirectional cross-attention mechanism to fuse features at the intermediate layer, learn the denoising distribution, and optimize the parameters by minimizing the variational lower bound: wherein: is the probability distribution of the latent representation at time step t given the latent representation is learned by a U-Net architecture with parameters ; is a mean function learned by a U-Net architecture with parameters taking the latent representation and the time step t as input; is a covariance function learned by a U-Net architecture with parameters taking the latent representation and the time step t as input; S2.4, from the noise distribution Start, by the reverse denoising process Step sample, generating a comprehensive representation , the noise distribution Subject to the standard normal distribution The sampling update formula in the reverse denoising process is: , , wherein: is a noise schedule coefficient is a cumulative product term, is an index of time step t, is a standard deviation parameter for time step t in the reverse denoising process; Repeat S times, take average to get the integrated representation and project back dimensions; wherein, represents the initial latent representation of the s'th sampled prediction.

5. The incomplete multi-omic cancer subtype data clustering method of claim 4, wherein, by minimizing a variational lower bound Parameter optimization is performed, specifically expressed as: in: This represents the expectation of the probability distribution of the forward diffusion process. The initial latent representation of the model prediction The probability distribution, through optimization This makes the model predict A data distribution that approximates the true value, i.e., the initial latent representation. ; Denotes KL divergence, This indicates that the forward diffusion process is at the final time step. Time potential representation probability distribution Represents a predefined prior distribution; This indicates that during the forward diffusion process, given the initial potential representation... and potential representation Under what conditions is the latent representation The probability distribution.

6. The incomplete multi-omic cancer subtype data clustering method of claim 4, wherein, The specific implementation of step S3 is as follows: S3.1, data augmentation: For the integrated representation a row vector corresponding to each sample , adding a random perturbation generates an enhanced representation : , wherein represents a random perturbation, , is a preset standard deviation; For the integrated representation a row vector corresponding to each sample , randomly selecting a column index to perform feature permutation and generate an enhanced representation , ; S3.2, extracting the enhanced representation corresponding to each sample by a multi-layer perceptron, MLP high-level features of the enhanced representation and the enhanced representation high-level features of the enhanced representation ; and computing a high-level contrastive loss : wherein: is a temperature parameter controlling the similarity distribution, is a high-level positive sample pair similarity matrix, , is a high-level negative sample pair similarity matrix, , denotes the cosine similarity: , wherein: is the transpose of is the L2 norm;​ S3.3 Extracting high-level features through cluster head mapping semantic representation and advanced features semantic representation And calculate semantic-level contrastive loss. : wherein is a semantic-level positive sample pair similarity matrix, , is a semantic-level negative sample pair similarity matrix, : , wherein is the transpose of ; S3.4, according to the high-level contrastive loss and semantic-level contrastive loss obtaining the contrastive loss : wherein is a control semantic loss weight; computing the gradient of the contrastive loss and performing an iterative update of the parameters ​ wherein is a parameter set including weight matrices of the multi-layer perceptron and weight matrices of the clustering heads; is a learning rate; complete set of parameters After iterative updates of the set of parameters to the optimized representation by backpropagation 7. The incomplete multi-omic cancer subtype data clustering method of claim 6, wherein, extracting, by a multi-layer perceptron (MLP), an enhanced representation corresponding to each sample high-level features of the enhanced representation high-level features of the enhanced representation high-level features of the enhanced representation as follows: and W1and W2denote the weight matrices of the first and second layer transformations of the multi-layer perceptron, respectively, , , is the dimension of the first hidden layer, is the dimension of the second hidden layer, is the ReLU activation function.

8. The incomplete multi-omic cancer subtype data clustering method of claim 6, wherein, The advanced features are extracted by mapping through clustering heads of semantic representations and advanced features of semantic representations as follows: wherein: is a weight matrix in the cluster head for implementing feature mapping and clustering decision, , is the number of clusters, is a Softmax activation function.

9. The incomplete multi-omic cancer subtype data clustering method of claim 8, wherein, The specific implementation of step S4 is as follows: S4.1, extending the optimal representation using the K-means algorithm selected initial cluster centers to optimize the initial clustering distribution, where c'=1,..., c' denotes the cluster index; S4.2, iterative optimization: S4.2.1, with optimized representation As initial clustering input, for each sample in the optimized representation a row vector is computed, which corresponds to the sample in the optimized representation The Euclidean distance between the row vector and all cluster centers is computed and the row vector is assigned to the cluster with the smallest distance: , denotes the cluster label of the i-th sample, is the squared distance of the row vector to the cluster center S4.2.2, updating the cluster center : where, denotes the summation operation over all row vectors belonging to the cluster , and is the number of samples in the cluster . S4.2.3, by clustering loss , reconstruction loss and semantic contrastive loss jointly optimize the clustering input, and obtain a clustering result according to the jointly optimized clustering input again; the total loss is represented as: wherein , denotes the reconstruction loss weights and semantic contrastive loss weights; S4.3 Output the clustering result after reaching the iteration termination condition, including the final clustering label of each sample and the cluster center of each cluster .

10. The incomplete multi-omic cancer subtype data clustering method of claim 9, wherein, clustering loss , reconstruction loss and semantic contrastive loss The specific calculation is as follows: Clustering loss: Reconstruction loss: wherein is the original ith sample data, is the result of reconstructing the row vector with the decoder of step S1, ; Semantic contrastive loss: wherein, and are two semantic representations of the row vector respectively obtained by the data augmentation method of step S3.1, the advanced feature extraction method of step S3.2 and the semantic representation extraction method of step S3.3.

Citation Information

Patent Citations

  • Multi-omics clustering method based on double subspace representation

    CN117577193A

  • Multi-omics data analysis-based pan cancer analysis method and system

    CN118136099A

  • Genomics cross-modal condition generation method and system based on potential diffusion model

    CN120183507A

  • Single-cell multi-omics data clustering method based on diffusion auto-encoder

    CN120541558A

  • Cancer subtype classification and prognosis prediction method based on generic cancer multi-omics data

    CN120670897A