Method and system for clustering transcriptome sequencing data
By combining the cluster analysis method of potential category models and sub-consortium division and the deep cluster prediction model based on the variational autoencoder and Gamma hybrid model, the limitations of traditional methods in terms of transcriptome data clustering and prediction are solved, and more stable and accurate clustering effects and higher prediction accuracy are achieved.
Patent Information
- Application Number
- CN202510619545.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional clustering methods have limitations in processing transcriptomic data, making it difficult to fully explore deep features and potential relationships in the data, resulting in poor clustering effect and ineffective prediction of prognostic risk and chemotherapy drug sensitivity in patients with cervical adenocarcinoma.
A cluster analysis method combining latent category model and sub-consortium division is used to model sample gene expression data through factor analysis and latent category model, and combined with sub-consortium division constrained by Nash stability, the gene expression similarity and potential category joint distribution are maximized. Then, a deep cluster prediction model was constructed based on the variant autoencoder and Gamma hybrid model to predict the prognostic risk of patients and chemotherapy sensitivity.
It improves the stability and accuracy of transcriptome data clustering, enhances the ability to model gene expression data, and significantly improves the predictive ability to prognostic risk and chemotherapy drug sensitivity in patients with cervical adenocarcinoma.
Smart Images

Figure CN120126581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a clustering method and system for transcriptome sequencing data. Background Art
[0002] Transcriptome sequencing technology can comprehensively and rapidly obtain the information of organism transcripts, and plays a key role in the field of tumor research, especially in the research of cervical adenocarcinoma. Through transcriptome sequencing, a large amount of gene expression data related to cervical adenocarcinoma can be obtained. These data contain rich biological information, which is of great significance for revealing the molecular mechanism of cervical adenocarcinoma, exploring potential therapeutic targets, predicting the prognosis risk of patients and the sensitivity of chemotherapy drugs, etc.
[0003] However, there are many problems in the raw data generated by transcriptome sequencing, such as containing low-quality bases, adapter sequences and contaminated data, and there is a lack of comparability between different sample data, which seriously affects the accuracy and reliability of subsequent data analysis. At the same time, how to effectively mine valuable information from the vast amount of gene expression data, accurately classify samples or genes, and discover potential data structures and patterns has always been a challenge in the field of bioinformatics.
[0004] Traditional clustering analysis methods, such as the K-Means clustering algorithm and the Gaussian mixture model, have limitations in processing complex high-dimensional transcriptome data, and it is difficult to fully mine the deep features and potential relationships in the data, resulting in poor clustering effects. In predicting the prognosis risk and chemotherapy drug sensitivity of cervical adenocarcinoma patients, the existing models and methods also cannot well adapt to the characteristics of transcriptome data, resulting in low prediction accuracy. Therefore, there is an urgent need for a new method to solve these problems to promote the in-depth development of cervical adenocarcinoma research and provide more powerful support for clinical diagnosis and treatment. Summary of the Invention
[0005] Aiming at the deficiencies of traditional clustering methods in processing transcriptome data, the present invention proposes a clustering method and system for transcriptome sequencing data.
[0006] To achieve the above object, it is realized through the following technical solutions:
[0007] According to the first aspect of the present invention, a clustering method for transcriptome sequencing data is provided, including the following steps:
[0008] Collect cervical adenocarcinoma and adjacent tissue specimens of several patients, perform transcriptome sequencing on the specimens and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix;
[0009] A clustering analysis method combining a latent class model and sub - coalition partitioning is used to perform clustering operations on the standardized gene expression matrix. By factor analysis and latent class model to model the sample gene expression data, combined with sub - coalition partitioning with Nash stability constraints, the gene expression similarity and the joint distribution of latent classes are maximized to obtain the clustered gene set;
[0010] Gene function annotation is performed on the clustered gene set, and the association analysis is carried out between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;
[0011] Based on the results of clustering analysis and association analysis, a deep clustering prediction model based on variational auto - encoder and Gamma mixture model is constructed to predict the prognosis risk and chemotherapy sensitivity of patients.
[0012] Furthermore, the clustering operation of the standardized gene expression matrix by the clustering analysis method combining the latent class model and sub - coalition partitioning includes:
[0013] Construct a gene expression data model using factor analysis:
[0014]
[0015] where \(X\) is the gene expression data, \(\Lambda\) is the factor loading matrix, \(\varepsilon\) is the error term, is the mixing coefficient of class \(m\), is the parameter of class \(m\); represents the true label of the \(i\) - th sample; \(M\) is the total number of classes;
[0016] Calculate the gene expression similarity between samples based on the cosine similarity metric function, and divide sub - coalitions according to the preset threshold \(\delta\);
[0017] During the sub - coalition partitioning process, each sample \(i\) belongs to the \(k\) - th sub - coalition , and its individual utility is defined as , where, represents the features of sample \(i\) and sample \(j\); represents the \(k\) - th sub - coalition; is the cosine similarity metric function between sample \(i\) and sample \(j\); : the individual utility of sample \(i\), that is, the sum of its similarities with all other samples in the same sub - coalition;
[0018] Ensure that the individual utility of samples in the sub - coalition is not lower than the individual utility after migrating to other sub - coalitions through Nash stability constraints;
[0019] The optimization objective function includes gene expression similarity and the joint distribution of latent classes.
[0020] Furthermore, the Nash stability condition for the sub - coalition division is as follows:
[0021]
[0022] That is, for any sample i, the individual utility of its current sub - coalition , where is the individual utility of sample i in its current affiliated sub - coalition ; is the individual utility of sample i if it migrates to other sub - coalition .
[0023] Furthermore, the correlation analysis between the clustered results after functional annotation and the clinical features of cervical adenocarcinoma is carried out, and the correlations between gene expression and clinical features include:
[0024] Pearson correlation coefficient is used to analyze the linear relationship between gene expression and clinical features;
[0025] t - test is used to compare the gene expression differences between different clinical groups.
[0026] Furthermore, based on the results of cluster analysis and correlation analysis, a deep clustering prediction model based on variational auto - encoder and Gamma mixture model is constructed to predict the prognosis risk and chemotherapy sensitivity of patients, including:
[0027] The variational auto - encoder is used to reduce the dimension of high - dimensional gene expression data to extract potential features, and a Gamma mixture model is used to model the asymmetric distribution in the latent space for clustering the latent space to identify molecular subgroups with different prognosis risks;
[0028] The Gamma distribution parameters are optimized through the reparameterization trick, and the model is trained by combining the cross - entropy loss and the objective function weighted by the maximum evidence lower bound (ELBO) to output the prediction results of the prognosis risk and chemotherapy sensitivity of patients.
[0029] Furthermore, the combination of the variational auto - encoder and the Gamma mixture model includes:
[0030] The probability density function of the Gamma mixture model is defined in the latent space of the variational auto - encoder; the input of the variational auto - encoder is high - dimensional gene expression data X, and the output is the distribution parameters of the latent embedding representation variable z;
[0031] The Gamma distribution sampling is differentiably optimized through the Gaussian distribution reparameterization trick, the Gamma distribution parameters are optimized to adapt to the deep learning framework, and the model is trained by combining the cross - entropy loss and the objective function weighted by the maximum evidence lower bound to output the prediction results of the prognosis risk and chemotherapy sensitivity of patients.
[0032] Further, the sampling of the Gamma distribution is differentiably optimized by the Gaussian distribution reparameterization technique to optimize the Gamma distribution parameters, including the following calculation formulas:
[0033]
[0034] where is the mean term after parameterization of the Gamma distribution; is the standard deviation term of the Gamma distribution; is the noise variable of the standard normal distribution; represents element-wise multiplication; represents the latent variable sample generated by the reparameterization technique of the Gaussian distribution;
[0035] Calculate the Gamma distribution parameters through the parameter conversion formula:
[0036]
[0037]
[0038] where is the mean output by the variational autoencoder; is the learnable parameter that controls the shape of the Gamma distribution; is the normalization term; respectively represent the shape parameter and rate parameter of the Gamma distribution.
[0039] Further, the decoder output of the variational autoencoder is mapped to the patient prognosis risk or chemotherapy sensitivity prediction label through a multi-layer perceptron (MLP):
[0040]
[0041] where is the predicted patient prognosis risk or chemotherapy drug sensitivity; is the multi-layer perceptron classifier for mapping z to the specific prediction label.
[0042] According to the second aspect of the present invention, there is also provided a clustering system for transcriptome sequencing data, including:
[0043] A data acquisition and preprocessing module, configured to collect cervical adenocarcinoma and adjacent tissue specimens of several patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix;
[0044] A clustering analysis module, which is used to perform clustering operations on the standardized gene expression matrix by using a clustering analysis method combining a latent class model and sub - coalition partitioning. By factor analysis and latent class model to model the sample gene expression data, combined with sub - coalition partitioning with Nash stability constraints, it maximizes gene expression similarity and latent class joint distribution to obtain the clustered gene set;
[0045] An association analysis module, which is used to perform gene function annotation on the clustered gene set, and perform association analysis on the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;
[0046] A deep clustering prediction module, which is used to construct a deep clustering prediction model based on a variational auto - encoder and a Gamma mixture model based on the results of clustering analysis and association analysis to predict the prognosis risk and chemotherapy sensitivity of patients.
[0047] According to the third aspect of the present invention, there is also provided an electronic device including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the steps of a clustering method for transcriptome sequencing data described in the first aspect.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] 1. The present invention integrates a clustering method of a latent class model and sub - coalition partitioning. By factor analysis to model the global distribution characteristics of gene expression, combined with sub - coalition partitioning based on cosine similarity to refine the local structure, and introducing Nash stability to constrain sample attribution, it solves the problem of poor clustering stability of high - dimensional transcriptome data.
[0050] 2. The present invention designs a deep joint model of a variational auto - encoder (VAE) and a Gamma mixture model. The variational auto - encoder is used for high - dimensional data dimensionality reduction to learn latent expression patterns, and the Gamma mixture model is used to model the asymmetric distribution in the latent space to make up for the deficiency of the traditional Gaussian mixture model in dealing with asymmetric data distributions. By optimizing the Gamma distribution parameters, the modeling ability of the model for gene expression data is enhanced, making the clustering results more stable, and improving the prediction ability for the prognosis risk and chemotherapy drug sensitivity of patients, overcoming the defect of poor adaptability of the traditional Gaussian mixture model to skewed data.
[0051] 3. The present invention constructs an end-to-end prognosis and chemotherapy sensitivity prediction framework. After associating the clustering results with clinical features and inputting them into a deep model, weighted optimization is performed through cross-entropy loss and maximized evidence lower bound loss. This loss function helps improve the clustering and prediction capabilities of the model on cervical adenocarcinoma gene expression data, achieving a prediction AUC of 0.85 for patient prognosis risk (0.68 for traditional logistic regression), and the F1 score of chemotherapy sensitivity is increased to 0.79.
[0052] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where:
[0054] Figure 1 is a schematic diagram of the specific steps of a clustering method for transcriptome sequencing data according to an embodiment of the present invention;
[0055] Figure 2 is a schematic flowchart of a clustering method for transcriptome sequencing data according to an embodiment of the present invention;
[0056] Figure 3 is a schematic diagram of the modules of a clustering system for transcriptome sequencing data according to an embodiment of the present invention;
[0057] Figure 4 is a performance comparison chart of the adjusted Rand index of different models;
[0058] Figure 5 is a performance comparison chart of the Davies-Bouldin index of different models;
[0059] Figure 6 is a performance comparison chart of the normalized mutual information of different models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.
[0061] In addition, the term "and / or" in this text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally indicates that the related objects before and after are in an "or" relationship.
[0062] Prognosis: Refers to the comprehensive assessment of the future disease course development, treatment effect, survival period, recurrence risk, etc. of a patient after the disease is diagnosed.
[0063] Prognosis prediction: Quantify the potential outcomes of patients through statistical models or machine learning methods, combined with multi-dimensional data (such as genes, pathology, clinical indicators).
[0064] Figure 1 The figure shows a schematic diagram of the specific steps of a clustering method for transcriptome sequencing data. Figure 2 The figure shows a schematic flow diagram of a clustering method for transcriptome sequencing data. As Figure 1 and Figure 2 shown, a clustering method 100 for transcriptome sequencing data includes the following steps:
[0065] S110: Collect cervical adenocarcinoma and adjacent tissue specimens from several patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix.
[0066] Data collection: From 21 - 75-year-old female patients with cervical adenocarcinoma in the research project "Exploring the Molecular Mechanisms and Potential Therapeutic Targets of Cervical Adenocarcinoma Based on Common Transcriptome Sequencing Data", collect 8 cervical adenocarcinoma and adjacent tissue specimens (8 specimens each for cancer and adjacent tissues). The specimens are collected for subsequent experiments after obtaining the informed consent of the patients. Perform transcriptome sequencing on the collected specimens to obtain the original transcriptome sequencing data (in FASTQ format), providing a basis for subsequent analysis.
[0067] Preprocessing of the original transcriptome sequencing data: Use professional software tools to perform quality control on the original sequencing data, remove low-quality bases, adapter sequences, and contaminated data to improve data quality. Perform standardization processing on the data after quality control to correct the differences between the data and ensure the comparability of data from different samples.
[0068] After the above preprocessing, a high-quality and standardized gene expression matrix is obtained (matrix dimension: sample × gene, and the value is the expression level).
[0069] S120: performing a clustering operation on the standardized gene expression matrix based on a clustering analysis method combining a latent class model with sub-coalition division, modeling sample gene expression data through factor analysis and latent class model, maximizing gene expression similarity and latent class joint distribution in combination with Nash stability constraint sub-coalition division, and obtaining a clustered gene set;
[0070] This step S120 is based on the clustering analysis of latent classes and sub-alliances to obtain the sub-alliances division results after clustering (the sub-alliances labels to which each sample belongs) and the latent class model parameters (to explain the latent structure of gene expression data). The latent class model and sub-alliances complement and cooperate with each other. The latent class model captures the global distribution, and the sub-alliances refine the local structure, thereby overcoming the sparsity of high-dimensional data.
[0071] Among them, sub-alliances are sample subsets divided based on gene expression similarity. Samples within the same sub-alliances have higher homogeneity in biological functions (such as gene co-expression and clinical phenotype), that is, a collection of samples with highly similar gene expression patterns, reflecting the molecular subtypes of cervical adenocarcinoma (such as immune infiltration type and metabolically active type). Difference from traditional clustering: Sub-alliances emphasize local similarities between samples (such as gene co-expression networks) rather than global distances (such as the Euclidean distance of K-Means).
[0072] The preprocessed data was input into the selected algorithm, and cluster analysis was performed on the standardized gene expression matrix and clinical feature data (such as tumor stage, prognostic label, chemotherapy sensitivity label) according to the clinical characteristics and gene expression patterns of cervical adenocarcinoma. The clustering analysis method that combines latent classes + sub-alliances was selected for clustering operations, and samples or genes were classified according to the similarity of gene expression to mine potential data structures and patterns. Among them, gene expression data It can be expressed using factor analysis and latent class models:
[0073]
[0074] Among them, X is the observed variable (gene expression data), Λ is the factor loading matrix, ε is the error term, is the mixing coefficient of class m, Yes Category Parameters; represents the true label of the i-th sample; M is the total number of categories; It is a latent class model, which aims to explain the joint distribution of indicators by determining latent classes under the premise that different latent classes present different distributions of indicators.
[0075] The goal of the sub-alliance is to divide multiple sub-alliances based on the similarity of gene expression Make the samples in the same sub - alliance more similar in gene expression and obtain a social weak order by aggregating individual preferences.
[0076] Define the gene expression similarity measure:
[0077]
[0078] Where, is the gene expression vector of the i - th sample; : the gene expression vector of the j - th sample; : measures the similarity of the gene expression patterns of sample i and sample j, calculated using cosine similarity. A value close to 1 indicates that the gene expressions of the two samples are very similar, close to 0 indicates no obvious correlation in gene expression, and close to - 1 indicates that the gene expression patterns are exactly opposite.
[0079] Here, Use cosine similarity to calculate the similarity of gene expression patterns. The closer the value is to 1, the more similar the gene expressions of sample i and sample j are. To screen for appropriate sub - alliances, the present invention sets a similarity threshold δ. If , then sample i and sample j belong to the same sub - alliance, indicating that their gene expression patterns are similar enough. If , then sample i and sample j do not belong to the same sub - alliance, indicating that there are significant differences in their gene expressions.
[0080] Among them, in the present invention, δ is optimized by grid search combined with survival analysis: test δ∈[0.7, 0.9] with an interval of 0.05 on the training set. Perform sub - alliance division for each δ value and calculate the difference in survival curves (Log - rank test P - value) of each sub - alliance. Select the δ value that minimizes the P - value (i.e., the most significant subtype prognosis differentiation), for example, take δ = 0.85.
[0081] During the sub - alliance division process, each sample i belongs to the k - th sub - alliance , and its individual utility can be defined as:
[0082]
[0083] Where, : represents the features of sample i and sample j (such as gene expression data); : represents the k - th sub - alliance; : the cosine similarity measure function between sample i and sample j; : the individual utility of sample i, that is, the sum of its similarities with all other samples in the same sub - alliance.
[0084] That is, the individual utility of sample \(i\) depends on the sum of its similarities with other samples within the same sub - coalition.
[0085] The goal of sub - coalition partitioning is to maximize the overall gene expression similarity:
[0086]
[0087] Among them, \(W(C)\): the overall utility function, representing the sum of the individual utilities of all samples in all sub - coalitions. The goal is to maximize this value to obtain a more compact sub - coalition partitioning. \(K\): the total number of sub - coalitions. At the same time, the Nash stability needs to be satisfied, that is:
[0088]
[0089] : the individual utility of sample \(i\) in its current sub - coalition ; : the individual utility of sample \(i\) if it migrates to sub - coalition .
[0090] This condition requires that the individual utility of each sample \(i\) in its own sub - coalition is not lower than its individual utility after migrating to other sub - coalitions, that is, an individual will not increase its utility by joining other sub - coalitions. Otherwise, it will migrate to a new sub - coalition until a stable state is reached.
[0091] The Nash stability constraint requires that the individual utility of a sample in its current sub - coalition is not lower than its individual utility after migration, thus preventing samples from frequently switching during the clustering process, ensuring that the membership of samples in sub - coalitions will not change due to small perturbations (such as noise or measurement errors), and improving the robustness of the clustering results. In cervical adenocarcinoma data, this constraint can improve the repeatability of subtype classification (such as multiple test results of the same patient being stably assigned to the same sub - coalition), enhance the reliability of clinical applications, be able to avoid samples from frequently switching sub - coalitions during the iterative process, ensure the stability of patient classification, and support clinical decisions (such as chemotherapy regimen selection).
[0092] Finally, the objective function for optimizing the clustering analysis method of latent class + sub - coalition includes gene expression similarity and the joint distribution of latent classes. The objective function \(C\) * is:
[0093]
[0094] Among them, \(W(C)\) guides sub - coalition partitioning by gene expression similarity measurement; \(C\) * : the final optimal sub - coalition partitioning scheme; \(\lambda\): the balance parameter, controlling the weights of gene expression similarity and the part of the class model calculating class probabilities; : the prior probability of the \(k\) - th sub - coalition; \(K\): the total number of sub - coalitions; \(N\): the total number of samples; : The conditional probability that sample i belongs to this category (i.e., belongs to the k-th sub - coalition) given the parameters of the k-th sub - coalition ; When : The probability of the calculated category by the latent class model, which measures the possibility that a sample belongs to different categories.
[0095] By optimizing the weight λ, optimizing the category parameters and the mixing coefficient to obtain the best classification of cervical adenocarcinoma gene expression.
[0096] This step S120 is based on sub - coalition partitioning. Samples with similar gene expression patterns are grouped into the same sub - coalition through the similarity threshold δ, and the co - expressed gene set within the sub - coalition is extracted; based on the latent class model, genes that contribute significantly to the category (high - loading genes) are identified through the factor loading matrix Λ to form a clustering gene set. Finally, a gene subset highly related to cervical adenocarcinoma (such as a differentially expressed gene cluster) can be identified through cluster analysis. These gene sets can be used as input features of the model to reduce the data dimension and filter out noise.
[0097] For example, assume that a cervical adenocarcinoma data set contains two types of patients. Sub - coalition partitioning: The patients are divided into sub - coalition A and B. The gene expressions of patients within A are highly similar (such as the activation of the EGFR pathway), and the similarity of patients within B is lower. Latent class model: It is found that the overall data can be divided into two types (such as "immune activation type" and "metabolic abnormality type"), and each type corresponds to a different survival prognosis. Joint optimization result: Patients in sub - coalition A may mainly belong to latent class 1 (immune activation type); patients in sub - coalition B may belong to latent class 2 (metabolic abnormality type). The final clustering result can reflect both local similarity and align with the global molecular subtypes.
[0098] S130: Perform gene function annotation on the clustered gene set, conduct correlation analysis between the functionally annotated clustering result and the clinical features of cervical adenocarcinoma, and obtain the correlation between gene expression and clinical features;
[0099] This step S130 is used to mine the association between gene expression patterns and clinical features through correlation analysis:
[0100] Perform functional annotation on the gene sets obtained by clustering in step S120. With the help of databases such as Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), clarify the biological processes, molecular functions, and signaling pathways involved by the genes, and deeply understand the functional significance of different clustered genes. If the genes are sorted according to the expression level, gene set enrichment analysis can be used for gene functional annotation. Perform correlation analysis on the clustering results and the clinical characteristics of cervical adenocarcinoma (such as tumor stage, patient prognosis risk, chemotherapy drug sensitivity, etc.), mine the potential relationship between the gene expression pattern and clinical characteristics, and provide guidance for clinical diagnosis and treatment.
[0101] To explore the correlation between gene expression and clinical characteristics (such as tumor stage, prognosis, chemotherapy sensitivity), the present invention can calculate the Pearson correlation coefficient (measuring the linear relationship between gene expression and the clinical variable Y):
[0102]
[0103] where : the gene expression value of the i-th sample; : the clinical variable corresponding to the i-th sample (such as tumor stage, prognosis, chemotherapy sensitivity, etc.); : the average gene expression of all samples; : the average clinical variable of all samples; n: the number of samples, that is, the number of patients.
[0104] If it is necessary to compare the differential expression levels of genes in two groups (such as chemotherapy-sensitive and drug-resistant), the t-test can be used:
[0105]
[0106] where respectively represent the average gene expression of the two groups of samples; respectively represent the standard deviations of the two groups of samples, that is, the degree of variation of the gene expression level; respectively represent the numbers of the two groups of samples.
[0107] In this step 130, according to the gene expression matrix and clinical characteristic data (such as stage, prognosis), through methods such as Pearson correlation coefficient (linear association) and t-test (difference between groups), the correlation between gene expression and clinical characteristics is finally obtained (such as gene A is significantly correlated with chemotherapy sensitivity).
[0108] S140: Based on the results of clustering analysis and correlation analysis, construct a deep clustering prediction model based on variational autoencoder and Gamma mixture model to predict the patient prognosis risk and chemotherapy sensitivity.
[0109] This step S140 is used to predict the prognosis risk and chemotherapy drug sensitivity of patients with cervical adenocarcinoma. Specifically, it includes:
[0110] S141: Construct a deep clustering prediction model
[0111] Based on the clustering analysis results and association analysis results of steps S120 and S130, construct a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model, that is, construct a model for predicting the prognosis risk or chemotherapy drug sensitivity of patients with cervical adenocarcinoma.
[0112] The model structure includes:
[0113] A variational autoencoder (VAE), including an encoder: reducing high-dimensional gene expression data to a latent embedding representation variable z; a decoder: reconstructing gene expression data from the latent space.
[0114] Gamma mixture model (GMM): Cluster samples in the latent embedding representation variable z and model the asymmetric distribution.
[0115] Input the standardized gene expression matrix (preprocessed high-dimensional data) and clinical labels (such as prognosis labels: high risk / low risk) into the model, use the reparameterization trick (introducing Gaussian noise) to solve the non-differentiable problem of Gamma distribution sampling and the loss function: cross-entropy loss (classification) + ELBO loss (reconstruction and distribution matching), and finally output the latent embedding representation variable z (low-dimensional representation for clustering), as well as the patient prognosis risk prediction result (such as high risk / low risk probability) and chemotherapy sensitivity prediction result (such as sensitive / resistant probability).
[0116] The latent space of the variational autoencoder (VAE) can be designed to capture the gene expression patterns after clustering. For example, during training, the encoder of the VAE can be guided by the clustering results to make the latent embedding representation variable z implicitly contain the structure information of sub-coalition division or latent categories.
[0117] Among them, this deep clustering prediction model uses a variational autoencoder to reduce the dimension of high-dimensional gene expression data to extract latent features, models the asymmetric distribution in the latent space with a Gamma mixture model, clusters the latent space to identify molecular subgroups with different prognosis risks; optimizes the Gamma distribution parameters through the reparameterization trick, combines the cross-entropy loss with the objective function weighted by maximizing the evidence lower bound (ELBO) to train the model, and outputs the patient prognosis risk and chemotherapy sensitivity prediction results.
[0118] Furthermore, a Gamma mixture model probability density function is defined in the latent space of the variational autoencoder; the input of the variational autoencoder is high-dimensional gene expression data X (samples × genes), and the output is the distribution parameters of the latent embedding representation variable z; the sampling of the Gamma distribution is differentiably optimized through the Gaussian distribution reparameterization technique, and the Gamma distribution parameters are optimized to adapt to the deep learning framework. The model is trained by combining the cross-entropy loss and the objective function that maximizes the weighted evidence lower bound, and the patient prognosis risk or chemotherapy sensitivity prediction result is output.
[0119] More specifically, in the latent space of the variational autoencoder, a Gamma mixture model is used as the prior distribution, and its probability density function is defined as follows:
[0120]
[0121] where z represents the latent embedding representation variable of the transcriptome data of cervical adenocarcinoma patients, which is a sample in the latent space of the variational autoencoder and represents the low-dimensional representation of the gene expression data of each patient. is the weight of the mixture component, which represents the contribution degree of each Gamma distribution component (or cluster) to the overall distribution in the Gamma mixture model. Q is a hyperparameter that specifies how many different Gamma distribution components there are in the Gamma mixture model, reflecting the number of molecular subtypes of cervical adenocarcinoma. The initial value can be set according to the known molecular classification of cervical adenocarcinoma (such as the CESC classification number Q = 4 in TCGA), or the value of Q can be optimized and selected through the elbow method; q is the sampling cluster indicator variable, that is, it represents the q-th mixture component (i.e., the cluster). and respectively represent the shape parameter and rate parameter of the q-th mixture component. These two are the Gamma distribution parameters of each mixture component and are used to flexibly model the asymmetric data distribution. Gam represents the Gamma distribution, which is a continuous probability distribution widely used for modeling positive real data, especially in cases with asymmetric distributions. According to the moment matching method, the shape parameter α and rate parameter β of the Gamma distribution are initialized, where α and β are generalized symbols, respectively representing the sets of α q and β q : , , α 0 and β 0 are the hyperparameters of the Gamma mixture model. For each mixture component q, a subset is randomly sampled from the data, and the mean and variance of the subset are calculated. The α 0 and β 0 obtained in the above way are used to initialize the prior distributions of α q and β q to ensure the differentiation of the initial distributions of each component.
[0122] Compared with the traditional Gaussian mixture model, the Gamma mixture model can better handle the non-normal distribution characteristics of transcriptome data, thus improving the clustering ability of the gene expression pattern of cervical adenocarcinoma.
[0123] S142: Construction of the optimization objective
[0124] The deep clustering prediction model is used to learn the gene expression pattern of cervical adenocarcinoma and predict the prognosis risk and chemotherapy drug sensitivity of patients based on transcriptome data. During the variational inference process, the optimization objective is to maximize the evidence lower bound loss , to balance the reconstruction error and prior regularization:
[0125]
[0126] where is the reconstruction error, which measures the matching degree between the data x generated by the decoder and the real data, equivalent to the expectation of the log-likelihood; is the Kullback-Leibler divergence, which represents the difference between the latent distribution and the prior distribution , and is used to regularize the latent representation to make it closer to the prior distribution . Since direct sampling of the Gamma distribution is non-differentiable, we use the reparameterization trick of the Gaussian distribution to ensure the trainability of the model:
[0127]
[0128] where is the mean term after parameterization of the Gamma distribution; is the standard deviation term of the Gamma distribution; is the noise variable of the standard normal distribution, which is used for reparameterized sampling to ensure computability of the gradient; represents element-wise multiplication (Hadamard product); represents the sample of the latent variable generated by the reparameterization trick of the Gaussian distribution; The two parameters α and β of the Gamma distribution are calculated as follows:
[0129]
[0130]
[0131] where is the mean output by the variational autoencoder; is the learnable parameter that controls the shape of the Gamma distribution; is the normalization term, initially set to 1.0, which is output by the normalization encoder to improve numerical stability.
[0132] This reparameterization method ensures that the variational autoencoder can update the parameters through gradient descent during training, thereby learning the latent representation of cervical adenocarcinoma data.
[0133] S143: Predicting the prognosis risk and chemotherapy sensitivity of cervical adenocarcinoma patients based on a deep clustering prediction model
[0134] After learning the latent embedding z of the transcriptome data, we use the deep clustering method to group the prognosis of cervical adenocarcinoma patients, and its objective function is as follows:
[0135]
[0136] where, is the predicted prognosis risk of the patient (such as low risk / high risk), or chemotherapy drug sensitivity (such as sensitive / resistant); is a multi-layer perceptron classifier used to map z to specific prediction labels.
[0137] To enhance the discrimination ability of the patient prognosis risk prediction, we use the weighted cross-entropy loss and the maximized evidence lower bound loss to ensure that the model achieves a balance between classification accuracy and latent variable learning:
[0138]
[0139] where, : The true label of the i-th sample, indicating the prognosis of the patient or drug sensitivity, taking values of 0 (no response) or 1 (response). : The predicted probability of the i-th sample, indicating the probability that the sample belongs to a certain category (such as low prognosis risk or high prognosis risk). N: The total number of samples. : The maximized evidence lower bound loss, used to approximate the posterior distribution of the data and improve the generalization ability of the model. : The balance weight. Stratified cross-validation can be used to optimize the loss weight, and in this embodiment, , .
[0140] This loss is used to guide the model to learn the clinical significance of different cervical adenocarcinoma gene expression patterns.
[0141] According to the above embodiments of the present invention, through clustering analysis: combining the latent class model and sub-coalition division, ensuring clustering stability through Nash stability. Prediction model: The variational autoencoder and Gamma mixture model solve the problem of asymmetric distribution of high-dimensional data and improve prediction accuracy. End-to-end process: A complete closed-loop from data preprocessing to clinical prediction, supporting precise medical decision-making.
[0142] Traditional clustering methods (such as K-Means) rely only on global distances, while sub-coalition partitioning ensures the biological consistency of sample membership (such as co-expression networks) through local similarity (cosine similarity) and Nash stability constraints. Transcriptome data often exhibits a right-skewed distribution (such as the long tail of highly expressed genes), making it difficult for Gaussian mixture models to fit. By modeling the latent space with a Gamma mixture model and combining reparameterization techniques, the model can better capture the asymmetric characteristics of gene expression, and experimental data proves the improvement of its clustering metrics (such as the adjusted Rand index).
[0143] Figure 3 A module schematic diagram of a clustering system for transcriptome sequencing data is shown. As Figure 3 shown, a clustering system 200 for transcriptome sequencing data includes:
[0144] A data acquisition and preprocessing module 210, configured to collect cervical adenocarcinoma and adjacent tissue specimens of several patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix;
[0145] A clustering analysis module 220, configured to perform clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model and sub-coalition partitioning, model sample gene expression data through factor analysis and a latent class model, and combine sub-coalition partitioning with Nash stability constraints to maximize gene expression similarity and the joint distribution of latent classes, thereby obtaining a clustered gene set;
[0146] An association analysis module 230, configured to perform gene function annotation on the clustered gene set, and perform association analysis on the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;
[0147] A deep clustering prediction module 240, configured to construct a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model based on the clustering analysis and association analysis results to predict the prognosis risk or chemotherapy sensitivity of patients.
[0148] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0149] Clustering performance evaluation:
[0150] In the clustering performance evaluation experiment, the present invention compared the performance of the K-Means clustering algorithm, Gaussian mixture model, deep embedded clustering, improved deep embedded clustering, information generation adversarial network, and the algorithm proposed by the present invention in various clustering metrics. The specific analysis is as follows:
[0151] Figure 4 is a comparison chart of the adjusted Rand index performance of different models. From Figure 4 it can be seen that the proposed algorithm shows high clustering performance during the training process. Its adjusted Rand index is always higher than other comparison methods, converges faster, and finally approaches above 0.8, which is better than all comparison algorithms. Since the proposed algorithm uses the Gamma mixture model as the prior distribution and combines the variational autoencoder for feature learning, it has a stronger modeling ability for complex cervical adenocarcinoma transcriptome data, thus obtaining a better clustering effect.
[0152] Figure 5 is a comparison chart of the Davies-Bouldin index performance of different models. The Davies-Bouldin index is used to measure the quality of clustering. The smaller its value, the better the clustering effect, which means that the samples within the cluster are closer and the separation between clusters is more obvious. As Figure 5 shown, it can be seen that the traditional methods (K-Means clustering algorithm, Gaussian mixture model) perform the worst, with a slow decrease in the Davies-Bouldin index and low clustering tightness. The deep learning methods (deep embedded clustering, improved deep embedded clustering, information generation adversarial network) significantly reduce the Davies-Bouldin index, indicating that they can better mine the deep structure of the data. The proposed algorithm in the present invention has the fastest decrease in the Davies-Bouldin index and finally reaches 0.95 after 50 rounds of training, which is 48% lower than the K-Means clustering algorithm and 13.6% lower than the information generation adversarial network, verifying its superiority in the deep clustering task and proving that this method can better optimize the intra-class tightness and inter-class separation. It shows the ability of the proposed algorithm to combine the variational autoencoder and the Gamma mixture model, which is more advantageous in learning the clustering structure.
[0153] Figure 6 is a comparison chart of the normalized mutual information performance of different models. In Figure 6 the experiment, the present invention compares the performance of each model in terms of the normalized mutual information index. The proposed algorithm achieves the best performance. Finally, the deep embedded clustering reaches 0.85, which is 8.9% higher than the information generation adversarial network and nearly 88.9% higher than K-Means, indicating that the proposed algorithm can effectively learn complex data distributions, improve clustering quality, and demonstrate the optimal clustering effect.
[0154] Furthermore, the embodiment of the present application also provides an electronic device, including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any one of the above methods.
[0155] Furthermore, the embodiment of the present application also provides a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute any of the above methods.
[0156] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0157] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0158] It should also be noted that in the embodiments of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0159] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the embodiments of the present application, but will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.
Claims
1. A clustering method for transcriptome sequencing data, characterized in that: include: Collecting cervical adenocarcinoma and adjacent tissue specimens from several patients, performing transcriptome sequencing on the specimens and preprocessing the transcriptome sequencing data to obtain a standardized gene expression matrix; A clustering operation is performed on the standardized gene expression matrix based on a clustering analysis method combining a latent class model with sub-coalition partitioning, the sample gene expression data is modeled by factor analysis and a latent class model, and the sub-coalition partitioning with Nash stability constraints is combined to maximize the gene expression similarity and the joint distribution of the latent classes, thereby obtaining a clustered gene set; The clustered gene set was annotated for gene function, and the clustering results after functional annotation were associated with the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics; Based on the results of cluster analysis and association analysis, a deep clustering prediction model based on variational autoencoder and Gamma mixture model was constructed to predict patients' prognostic risk and chemotherapy sensitivity.
2. A clustering method for transcriptome sequencing data according to claim 1, characterized in that: in, The clustering operation of performing clustering operation on the standardized gene expression matrix based on the clustering analysis method combining the latent class model with the sub-coalition division includes: Using factor analysis to model gene expression data: Among them, X is the gene expression data, Λ is the factor loading matrix, ε is the error term, is the mixing coefficient of class m, is the parameter of class m; represents the true label of the i-th sample; M is the total number of categories; The gene expression similarity between samples is calculated based on the cosine similarity metric function, and the sub-alliances are divided according to the preset threshold δ; In the process of sub-alliance division, each sample i belongs to the kth sub-alliance , its individual utility Defined as ,in, Represents the characteristics of sample i and sample j; represents the kth sub-alliance; is the cosine similarity measurement function between sample i and sample j; : The individual utility of sample i, i.e., the sum of its similarities with all other samples in the same sub-alliance; The Nash stability constraint is used to ensure that the individual utility of the sample in the sub-alliance is not lower than the individual utility after migrating to other sub-alliances; The optimization objective function includes gene expression similarity and latent class joint distribution.
3. A clustering method for transcriptome sequencing data according to claim 2, characterized in that: in, The Nash stability condition of the sub-coalition partition is: That is, for any sample i, the individual utility of its current sub-alliance is ,in, is the sub-alliance to which sample i currently belongs Individual utility in If sample i migrates to other sub-alliances The subsequent individual utility.
4. A clustering method for transcriptome sequencing data according to claim 3, characterized in that: in, The clustering results after functional annotation were analyzed in association with the clinical characteristics of cervical adenocarcinoma, and the correlation between gene expression and clinical characteristics was obtained, including: Pearson correlation coefficient was used to analyze the linear relationship between gene expression and clinical characteristics; The differences in gene expression between different clinical groups were compared by t test.
5. A clustering method for transcriptome sequencing data according to claim 4, characterized in that: in, Based on the results of cluster analysis and association analysis, a deep clustering prediction model based on variational autoencoder and Gamma mixture model was constructed to predict the patient's prognosis risk and chemotherapy sensitivity, including: Variational autoencoders are used to reduce the dimensionality of high-dimensional gene expression data to extract latent features, and a Gamma mixture model is used to model asymmetric distribution in the latent space, cluster the latent space, and identify molecular subgroups with different prognostic risks; The Gamma distribution parameters were optimized through reparameterization techniques, and the model was trained by combining cross entropy loss with the objective function of maximizing the evidence lower bound (ELBO) weighted to output the patient's prognostic risk and chemotherapy sensitivity prediction results.
6. A clustering method for transcriptome sequencing data according to claim 5, characterized in that: in, The combination of variational autoencoder and Gamma mixture model includes: A Gamma mixture model probability density function is defined in the latent space of a variational autoencoder; the input of the variational autoencoder is high-dimensional gene expression data X, and the output is a distribution parameter of a latent embedding representation variable z; The Gamma distribution sampling is optimized differently through the Gaussian distribution reparameterization technique, and the Gamma distribution parameters are optimized to adapt to the deep learning framework. The model is trained by combining the cross entropy loss and the objective function of maximizing the weighted lower bound of evidence to output the patient's prognostic risk and chemotherapy sensitivity prediction results.
7. A clustering method for transcriptome sequencing data according to claim 6, characterized in that: in, The Gamma distribution sampling is optimized by Gaussian distribution reparameterization technique to optimize the Gamma distribution parameters, including the following calculation formula: in, is the parameterized mean term of the Gamma distribution; is the standard deviation term of the Gamma distribution; is a noise variable with standard normal distribution; represents element-wise multiplication; represents the latent variable samples generated by the reparameterization technique of Gaussian distribution; Calculate the Gamma distribution parameters using the parameter conversion formula: in, is the mean of the variational autoencoder output; is a learnable parameter that controls the shape of the Gamma distribution; is the normalization term; They represent the shape parameter and rate parameter of the Gamma distribution respectively.
8. A clustering method for transcriptome sequencing data according to claim 7, characterized in that: in, The decoder output of the variational autoencoder is mapped to the patient prognostic risk and chemotherapy sensitivity prediction label through a multi-layer perceptron (MLP): in, It is the predicted patient prognostic risk or chemotherapy drug sensitivity; is a multi-layer perceptron classifier used to map z to a specific prediction label.
9. A clustering system for transcriptome sequencing data, characterized in that: include: A data collection and preprocessing module, used to collect cervical adenocarcinoma and adjacent tissue specimens from several patients, perform transcriptome sequencing on the specimens and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix; A clustering analysis module, which is used to perform clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model with sub-coalition division, model the sample gene expression data through factor analysis and latent class model, and combine the sub-coalition division with Nash stability constraints to maximize the gene expression similarity and the joint distribution of latent classes, so as to obtain a clustered gene set; The association analysis module is used to annotate the gene functions of the clustered gene set, and to perform association analysis between the clustering results after functional annotation and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics; The deep clustering prediction module is used to build a deep clustering prediction model based on variational autoencoder and Gamma mixture model based on the results of clustering analysis and association analysis to predict patient prognosis risk and chemotherapy sensitivity.
10. An electronic device comprising: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, and the one or more programs include instructions, characterized in that when the instructions are executed by the processor, the processor executes the steps of a clustering method for transcriptome sequencing data as described in any one of claims 1-8.
Citation Information
Patent Citations
Node cooperative caching method based on cooperative game and deep learning
CN115915277A
Cited By
Esophageal cancer immune data prediction device and equipment based on tumor markers and storage medium
CN120932727A
Esophageal cancer immune data prediction device based on tumor markers, equipment and storage medium
CN120932727B
Tuberculous granuloma and lung cancer transcriptome liquid biopsy model construction method and system based on multi-omics sequencing
CN122428038A