A Clustering Method and System for Transcriptome Sequencing Data

By combining the cluster analysis method of potential category models and sub-consortium division and the deep cluster prediction model of the mixed model of the variant autoencoder and Gamma, the cluster stability and prediction accuracy of transcriptome data are solved, and accurate prediction of prognostic risk and chemotherapy sensitivity in cervical adenocarcinoma patients is achieved.

CN120126581BActive Publication Date: 2025-07-18THE AFFILIATED HOSPITAL OF SOUTHWEST MEDICAL UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510619545.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing cluster analysis methods have problems with poor clustering stability and low prediction accuracy when processing transcriptomic data. Especially in cervical adenocarcinoma studies, it is difficult to effectively explore deep features and potential relationships in the data, resulting in poor predicting patient prognosis risks and chemotherapy drugs sensitivity.

Method used

A cluster analysis method combining latent category model and sub-consortium division is adopted, combining factor analysis and Nash stability constraints, and a deep cluster prediction model is constructed through a variational autoencoder and Gamma hybrid model, and the objective function is optimized to improve the stability and prediction ability of clustering results.

Benefits of technology

The clustering stability and prediction accuracy of transcriptomic data were improved, the AUC predicted by patients with prognostic risk reached 0.85, and the F1 score of chemotherapy sensitivity increased to 0.79, significantly better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126581B_ABST
    Figure CN120126581B_ABST
Patent Text Reader

Abstract

The present invention provides a clustering method and system for transcriptome sequencing data. Applied to the technical field of data processing, the method includes: collecting cervical adenocarcinoma and adjacent tissue specimens of a number of patients, and performing data preprocessing to obtain a standardized gene expression matrix; performing clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model and sub-coalition partitioning to obtain a clustered gene set; performing gene function annotation on the clustered gene set, and performing correlation analysis between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics; constructing a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model based on the clustering analysis and correlation analysis results to predict the prognosis risk or chemotherapy sensitivity of patients. The present invention solves the problems of poor clustering stability of high-dimensional transcriptome data and low prognostic prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a clustering method and system for transcriptome sequencing data. Background Art

[0002] Transcriptome sequencing technology can comprehensively and rapidly obtain information on the transcripts of an organism, and plays a key role in the field of tumor research, especially in the research of cervical adenocarcinoma. Through transcriptome sequencing, a large amount of gene expression data related to cervical adenocarcinoma can be obtained. These data contain rich biological information and are of great significance for revealing the molecular mechanism of cervical adenocarcinoma, exploring potential therapeutic targets, predicting the prognosis risk of patients, and the sensitivity of chemotherapy drugs.

[0003] However, the raw data generated by transcriptome sequencing has many problems, such as containing low-quality bases, adapter sequences, and contaminated data, and there is a lack of comparability between different sample data, which seriously affects the accuracy and reliability of subsequent data analysis. At the same time, how to effectively mine valuable information from the vast amount of gene expression data, accurately classify samples or genes, and discover potential data structures and patterns has always been a challenge in the field of bioinformatics.

[0004] Traditional clustering analysis methods, such as the K-Means clustering algorithm and the Gaussian mixture model, have limitations in processing complex high-dimensional transcriptome data, and it is difficult to fully mine the deep features and potential relationships in the data, resulting in poor clustering effects. In predicting the prognosis risk and chemotherapy drug sensitivity of cervical adenocarcinoma patients, existing models and methods also cannot well adapt to the characteristics of transcriptome data, resulting in low prediction accuracy. Therefore, there is an urgent need for a new method to solve these problems to promote the in-depth development of cervical adenocarcinoma research and provide more powerful support for clinical diagnosis and treatment. Summary of the Invention

[0005] In view of the deficiencies of traditional clustering methods in processing transcriptome data, the present invention proposes a clustering method and system for transcriptome sequencing data.

[0006] To achieve the above object, it is realized through the following technical solutions:

[0007] According to the first aspect of the present invention, a clustering method for transcriptome sequencing data is provided, including the following steps:

[0008] Collect cervical adenocarcinoma and adjacent normal tissue specimens from several patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix;

[0009] A clustering analysis method combining a latent class model and sub - coalition partitioning is used to perform clustering operations on the standardized gene expression matrix. By factor analysis and latent class model to model the sample gene expression data, combined with sub - coalition partitioning with Nash stability constraints, the gene expression similarity and the joint distribution of latent classes are maximized to obtain the clustered gene set;

[0010] Gene function annotation is performed on the clustered gene set, and the correlation analysis is carried out between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;

[0011] Based on the results of clustering analysis and correlation analysis, a deep clustering prediction model based on variational auto - encoder and Gamma mixture model is constructed to predict the prognosis risk and chemotherapy sensitivity of patients.

[0012] Furthermore, the clustering operation of the standardized gene expression matrix by the clustering analysis method combining the latent class model and sub - coalition partitioning includes:

[0013] Construct a gene expression data model using factor analysis:

[0014]

[0015] where \(X\) is the gene expression data, \(\Lambda\) is the factor loading matrix, \(\varepsilon\) is the error term, is the mixing coefficient of class \(m\), is the parameter of class \(m\); represents the true label of the \(i\) - th sample; \(M\) is the total number of classes;

[0016] Calculate the gene expression similarity between samples based on the cosine similarity metric function, and divide sub - coalitions according to the preset threshold \(\delta\);

[0017] During the sub - coalition partitioning process, each sample \(i\) belongs to the \(k\) - th sub - coalition , and its individual utility is defined as , where, represents the features of sample \(i\) and sample \(j\); represents the \(k\) - th sub - coalition; is the cosine similarity metric function between sample \(i\) and sample \(j\); : the individual utility of sample \(i\), that is, the sum of its similarities with all other samples in the same sub - coalition;

[0018] Ensure that the individual utility of samples in the sub - coalition is not lower than the individual utility after migrating to other sub - coalitions through Nash stability constraints;

[0019] The optimization objective function includes gene expression similarity and the joint distribution of latent classes.

[0020] Furthermore, the Nash stability condition for the sub - coalition division is as follows:

[0021]

[0022] That is, for any sample i, the individual utility of its current sub - coalition , where is the individual utility of sample i in its current affiliated sub - coalition ; is the individual utility of sample i if it migrates to another sub - coalition .

[0023] Furthermore, where the clustering results after functional annotation are correlated with the clinical features of cervical adenocarcinoma, and the correlations between gene expression and clinical features obtained include:

[0024] Using the Pearson correlation coefficient to analyze the linear relationship between gene expression and clinical features;

[0025] Comparing the gene expression differences between different clinical groups through t - test.

[0026] Furthermore, where based on the results of clustering analysis and correlation analysis, a deep clustering prediction model based on variational auto - encoder and Gamma mixture model is constructed to predict the prognosis risk and chemotherapy sensitivity of patients, including:

[0027] Using the variational auto - encoder to reduce the dimensionality of high - dimensional gene expression data to extract potential features, and modeling the asymmetric distribution with the Gamma mixture model in the latent space, clustering the latent space to identify molecular subgroups with different prognosis risks;

[0028] Optimizing the Gamma distribution parameters through the reparameterization trick, training the model with the objective function weighted by cross - entropy loss and maximizing the evidence lower bound (ELBO), and outputting the prediction results of the patient's prognosis risk and chemotherapy sensitivity.

[0029] Furthermore, the combination of the variational auto - encoder and the Gamma mixture model includes:

[0030] Defining the probability density function of the Gamma mixture model in the latent space of the variational auto - encoder; the input of the variational auto - encoder is high - dimensional gene expression data X, and the output is the distribution parameters of the latent embedding representation variable z;

[0031] Performing differentiable optimization on the Gamma distribution sampling through the Gaussian distribution reparameterization trick, optimizing the Gamma distribution parameters to adapt to the deep learning framework, training the model with the objective function weighted by cross - entropy loss and maximizing the evidence lower bound, and outputting the prediction results of the patient's prognosis risk and chemotherapy sensitivity.

[0032] Further, the sampling of the Gamma distribution is differentiably optimized by the Gaussian distribution reparameterization technique to optimize the Gamma distribution parameters, including the following calculation formulas:

[0033]

[0034] where, is the mean term after parameterization of the Gamma distribution; is the standard deviation term of the Gamma distribution; is the noise variable of the standard normal distribution; represents element-wise multiplication; represents the latent variable sample generated by the reparameterization technique of the Gaussian distribution;

[0035] Calculate the Gamma distribution parameters through the parameter conversion formula:

[0036]

[0037]

[0038] where, is the mean output by the variational autoencoder; is the learnable parameter that controls the shape of the Gamma distribution; is the normalization term; respectively represent the shape parameter and rate parameter of the Gamma distribution.

[0039] Further, the decoder output of the variational autoencoder is mapped to the patient prognosis risk or chemotherapy sensitivity prediction label through a multi-layer perceptron (MLP):

[0040]

[0041] where, is the predicted patient prognosis risk or chemotherapy drug sensitivity; is the multi-layer perceptron classifier for mapping z to the specific prediction label.

[0042] According to the second aspect of the present invention, there is also provided a clustering system for transcriptome sequencing data, including:

[0043] A data acquisition and preprocessing module for collecting cervical adenocarcinoma and adjacent tissue specimens of several patients, performing transcriptome sequencing on the specimens, and preprocessing the transcriptome sequencing data to obtain a standardized gene expression matrix;

[0044] A clustering analysis module, which is used to perform clustering operations on the standardized gene expression matrix based on a clustering analysis method that combines a latent class model and sub - coalition partitioning. By factor analysis and latent class model to model the sample gene expression data, combined with sub - coalition partitioning with Nash stability constraints, it maximizes gene expression similarity and latent class joint distribution to obtain the clustered gene set;

[0045] An association analysis module, which is used to perform gene function annotation on the clustered gene set, and perform association analysis on the clustered results after function annotation and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;

[0046] A deep clustering prediction module, which is used to construct a deep clustering prediction model based on a variational auto - encoder and a Gamma mixture model based on the results of clustering analysis and association analysis to predict the prognosis risk and chemotherapy sensitivity of patients.

[0047] According to the third aspect of the present invention, there is also provided an electronic device including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the steps of a clustering method for transcriptome sequencing data described in the first aspect.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention integrates a clustering method of a latent class model and sub - coalition partitioning. By factor analysis to model the global distribution characteristics of gene expression, combined with sub - coalition partitioning based on cosine similarity to refine the local structure, and introducing Nash stability to constrain sample attribution, it solves the problem of poor clustering stability of high - dimensional transcriptome data.

[0050] 2. The present invention designs a deep joint model of a variational auto - encoder (VAE) and a Gamma mixture model. The variational auto - encoder is used for high - dimensional data dimensionality reduction to learn latent expression patterns, and the Gamma mixture model is used to model the asymmetric distribution in the latent space to make up for the deficiency of the traditional Gaussian mixture model in dealing with asymmetric data distributions. By optimizing the Gamma distribution parameters, it enhances the model's ability to model gene expression data, makes the clustering results more stable, and improves the prediction ability of the prognosis risk and chemotherapy drug sensitivity of patients, overcoming the defect of poor adaptability of the traditional Gaussian mixture model to skewed data.

[0051] 3. The present invention constructs an end-to-end prognosis and chemotherapy sensitivity prediction framework. After associating the clustering results with clinical features and inputting them into a deep model, weighted optimization is performed through cross-entropy loss and maximized evidence lower bound loss. This loss function helps improve the clustering and prediction capabilities of the model on cervical adenocarcinoma gene expression data, achieving a prognostic risk prediction AUC of 0.85 (0.68 for traditional logistic regression) and boosting the F1 score of chemotherapy sensitivity to 0.79.

[0052] It should be understood that the content described in the summary of the invention is not intended to define the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0054] Figure 1 is a schematic diagram of the specific steps of a clustering method for transcriptome sequencing data according to an embodiment of the present invention;

[0055] Figure 2 is a schematic flowchart of a clustering method for transcriptome sequencing data according to an embodiment of the present invention;

[0056] Figure 3 is a schematic diagram of the modules of a clustering system for transcriptome sequencing data according to an embodiment of the present invention;

[0057] Figure 4 is a performance comparison chart of the adjusted Rand index for different models;

[0058] Figure 5 is a performance comparison chart of the Davies-Bouldin index for different models;

[0059] Figure 6 is a performance comparison chart of the normalized mutual information for different models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0061] In addition, the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0062] Prognosis: Refers to the comprehensive evaluation of the future course of the disease, treatment effect, survival period, recurrence risk, etc. of a patient after the disease is diagnosed.

[0063] Prognosis prediction: Quantify the potential outcomes of patients through statistical models or machine learning methods, combined with multi-dimensional data (such as genes, pathology, clinical indicators).

[0064] Figure 1 The figure shows a schematic diagram of the specific steps of a clustering method for transcriptome sequencing data. Figure 2 The figure shows a schematic flow diagram of a clustering method for transcriptome sequencing data. As Figure 1 and Figure 2 shown, a clustering method 100 for transcriptome sequencing data includes the following steps:

[0065] S110: Collect cervical adenocarcinoma and adjacent tissue specimens from a number of patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix.

[0066] Data collection: From female cervical adenocarcinoma patients aged 21 - 75 in the research project "Exploring the Molecular Mechanisms and Potential Therapeutic Targets of Cervical Adenocarcinoma Based on Common Transcriptome Sequencing Data", collect 8 cervical adenocarcinoma and adjacent tissue specimens (8 specimens each for cancer and adjacent tissues) from 8 patients. The specimens are collected for subsequent experiments after obtaining the informed consent of the patients. Perform transcriptome sequencing on the collected specimens to obtain the original transcriptome sequencing data (in FASTQ format), providing a basis for subsequent analysis.

[0067] Preprocessing of the original transcriptome sequencing data: Use professional software tools to perform quality control on the original sequencing data, remove low-quality bases, adapter sequences, and contaminated data to improve data quality. Perform standardization processing on the data after quality control to correct the differences between the data and ensure the comparability of data from different samples.

[0068] After the above preprocessing, a high-quality and standardized gene expression matrix (matrix dimension: sample × gene, with the value being the expression level) is obtained.

[0069] S120: performing a clustering operation on the standardized gene expression matrix based on a clustering analysis method combining a latent class model with sub-coalition division, modeling sample gene expression data through factor analysis and latent class model, maximizing gene expression similarity and latent class joint distribution in combination with Nash stability constraint sub-coalition division, and obtaining a clustered gene set;

[0070] This step S120 is based on the clustering analysis of latent classes and sub-alliances to obtain the sub-alliances division results after clustering (the sub-alliances labels to which each sample belongs) and the latent class model parameters (to explain the latent structure of gene expression data). The latent class model and sub-alliances complement and cooperate with each other. The latent class model captures the global distribution, and the sub-alliances refine the local structure, thereby overcoming the sparsity of high-dimensional data.

[0071] Among them, sub-alliances are sample subsets divided based on gene expression similarity. Samples within the same sub-alliances have higher homogeneity in biological functions (such as gene co-expression and clinical phenotype), that is, a collection of samples with highly similar gene expression patterns, reflecting the molecular subtypes of cervical adenocarcinoma (such as immune infiltration type and metabolically active type). Difference from traditional clustering: Sub-alliances emphasize local similarities between samples (such as gene co-expression networks) rather than global distances (such as the Euclidean distance of K-Means).

[0072] The preprocessed data was input into the selected algorithm, and cluster analysis was performed on the standardized gene expression matrix and clinical feature data (such as tumor stage, prognostic label, chemotherapy sensitivity label) according to the clinical characteristics and gene expression patterns of cervical adenocarcinoma. The clustering analysis method that combines latent classes + sub-alliances was selected for clustering operations, and samples or genes were classified according to the similarity of gene expression to mine potential data structures and patterns. Among them, gene expression data It can be expressed using factor analysis and latent class models:

[0073]

[0074] Among them, X is the observed variable (gene expression data), Λ is the factor loading matrix, ε is the error term, is the mixing coefficient of class m, Yes Category Parameters; represents the true label of the i-th sample; M is the total number of categories; It is a latent class model, which aims to explain the joint distribution of indicators by determining latent classes under the premise that different latent classes present different distributions of indicators.

[0075] The goal of the sub-alliance is to divide multiple sub-alliances based on the similarity of gene expression Making the samples in the same sub - coalition more similar in gene expression and obtaining a social weak order by aggregating individual preferences.

[0076] Define the gene expression similarity measure:

[0077]

[0078] Where, is the gene expression vector of the \(i\) - th sample; : the gene expression vector of the \(j\) - th sample; : measures the similarity of the gene expression patterns of sample \(i\) and sample \(j\), calculated using cosine similarity. A value close to 1 indicates that the gene expressions of the two samples are very similar, close to 0 indicates no obvious correlation in gene expression, and close to - 1 indicates that the gene expression patterns are exactly opposite.

[0079] Here, The cosine similarity is used to calculate the similarity of the gene expression pattern. The closer the value is to 1, the more similar the gene expressions of sample \(i\) and sample \(j\) are. To screen for suitable sub - coalitions, the present invention sets a similarity threshold \(\delta\). If , then sample \(i\) and sample \(j\) belong to the same sub - coalition, indicating that their gene expression patterns are similar enough. If , then sample \(i\) and sample \(j\) do not belong to the same sub - coalition, indicating that there are significant differences in their gene expressions.

[0080] Among them, in the present invention, \(\delta\) is optimized by grid search combined with survival analysis: test \(\delta\in[0.7,0.9]\) with an interval of 0.05 on the training set. Perform sub - coalition division for each \(\delta\) value and calculate the difference in survival curves (Log - rank test \(P\) - value) of each sub - coalition. Select the \(\delta\) value that minimizes the \(P\) - value (i.e., the most significant subtype prognosis differentiation), for example, take \(\delta = 0.85\).

[0081] During the sub - coalition division process, each sample \(i\) belongs to the \(k\) - th sub - coalition , and its individual utility can be defined as:

[0082]

[0083] Where, : represents the features of sample \(i\) and sample \(j\) (such as gene expression data); : represents the \(k\) - th sub - coalition; : the cosine similarity measure function between sample \(i\) and sample \(j\); : the individual utility of sample \(i\), that is, the sum of its similarities with all other samples within the same sub - coalition.

[0084] That is, the individual utility of sample \(i\) depends on the sum of its similarities with other samples within the same sub - coalition.

[0085] The goal of sub - coalition partitioning is to maximize the overall gene expression similarity:

[0086]

[0087] Where \(W(C)\): the overall utility function, representing the sum of the individual utilities of all samples in all sub - coalitions. The goal is to maximize this value to obtain a more compact sub - coalition partitioning. \(K\): the total number of sub - coalitions. At the same time, Nash stability needs to be satisfied, that is:

[0088]

[0089] : the individual utility of sample \(i\) in its current affiliated sub - coalition ; : the individual utility of sample \(i\) if it migrates to sub - coalition .

[0090] This condition requires that the individual utility of each sample \(i\) in its own sub - coalition is not lower than its individual utility after migrating to other sub - coalitions, that is, an individual will not increase its utility by joining other sub - coalitions. Otherwise, it will migrate to a new sub - coalition until a stable state is reached.

[0091] The Nash stability constraint requires that the individual utility of a sample in its current sub - coalition is not lower than its individual utility after migration, thus preventing samples from frequently switching during the clustering process, ensuring that the membership of samples in sub - coalitions will not change due to minor perturbations (such as noise or measurement errors), and improving the robustness of the clustering results. In cervical adenocarcinoma data, this constraint can improve the repeatability of subtype classification (e.g., multiple test results of the same patient are stably classified into the same sub - coalition), enhance the reliability of clinical applications, be able to avoid samples from frequently switching sub - coalitions during the iteration process, ensure the stability of patient classification, and support clinical decisions (such as chemotherapy regimen selection).

[0092] Finally, the objective function for optimizing the clustering analysis method of latent class + sub - coalition includes gene expression similarity and the joint distribution of latent classes. The objective function \(C\) * is:

[0093]

[0094] Where \(W(C)\) guides sub - coalition partitioning by gene expression similarity measure; \(C\) * : the final optimal sub - coalition partitioning scheme; \(\lambda\): the balance parameter, controlling the weights of gene expression similarity and the part of calculating class probabilities in the class model; : the prior probability of the \(k\) - th sub - coalition; \(K\): the total number of sub - coalitions; \(N\): the total number of samples; : The conditional probability that sample i belongs to this class (i.e., belongs to the k-th sub-coalition) given the parameters of the k-th sub-coalition ; : The class probabilities calculated by the latent class model, which measure the likelihood of a sample belonging to different classes.

[0095] By optimizing the weight λ, optimizing the class parameters and the mixing coefficients to obtain the best classification of cervical adenocarcinoma gene expression.

[0096] This step S120 is based on sub-coalition partitioning. Samples with similar gene expression patterns are grouped into the same sub-coalition through the similarity threshold δ, and the co-expressed gene set within the sub-coalition is extracted; based on the latent class model, genes that contribute significantly to the class (high-loading genes) are identified through the factor loading matrix Λ to form a clustering gene set. Finally, a gene subset highly related to cervical adenocarcinoma (such as a differentially expressed gene cluster) can be identified through cluster analysis. These gene sets can be used as input features of the model to reduce the data dimension and filter out noise.

[0097] For example, assume that a cervical adenocarcinoma data set contains two types of patients. Sub-coalition partitioning: The patients are divided into sub-coalitions A and B. The gene expressions of patients within A are highly similar (such as the activation of the EGFR pathway), and the similarity of patients within B is lower. Latent class model: It is found that the overall data can be divided into two types (such as "immune activation type" and "metabolic abnormality type"), and each type corresponds to a different survival prognosis. Joint optimization result: Patients in sub-coalition A may mainly belong to latent class 1 (immune activation type); patients in sub-coalition B may belong to latent class 2 (metabolic abnormality type). The final clustering result can reflect both local similarity and align with the global molecular subtypes.

[0098] S130: Perform gene function annotation on the clustered gene set, and conduct correlation analysis between the functionally annotated clustering result and the clinical features of cervical adenocarcinoma to obtain the correlation between gene expression and clinical features;

[0099] This step S130 is used to mine the association between gene expression patterns and clinical features through correlation analysis:

[0100] Perform functional annotation on the gene sets obtained by clustering in step S120. With the help of databases such as Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), clarify the biological processes, molecular functions, and signaling pathways involved by the genes, and deeply understand the functional significance of different clustered genes. If the genes are sorted according to the expression level, gene set enrichment analysis can be used for gene functional annotation. Perform correlation analysis on the clustering results and the clinical characteristics of cervical adenocarcinoma (such as tumor stage, patient prognosis risk, chemotherapy drug sensitivity, etc.), mine the potential connections between gene expression patterns and clinical characteristics, and provide guidance for clinical diagnosis and treatment.

[0101] To explore the correlation between gene expression and clinical characteristics (such as tumor stage, prognosis, chemotherapy sensitivity), the present invention can calculate the Pearson correlation coefficient (measuring the linear relationship between gene expression and the clinical variable Y):

[0102]

[0103] where : the gene expression value of the i-th sample; : the clinical variable corresponding to the i-th sample (such as tumor stage, prognosis, chemotherapy sensitivity, etc.); : the average gene expression of all samples; : the average clinical variable of all samples; n: the number of samples, that is, the number of patients.

[0104] If you want to compare the differential expression levels of genes in two groups (such as chemotherapy-sensitive and drug-resistant), a t-test can be used:

[0105]

[0106] where respectively represent the average gene expression of the two groups of samples; respectively represent the standard deviations of the two groups of samples, that is, the degree of variation of the gene expression level; respectively represent the numbers of the two groups of samples.

[0107] This step 130 finally obtains the correlation between gene expression and clinical characteristics (such as gene A is significantly correlated with chemotherapy sensitivity) through methods of Pearson correlation coefficient (linear association) and t-test (difference between groups) based on the gene expression matrix and clinical characteristic data (such as stage, prognosis).

[0108] S140: Based on the results of clustering analysis and correlation analysis, construct a deep clustering prediction model based on variational autoencoder and Gamma mixture model to predict the prognosis risk and chemotherapy sensitivity of patients.

[0109] This step S140 is used to predict the prognosis risk and chemotherapy drug sensitivity of patients with cervical adenocarcinoma. Specifically, it includes:

[0110] S141: Construct a deep clustering prediction model

[0111] Based on the clustering analysis results and association analysis results of steps S120 and S130, construct a deep clustering prediction model based on variational autoencoder and Gamma mixture model, that is, construct a model for predicting the prognosis risk or chemotherapy drug sensitivity of patients with cervical adenocarcinoma.

[0112] The model structure includes:

[0113] Variational autoencoder (VAE), including an encoder: reducing high-dimensional gene expression data to a latent embedding representation variable z; a decoder: reconstructing gene expression data from the latent space.

[0114] Gamma mixture model (GMM): Cluster samples in the latent embedding representation variable z and model the asymmetric distribution.

[0115] Input the standardized gene expression matrix (preprocessed high-dimensional data) and clinical labels (such as prognosis labels: high risk / low risk) into the model, use the reparameterization trick (introducing Gaussian noise) to solve the non-differentiable problem of Gamma distribution sampling and the loss function: cross-entropy loss (classification) + ELBO loss (reconstruction and distribution matching), and finally output the latent embedding representation variable z (low-dimensional representation for clustering), as well as the patient prognosis risk prediction result (such as high risk / low risk probability) and chemotherapy sensitivity prediction result (such as sensitive / resistant probability).

[0116] The latent space of the variational autoencoder (VAE) can be designed to capture the gene expression patterns after clustering. For example, during the training process, the encoder of the VAE can be guided by the clustering results, so that the latent embedding representation variable z implicitly contains the structure information of sub-coalition division or latent classes.

[0117] Among them, this deep clustering prediction model uses a variational autoencoder to reduce the dimension of high-dimensional gene expression data to extract latent features, models the asymmetric distribution with a Gamma mixture model in the latent space, clusters the latent space, and identifies molecular subgroups with different prognosis risks; optimizes the Gamma distribution parameters through the reparameterization trick, combines the cross-entropy loss with the objective function weighted by maximizing the evidence lower bound (ELBO) to train the model, and outputs the patient prognosis risk and chemotherapy sensitivity prediction results.

[0118] Furthermore, a Gamma mixture model probability density function is defined in the latent space of the variational autoencoder; the input of the variational autoencoder is high-dimensional gene expression data X (samples × genes), and the output is the distribution parameters of the latent embedding representation variable z; the Gamma distribution sampling is differentiably optimized through the Gaussian distribution reparameterization technique, the Gamma distribution parameters are optimized to adapt to the deep learning framework, and the model is trained by combining the cross-entropy loss and the objective function of maximizing the evidence lower bound weighted, and the patient prognosis risk or chemotherapy sensitivity prediction result is output.

[0119] More specifically, in the latent space of the variational autoencoder, a Gamma mixture model is used as the prior distribution, and its probability density function is defined as follows:

[0120]

[0121] where z represents the latent embedding representation variable of the transcriptome data of cervical adenocarcinoma patients, which is a sample in the latent space of the variational autoencoder and represents the low-dimensional representation of the gene expression data of each patient. is the weight of the mixture component, which represents the contribution degree of each Gamma distribution component (or cluster) to the overall distribution in the Gamma mixture model. Q is a hyperparameter that specifies how many different Gamma distribution components there are in the Gamma mixture model, reflecting the number of molecular subtypes of cervical adenocarcinoma. The initial value can be set according to the known cervical adenocarcinoma molecular typing (such as the CESC typing number Q = 4 of TCGA), or the value of Q can be optimized and selected through the elbow method; q is the sampling cluster indicator variable, that is, it represents the q-th mixture component (i.e., the cluster); and respectively represent the shape parameter and rate parameter of the q-th mixture component. These two are the Gamma distribution parameters of each mixture component and are used to flexibly model the asymmetric data distribution. Gam represents the Gamma distribution, which is a continuous probability distribution widely used to model positive real number data, especially in cases with asymmetric distributions. The shape parameter α and rate parameter β of the Gamma distribution are initialized according to the moment matching method, where α and β are generalized symbols representing the sets of α q and β q respectively: , , α0 and β0 are the hyperparameters of the Gamma mixture model. For each mixture component q, a subset is randomly sampled from the data, the mean and variance of the subset are calculated, and α0 and β0 obtained in the above manner are used to initialize the prior distributions of α q and β q to ensure the differentiation of the initial distributions of each component.

[0122] Compared with the traditional Gaussian mixture model, the Gamma mixture model can better handle the non-normal distribution characteristics of transcriptome data, thus improving the clustering ability of the gene expression patterns of cervical adenocarcinoma.

[0123] S142: Construction of the optimization objective

[0124] The deep clustering prediction model is used to learn the gene expression patterns of cervical adenocarcinoma and predict the prognosis risk and chemotherapy drug sensitivity of patients based on transcriptome data. During the variational inference process, the optimization objective is to maximize the evidence lower bound loss , to balance the reconstruction error and the prior regularization:

[0125]

[0126] where is the reconstruction error, which measures the matching degree between the data x generated by the decoder and the real data, equivalent to the expectation of the log-likelihood; is the Kullback-Leibler divergence, representing the difference between the latent distribution and the prior distribution , which is used to regularize the latent representation to make it closer to the prior distribution . Since direct sampling of the Gamma distribution is non-differentiable, we use the reparameterization trick of the Gaussian distribution to ensure the model is trainable:

[0127]

[0128] where is the mean term after parameterizing the Gamma distribution; is the standard deviation term of the Gamma distribution; is the noise variable of the standard normal distribution, used for reparameterized sampling to ensure the gradient is computable; represents element-wise multiplication (Hadamard product); represents the sample of the latent variable generated by the reparameterization trick of the Gaussian distribution; The two parameters α and β of the Gamma distribution are calculated as follows:

[0129]

[0130]

[0131] where is the mean output by the variational autoencoder; is the learnable parameter that controls the shape of the Gamma distribution; is the normalization term, initially set to 1.0, output by the normalization encoder to improve numerical stability.

[0132] This reparameterization method ensures that the variational autoencoder can update its parameters through gradient descent during training, thereby learning the latent representation of cervical adenocarcinoma data.

[0133] S143: Predicting the prognosis risk and chemotherapy sensitivity of cervical adenocarcinoma patients based on a deep clustering prediction model

[0134] After learning the latent embedding z of the transcriptome data, we use the deep clustering method to group the prognosis of cervical adenocarcinoma patients. The objective function is as follows:

[0135]

[0136] where, is the predicted prognosis risk of the patient (such as low risk / high risk), or chemotherapy drug sensitivity (such as sensitive / resistant); is a multi-layer perceptron classifier used to map z to specific prediction labels.

[0137] To enhance the discrimination ability of the patient prognosis risk prediction, we use the weighted cross-entropy loss and the maximized evidence lower bound loss to ensure that the model achieves a balance between classification accuracy and latent variable learning:

[0138]

[0139] where, : The true label of the i-th sample, indicating the prognosis of the patient or drug sensitivity, taking values of 0 (no response) or 1 (response). : The predicted probability of the i-th sample, indicating the probability that the sample belongs to a certain category (such as low prognosis risk or high prognosis risk). N: The total number of samples. : The maximized evidence lower bound loss, used to approximate the posterior distribution of the data and improve the generalization ability of the model. : The balance weight. Stratified cross-validation can be used to optimize the loss weight. In this embodiment, , .

[0140] This loss is used to guide the model to learn the clinical significance of different cervical adenocarcinoma gene expression patterns.

[0141] According to the above embodiments of the present invention, through clustering analysis: combining the latent class model with sub-coalition division, ensuring clustering stability through Nash stability. Prediction model: The variational autoencoder and the Gamma mixture model solve the problem of asymmetric distribution of high-dimensional data and improve prediction accuracy. End-to-end process: A complete closed-loop from data preprocessing to clinical prediction, supporting precise medical decision-making.

[0142] Traditional clustering methods (such as K-Means) only rely on global distances, while sub-coalition partitioning ensures the biological consistency of sample membership (such as co-expression networks) through local similarity (cosine similarity) and Nash stability constraints. Transcriptome data often exhibits a right-skewed distribution (such as the long tail of highly expressed genes), making it difficult for Gaussian mixture models to fit. By modeling the latent space with a Gamma mixture model and combining reparameterization techniques, the model can better capture the asymmetric characteristics of gene expression, and experimental data demonstrates the improvement of its clustering metrics (such as the adjusted Rand index).

[0143] Figure 3 The module schematic diagram of a clustering system for transcriptome sequencing data is shown. As Figure 3 shown, a clustering system 200 for transcriptome sequencing data includes:

[0144] A data acquisition and preprocessing module 210, configured to collect cervical adenocarcinoma and adjacent tissue specimens of several patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix;

[0145] A clustering analysis module 220, configured to perform clustering operations on the standardized gene expression matrix based on a clustering analysis method that combines a latent class model and sub-coalition partitioning, model the sample gene expression data through factor analysis and a latent class model, and combine sub-coalition partitioning with Nash stability constraints to maximize gene expression similarity and the joint distribution of latent classes, thereby obtaining a clustered gene set;

[0146] An association analysis module 230, configured to perform gene function annotation on the clustered gene set, and perform association analysis on the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics;

[0147] A deep clustering prediction module 240, configured to construct a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model based on the clustering analysis and association analysis results to predict the prognosis risk or chemotherapy sensitivity of patients.

[0148] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0149] Clustering performance evaluation:

[0150] In the clustering performance evaluation experiment, the present invention compared the performance of the K-Means clustering algorithm, Gaussian mixture model, deep embedded clustering, improved deep embedded clustering, information generative adversarial network, and the algorithm proposed by the present invention in various clustering metrics. The specific analysis is as follows:

[0151] Figure 4 is a performance comparison chart of the adjusted Rand index for different models. From Figure 4 it can be seen that the proposed algorithm shows high clustering performance during the training process. Its adjusted Rand index is always higher than other comparison methods, and it converges faster, finally approaching above 0.8, which is better than all comparison algorithms. Since the proposed algorithm uses the Gamma mixture model as the prior distribution and combines the variational autoencoder for feature learning, it has a stronger modeling ability for complex cervical adenocarcinoma transcriptome data, thus obtaining a better clustering effect.

[0152] Figure 5 is a performance comparison chart of the Davies-Bouldin index for different models. The Davies-Bouldin index is used to measure the quality of clustering. The smaller its value, the better the clustering effect, which means that the samples within the clusters are closer and the separation between clusters is more obvious. As Figure 5 shown, it can be seen that the traditional methods (K-Means clustering algorithm, Gaussian mixture model) perform the worst, with a slow decrease in the Davies-Bouldin index and low clustering tightness. The deep learning methods (deep embedded clustering, improved deep embedded clustering, information generative adversarial network) significantly reduce the Davies-Bouldin index, indicating that they can better mine the deep structure of the data. The proposed algorithm in the present invention has the fastest decrease in the Davies-Bouldin index and finally reaches 0.95 after 50 rounds of training, which is 48% lower than the K-Means clustering algorithm and 13.6% lower than the information generative adversarial network, verifying its superiority in deep clustering tasks and proving that this method can better optimize the intra-class compactness and inter-class separation. It shows the ability of the proposed algorithm to combine the variational autoencoder and the Gamma mixture model, which is more advantageous in learning the clustering structure.

[0153] Figure 6 is a performance comparison chart of the normalized mutual information for different models. In Figure 6 the experiment, the present invention compares the performance of each model on the normalized mutual information index. The proposed algorithm achieves the best performance. Finally, the deep embedded clustering reaches 0.85, which is 8.9% higher than the information generative adversarial network and nearly 88.9% higher than K-Means, indicating that the proposed algorithm can effectively learn complex data distributions, improve clustering quality, and demonstrate the optimal clustering effect.

[0154] Furthermore, the embodiment of the present application also provides an electronic device, including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and the instructions, when executed by the processor, cause the processor to execute any one of the above methods.

[0155] Furthermore, an embodiment of the present application also provides a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute any of the above methods.

[0156] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0157] It should be noted that the embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0158] It should also be noted that in the embodiments of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0159] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the embodiments of the present application, but will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A clustering method for transcriptome sequencing data, characterized in that, Comprising: Collecting cervical adenocarcinoma and adjacent tissue specimens from a number of patients, performing transcriptome sequencing on the specimens and preprocessing the transcriptome sequencing data to obtain a standardized gene expression matrix; Performing clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model and sub - coalition partitioning. Modeling the sample gene expression data through factor analysis and the latent class model, and combining the sub - coalition partitioning with Nash stability constraints to maximize gene expression similarity and the joint distribution of latent classes, to obtain a clustered gene set; Among them, performing clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model and sub - coalition partitioning includes: Constructing a gene expression data model using factor analysis; ; where X is gene expression data, Λ is the factor loading matrix, ε is the error term, is the mixing coefficient for class m, is the parameter for class m; represents the true label of the i-th sample; M is the total number of classes; Calculating the gene expression similarity between samples based on the cosine similarity metric function, and partitioning sub - coalitions according to a preset threshold δ; During the sub - coalition division process, each sample i belongs to the k - th sub - coalition , and its individual utility is defined as , where represents the features of sample i and sample j; represents the k - th sub - coalition; is the cosine similarity metric function between sample i and sample j; : the individual utility of sample i, that is, the sum of its similarities with all other samples within the same sub - coalition; Ensuring that the individual utility of samples in the sub - coalition is not lower than the individual utility after migrating to other sub - coalitions through Nash stability constraints; The optimization objective function includes gene expression similarity and the joint distribution of latent classes, and the objective function is as follows: ; Among them, W(C) is guided by gene expression similarity metrics to divide sub - alliances; : The final optimal sub - alliance division scheme; λ: The balance parameter that controls the weights of gene expression similarity and the part of calculating class probabilities in the class model; : The prior probability of the k - th sub - alliance; K: The total number of sub - alliances; N: The total number of samples; : The parameter given for the k - th sub - alliance When, the conditional probability that sample i belongs to this class (i.e., belongs to the k - th sub - alliance); : Calculating class probabilities by the latent class model, measuring the possibility that a sample belongs to different classes; By optimizing the weight λ and the class parameters and the mixing coefficient to obtain the best classification of cervical adenocarcinoma gene expression; Performing gene function annotation on the clustered gene set, and performing correlation analysis between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics; Based on the results of clustering analysis and correlation analysis, constructing a deep clustering prediction model based on a variational auto - encoder and a Gamma mixture model to predict the prognosis risk and chemotherapy sensitivity of patients.

2. The clustering method for transcriptome sequencing data according to claim 1, wherein Among them, The Nash stability condition for the sub - coalition partitioning is: ; That is, for any sample i, its current individual utility in the sub - coalition , where is the individual utility of sample i in its current sub - coalition; is the individual utility of sample i if it migrates to other sub - coalitions after that.

3. A clustering method for transcriptome sequencing data according to claim 2, characterized in that, Among them, Performing correlation analysis between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics includes: Analyzing the linear relationship between gene expression and clinical characteristics using the Pearson correlation coefficient; Comparing the gene expression differences between different clinical groups through t - tests.

4. A clustering method for transcriptome sequencing data according to claim 3, wherein, Among them, Based on the results of clustering analysis and correlation analysis, constructing a deep clustering prediction model based on a variational auto - encoder and a Gamma mixture model to predict the prognosis risk and chemotherapy sensitivity of patients, includes: Using a variational auto - encoder to reduce the dimension of high - dimensional gene expression data to extract latent features, and modeling the asymmetric distribution in the latent space with a Gamma mixture model, clustering the latent space to identify molecular subgroups with different prognosis risks; Optimizing the Gamma distribution parameters through the reparameterization trick, training the model by combining the cross - entropy loss and the objective function weighted by the maximum evidence lower bound (ELBO), and outputting the prediction results of the prognosis risk and chemotherapy sensitivity of patients.

5. A clustering method for transcriptome sequencing data according to claim 4, characterized in that, Among them, The combination of the variational auto - encoder and the Gamma mixture model includes: Defining the probability density function of the Gamma mixture model in the latent space of the variational auto - encoder; the input of the variational auto - encoder is high - dimensional gene expression data X, and the output is the distribution parameters of the latent embedding representation variable z; Performing differentiable optimization on the Gamma distribution sampling through the Gaussian distribution reparameterization trick, optimizing the Gamma distribution parameters to adapt to the deep learning framework, training the model by combining the cross - entropy loss and the objective function weighted by the maximum evidence lower bound, and outputting the prediction results of the prognosis risk and chemotherapy sensitivity of patients.

6. A clustering method for transcriptome sequencing data according to claim 5, wherein Among them, Differentiable optimization of Gamma distribution sampling is performed through the Gaussian distribution reparameterization technique to optimize the Gamma distribution parameters, including the following calculation formulas: ; Among them, is the mean term after parameterizing the Gamma distribution; is the standard deviation term of the Gamma distribution; is the noise variable of the standard normal distribution; represents element-wise multiplication; represents a sample of latent variables generated by the reparameterization trick of the Gaussian distribution; Calculate the Gamma distribution parameters through the parameter conversion formula: ; ; Among them, is the mean output by the variational autoencoder; is a learnable parameter that controls the shape of the Gamma distribution; is the normalization term; respectively represent the shape parameter and rate parameter of the Gamma distribution.

7. A clustering method for transcriptome sequencing data according to claim 6, wherein Wherein, The decoder output of the variational autoencoder is mapped to the patient prognosis risk and chemotherapy sensitivity prediction labels through a multi-layer perceptron (MLP): ; wherein, is the predicted patient prognosis risk or chemotherapy drug sensitivity; is a multi-layer perceptron classifier for mapping to specific prediction labels.

8. A clustering system for transcriptome sequencing data, which is used to implement a clustering method for transcriptome sequencing data as described in any one of claims 1-7, characterized in that, Including: A data acquisition and preprocessing module, configured to acquire cervical adenocarcinoma and adjacent tissue specimens of a number of patients, perform transcriptome sequencing on the specimens, and preprocess the transcriptome sequencing data to obtain a standardized gene expression matrix; A clustering analysis module, configured to perform clustering operations on the standardized gene expression matrix through a clustering analysis method combining a latent class model and sub-coalition partitioning, model the sample gene expression data through factor analysis and a latent class model, and combine the sub-coalition partitioning with Nash stability constraints to maximize the gene expression similarity and the joint distribution of latent classes, to obtain a clustered gene set; An association analysis module, configured to perform gene function annotation on the clustered gene set, and perform association analysis on the functionally annotated clustering results and the clinical features of cervical adenocarcinoma to obtain the association between gene expression and clinical features; A deep clustering prediction module, configured to construct a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model based on the clustering analysis and association analysis results to predict the patient prognosis risk and chemotherapy sensitivity.

9. An electronic device, comprising: A processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is used to store one or more programs, and the one or more programs include instructions, wherein the instructions, when executed by the processor, cause the processor to execute the steps of a clustering method for transcriptome sequencing data according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Node cooperative caching method based on cooperative game and deep learning

    CN115915277A