Methods and models based on transcriptome and intratumoral microorganisms

By constructing a computational model that integrates gene expression and microbial abundance characteristics, the low sensitivity and low accuracy of cervical cancer lymph node metastasis diagnosis is solved, and high-precision early risk assessment and personalized therapeutic support are achieved.

CN120199322BActive Publication Date: 2025-08-01ZHEJIANG CANCER HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510671977.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-01
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art has low sensitivity and accuracy in the diagnosis of lymph node metastasis in cervical cancer, resulting in many patients not being identified and lack of effective biomarkers for early warning and personalized treatment.

Method used

A calculation model that integrates gene expression characteristics and microbial abundance characteristics is constructed. Through multi-dimensional biological information feature mining and machine learning algorithms, a classification prediction model is constructed to evaluate the risk of lymph node metastasis in patients with cervical cancer.

Benefits of technology

It significantly improves the prediction accuracy of lymph node metastasis in cervical cancer, realizes support for early risk stratified assessment and personalized treatment, has high-throughput technical support and broad applicability, and provides interpretable predictive results and integrated evaluation tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199322B_ABST
    Figure CN120199322B_ABST
Patent Text Reader

Abstract

The present invention provides a method and a model based on transcriptome and intratumoral microorganisms, specifically relating to a computational model for evaluating the risk of lymph node metastasis in cervical cancer and related methods, programs, systems and devices. This model integrates the abundance information of multiple gene markers (including CARD9, MNX1_AS2, MRAS, OLFML2A and RPS28) and multiple microbial genera (including Dialister, Catonella and Campylobacter) as input features, and constructs a risk prediction model through machine learning algorithms, so as to realize the quantitative evaluation of the risk of lymph node metastasis in cervical cancer patients. The present invention improves the prediction accuracy of lymph node metastasis in cervical cancer, provides a basis for clinical auxiliary diagnosis and decision-making, and has important medical application value and industrial transformation potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical diagnosis, and specifically to a diagnostic model for cervical cancer lymph node metastasis, a construction method thereof, and an application thereof in the preparation of a diagnostic product for cervical cancer lymph node metastasis. Background Art

[0002] Cervical cancer (CC) ranks fourth among female malignant tumors and is a common malignant tumor of the female reproductive system, seriously threatening the life, physical and mental health of women. One of the important factors affecting the prognosis of CC patients is lymph node metastasis. Studies have confirmed that early CC with lymph node metastasis will seriously affect its 5-year survival rate. Whether the lymph nodes are metastasized not only affects the prognosis of cervical cancer patients, but also plays a decisive role in individualized treatment. For early CC patients, systematic lymph node resection is mostly used to determine the occurrence of lymph node metastasis and provide a basis for postoperative adjuvant treatment. However, at the same time, a certain number of CC patients without lymph node metastasis undergo unnecessary surgeries, thus leading to unnecessary surgical risks. Therefore, accurately judging the lymph node status of CC patients and accurately assessing whether lymph node metastasis has occurred are crucial for clarifying the stage, judging the prognosis and formulating treatment plans. At present, the clinical diagnosis of CC lymph node metastasis mainly relies on magnetic resonance, CT or PET / CT imaging to examine the morphological characteristics of lymph nodes. The most commonly used criterion is whether the short diameter of the lymph node exceeds 10 mm. Although this method has acceptable specificity, its sensitivity and accuracy are poor. As a result, a large proportion of CC patients with lymph node metastasis are not identified. Therefore, there is still a need to find new biomarkers (or key genes) that can predict CC lymph node metastasis. The intratumoral microbiome refers to the genomes of microorganisms (bacteria, archaea, fungi, and viruses) present in the tumor parenchyma and the tumor microenvironment. The discovery of intratumoral microorganisms has greatly increased people's understanding of the complexity of the tumor microenvironment. Initially, these microorganisms were regarded as bystanders in tumor development. However, with the in-depth study, people began to realize that microorganisms and their metabolites in the tumor microenvironment play a key role in the pathogenesis of cancer. They not only promote or inhibit the occurrence and development of tumors, but also significantly affect the metabolism and therapeutic effects of related drugs in the body, thus gradually becoming new tumor treatment targets. In addition, previous studies have found that intratumoral microorganisms in CC are related to CC metastasis and affect the disease prognosis, yet their specific mechanism of action has not been fully elucidated.

[0003] In recent years, with the emergence and rapid development of high-throughput sequencing technology, more and more CC genes and epigenetic features have been discovered. However, few studies have been conducted on intratumoral microbiomics based on clinical CC lymph node metastasis and CC primary site patient samples. Multi-omics research is a method to explore the interactions between various substances in biological systems, including transcriptomics, microbiomics, proteomics, metabolomics, etc. These substances jointly affect the phenotypes, traits, etc. of the life system. Each type of omics data usually provides a list of differential factors that may be related to diseases, and this data can be used as disease markers. Multi-omics data helps to explore different levels of systems biology from an overall perspective. Integrating different types of omics data may help to clarify the potential pathogenic changes leading to diseases, or can be used to identify potential therapeutic targets for further molecular research. Summary of the Invention

[0004] The object of the present invention is to provide a computational model that integrates gene expression features and microbial abundance features for evaluating the lymph node metastasis risk of cervical cancer patients. This model is based on multi-group clinical sample data, through the mining and feature screening of multi-dimensional bioinformatics features, and combines machine learning algorithms to construct a classification prediction model to achieve early warning of high-risk metastatic patients.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] On the one hand, the present invention provides a computational model, whose input is the expression levels of specific gene markers (CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28) in cervical cancer samples, and the relative abundances of specific microbial genera (Dialister, Catonella, and Campylobacter), and the output is the lymph node metastasis risk score of the patient.

[0007] The model is obtained by training on a large sample data set. The model algorithms include but are not limited to random forest, logistic regression, LASSO regression, COX regression, artificial neural network, decision tree, support vector machine, naive Bayes, K-nearest neighbor algorithm, gradient boosting tree, Adaboost algorithm, XGBoost algorithm, LightGBM algorithm, CatBoost algorithm, multi-layer perceptron, convolutional neural network, recurrent neural network, long short-term memory network, gated recurrent unit network, etc. The robustness and generalization ability of the model are ensured through cross-validation and feature importance analysis.

[0008] The model described in the present invention can further be implemented by a set of computer programs, which include the following steps when executed by a processor: receiving the raw data of the sample to be tested, performing data preprocessing and standardization, inputting the trained model, and outputting the metastasis risk score and prediction probability.

[0009] The present invention also provides a method for constructing the model, and the steps include: collecting cervical cancer tissue samples with and without lymph node metastasis, respectively performing transcriptome and microbiome detection; using statistical and machine learning methods for differential analysis and feature screening; constructing a prediction model and an optional nomogram based on the screened key features.

[0010] Furthermore, the present invention provides a software system and hardware device integrating the model. Users can input the detection data of patients and automatically output the metastasis risk score for clinical assistant diagnosis.

[0011] The evaluation model proposed by the present invention has one of the following significant technical advantages and / or clinical values:

[0012] Multi-dimensional feature integration: The present invention first incorporates gene expression information highly related to cervical cancer metastasis and cervical-related microecological microbial genera into the modeling input, captures the comprehensive biological characteristics of cancer progression, and improves the prediction accuracy.

[0013] High-throughput technology support: The model is based on multi-omics data (transcriptome and microbiome) of a large sample of real-world patients, and has both clinical representativeness and wide applicability.

[0014] Driven by artificial intelligence algorithms: Multiple machine learning algorithms are used for comparison and optimization in model construction, significantly improving the model's ability to distinguish between metastasis and non-metastasis, and supporting the migration and application of the model on different data sets.

[0015] Early warning: The model can assist doctors in achieving risk stratification before treatment.

[0016] Standardized evaluation tool: The present invention also constructs a diagnostic nomogram tool, which can intuitively display the contribution of each feature to the score and provide interpretable prediction results for clinicians.

[0017] System equipment can be integrated: By integrating computer programs and evaluation devices, the integration of risk assessment can be realized, and it has the potential for industrial transformation and popularization and application.

[0018] The present invention proposes a machine learning model integrating molecular biology information and microecological information for the prediction and risk assessment of cervical cancer lymph node metastasis.

[0019] In the present disclosure, a "gene marker" refers to a gene whose expression level in tumor tissue is highly related to the disease state (such as metastasis).

[0020] "Genus of microorganism" refers to a group of microorganisms at the taxonomic level of "genus" identified by, for example, microbial 16S rRNA or metagenomic detection.

[0021] "Abundance" means that the microbial abundance represents the relative occurrence frequency of a genus in the total sample, often expressed as a percentage.

[0022] "Nomogram" refers to a visualization tool for presenting a regression model, which can map multiple predictors to a score interval to estimate the result probability.

[0023] In order to identify the differential genes between the MN sample and the MP sample, the present invention analyzed the differential genes between the MP and MN samples (MP group VS MN group) in the transcriptome dataset through the "DESeq2" package, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC| > 0.5, p-value < 0.05); subsequently, the volcano plot and heat map of the DEGs were drawn using the R packages "ggplot2" and "ComplexHeatmap" (the top 10 up-regulated genes with the largest change in |log2FC|: SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, TEKT4; the top 10 down-regulated genes: FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, PAK5).

[0024] The gene numbers of the top 10 up-regulated genes with the largest change in |log2FC| are respectively: SOHLH1 (ID: 402381), MUCL3 (ID: 283232), PNMA5 (ID: 114836), REG1A (ID: 5967), FGA (ID: 2243), PGC (ID: 5225), DSCAM-AS1 (ID: 102723407), IGHV3-22 (ID: 28450), LINC01320 (ID: 100507487), TEKT4 (ID: 150465).

[0025] The top 10 downregulated genes with the largest change in |log2FC| are FOLR1P1 (ID: 107986793), TLX1 (ID: 3195), MGAM2 (ID: 9342), KRT1 (ID: 3848), KRTDAP (ID: 200172), LINC00167 (ID: 100507065), MUCL1 (ID: 118430), LINC00457 (ID: 285237), BPIFC (ID: 653145), and PAK5 (ID: 57144).

[0026] Based on the above analysis, 1057 upregulated genes were obtained in the samples of patients with cervical cancer lymph node metastasis. Among the top 10 upregulated genes with the largest change in |log2FC|, there is at least one of SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, and TEKT4.

[0027] Based on the above analysis, 317 downregulated genes were obtained in the samples of patients with cervical cancer lymph node metastasis. Among the genes with the largest change in |log2FC|, there is at least one of FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5.

[0028] Preferably, in step 4, in order to obtain key genes, bioinformatics analysis methods are performed, including systematically performing correlation analysis, machine learning, ROC analysis, etc. on the differential genes.

[0029] The present invention relates to a method for evaluating the risk of lymph node metastasis in cervical cancer patients, in particular to a nomogram prediction model constructed based on key genes and microbiome characteristics and its construction and validation methods. This method combines the transcriptome expression information of patients with the compositional characteristics of the intestinal or cervical local microbiome, and improves the accuracy and practicality of the clinical risk stratification of cervical cancer through statistical modeling and visualization tools.

[0030] In a preferred embodiment of the present invention, first, mRNA expression profile data of cervical cancer tissues are obtained through transcriptome sequencing technology, and the DESeq2 software is used to perform differential expression analysis on patient samples with and without lymph node metastasis, and candidate genes significantly related to lymph node metastasis are screened out from them. Subsequently, combined with variable selection methods such as LASSO, a key gene set with high predictive value is further identified.

[0031] In the process of constructing the prediction model, it is preferred to divide the original sample data set into a training set and a validation set. Based on the training set, a multi-factor regression analysis method (preferably the Logistic regression model) is used to evaluate the correlation between the expression levels of the identified key genes and lymph node metastasis, and a nomogram is drawn based on the regression results. Preferably, the rms package in R language is used to construct the nomogram model. The nomogram can map the expression level of each key gene to a risk score, and predict the probability (Pr(Y)) of lymph node metastasis for each patient by calculating the total points.

[0032] To improve the interpretability and prediction performance of the model, the present invention further uses multiple groups of key gene combinations (including any two, three, four or all five genes) for nomogram modeling respectively, and selects the best combination by calculating the area under the ROC curve (AUC) of each combined model. The performance evaluation of the nomogram model also includes the C-index (Harrell concordance index), calibration curve (preferably drawn using the rms package), decision curve analysis (DCA, preferably drawn using the ggDCA package), and receiver operating characteristic curve (ROC, preferably drawn using the pROC package) to comprehensively reflect the goodness of fit, clinical applicability and discriminative ability of the model.

[0033] The present invention further preferably combines microbiome information, and obtains the composition data of the local cervix or intestinal microbiota of each patient through sequencing analysis. Based on the key microbial genera that have been identified as being closely related to lymph node metastasis (such as Dialister, Catonella, and Campylobacter), their relative abundances are extracted as modeling features, and jointly constitute combined features with the key gene expression values and input into the model to construct a more comprehensive nomogram prediction tool. This combined model significantly improves the AUC value of the model, indicating the important role of microbiome data in predicting the progression of cervical cancer.

[0034] In the process of model construction, the top layer of the nomogram is used to display the scores (Points) of each key gene and microbial genus. The corresponding scores are obtained according to the specific values of each variable and summed to obtain the "Total Points", and the corresponding probability axis below is the quantitative probability of predicting the risk of lymph node metastasis.

[0035] The key genes of the present invention preferably include at least one of CARD9 (Gene ID: 64170), MNX1_AS2 (Gene ID: 100873971), MRAS (Gene ID: 22808), OLFML2A (Gene ID: 79589), and RPS28 (Gene ID: 6234), or any combination of two, three, or four of them. More preferably, it is the complete combination of the above five genes. The key microorganisms of the present invention preferably include at least one of the key microorganism genera closely related to lymph node metastasis, such as Dialister, Catonella, and Campylobacter genera, or any combination of two or three of them. More preferably, it is the complete combination of the above three microorganism genera. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1A and Figure 1B are the volcano plot ( Figure 1A ), and the expression heatmap ( Figure 1B ) of the differential genes between the MP group and the MN group drawn using the R packages "ggplot2" and "ComplexHeatmap" in the preferred Embodiment 1 of the present invention. Among them, the volcano plot shows the TOP10 genes with the highest up- and down-regulation fold changes. Red indicates high expression, green indicates low expression, and gray indicates genes with no significant difference. The heatmap consists of two parts: the upper part is the density heatmap of the expression levels of the differential genes, showing the lines of five quantiles and the average value; the lower part is the expression heatmap of the differential genes. Each row represents the expression level of each gene in different group samples, and each column represents the expression levels of the top 10 up- and down-regulated DEGs in each sample. The color represents the magnitude of the gene expression level. The darker the color, the higher the expression level of the DEGs (red for high expression, cyan for low expression).

[0038] Figure 2 : The histogram of differential microorganisms between the MP group and the MN group. Among them, the X-axis is the LDA score, indicating the contribution degree of different genera to the inter-group difference; the Y-axis lists the names of each genus. The larger the LDA value, the more significant the difference of the genus between different groups.

[0039] Figure 3: Diagnostic nomogram for constructing a model with key genes and microorganisms. Among them, the top layer of the nomogram represents the scores (Points) of each key gene and key microorganism. The corresponding scores are obtained according to the values of the key genes and microorganisms, and then the scores are added together to get "Total points". The higher the total score, the higher the predicted risk. The risk probability (Pr(Y)) calculated based on the total score is shown at the bottom.

[0040] Figure 4 : Calibration curve of the model nomogram in the training set. Among them, the X-axis represents the predicted probability, and the Y-axis represents the actual probability. The blue curve is the actual prediction result (Apparent), the black curve is the result after bias correction (Bias-corrected), and the dashed line is the ideal situation (Ideal), that is, the reference line where the model prediction is exactly the same as the actual. The p-value of the Hosmer-Lemeshow test is 0.726, indicating that the calibration performance of the model is good and does not deviate significantly from the ideal situation.

[0041] Figure 5 : DCA curve of the model nomogram in the training set. Among them, the X-axis represents the risk threshold, and the Y-axis represents the net benefit. Curves of different colors represent different prediction models, including the nomogram model (Nomogram), the control lines of all (All) and none (None). The higher the net benefit of the curve, the greater the decision value of the model at the corresponding risk threshold. The dashed line represents the control situation without any intervention in all samples.

[0042] Figure 6 : ROC curve of the diagnostic model in the training set. Among them, the ROC (Receiver Operating Characteristic) curve is used to evaluate the performance of the diagnostic model, where the horizontal axis is 1 - Specificity (false positive rate), and the vertical axis is Sensitivity (true positive rate). The gray dashed line represents the baseline of random prediction. The red curve is the ROC curve of the diagnostic model, and the area under the curve (AUC, Area Under the Curve) is 0.8924, indicating that the model has good discrimination ability.

[0043] Figure 7 : Flowchart of the operation of the computer system in an embodiment of the present invention.

[0044] Figure 8 : Structural block diagram of the intelligent device in an embodiment of the present invention. Detailed implementation mode

[0045] Example 1: Screening and modeling of key genes and microbial genera

[0046] In this study, transcriptome and intratumoral microbiome sequencing data of cervical tissues from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis were used. Through bioinformatics methods, the molecular changes and pathogenesis of CC lymph node metastasis were comprehensively and deeply explored. Based on a variety of bioinformatics analysis methods, the diversity, community composition and differential microorganisms of intratumoral microorganisms in CC lymph node metastasis were systematically explored, and their correlation analysis, machine learning and ROC analysis were performed with differential genes to screen out key genes.

[0047] Total sample RNA was isolated and purified using TRIzol (Thermofisher, 15596018) according to the manufacturer's protocol. Then, the quantity and purity of the total RNA were quality controlled using a NanoDrop ND-1000 (NanoDrop, Wilmington, DE, USA), and the integrity of the RNA was detected using a Bioanalyzer 2100 (Agilent, CA, USA); a concentration > 50 ng / μL, RIN value > 7.0, and total RNA > 1 μg were required to meet the downstream experiment requirements. Poly(A)-tailed mRNA was specifically captured using oligo(dT) magnetic beads (Dynabeads Oligo(dT), cat. 25-61005, Thermo Fisher, USA) through two rounds of purification. The captured mRNA was fragmented at high temperature using a magnesium ion fragmentation kit (NEBNext Magnesium RNA Fragmentation Module, cat. E6150S, USA) at 94°C for 5 - 7 minutes. The fragmented RNA was reverse transcribed into cDNA using reverse transcriptase (Invitrogen SuperScriptTM II Reverse Transcriptase, cat. 1896649, CA, USA). Then, E. coli DNA polymerase I (NEB, cat. m0209, USA) and RNase H (NEB, cat. m0297, USA) were used for second-strand synthesis to convert the DNA-RNA hybrid double-strand into a DNA double-strand, and dUTP Solution (Thermo Fisher, cat. R0133, CA, USA) was incorporated into the second strand to fill in the ends of the double-stranded DNA to make them blunt ends, and then an A base was added to each end to enable ligation to an adapter with a T base at the end, and magnetic beads were used to screen and purify the fragment sizes. The second strand was digested with UDG enzyme (NEB, cat. m0280, MA, US), and then PCR was performed - pre-denaturation at 95°C for 3 minutes, 98°C denaturation for a total of 8 cycles, 15 seconds each, annealing to 60°C for 15 seconds, extension at 72°C for 30 seconds, and finally extension at 72°C for 5 minutes to form a library with a fragment size of 300 bp ± 50 bp (strand-specific library). Finally, double-end sequencing was performed on it using an illumina NovaseqTM 6000 according to the standard operation, and the sequencing mode was PE150.Valid Data that can be aligned to the reference genome can be defined as being aligned to exons, introns, and intergenic regions according to the regional information of the reference genome. Generally, for species with relatively well-annotated genomes (such as model species like humans and Arabidopsis thaliana), the percentage content of the sequencing sequence mapped to the exon region should be the highest. The reads mapped to the intron and intergenic regions may be due to precursor mRNA splicing events, ncRNAs (non-coding RNAs), imperfect genome annotation, DNA contamination, and background noise, etc. It was found that the content of all samples mapped to the exon region reached over 90%, indicating that the sequencing data quality was good and the sequencing results were reliable.

[0048] To determine whether there are clusters or outliers in the samples, the FactoMineR package was used to perform principal component analysis (PCA) for data dimensionality reduction on the transcriptome sequencing dataset, in order to identify the dispersion of samples in the metastasis-negative (MN) group and metastasis-positive (MP) group of lymph nodes through patternization and visualization. It was found that PC1 explained 22.37% of the variance, which means that the first principal component can capture relatively high variability in the data. The combined variance explained by PC1 and PC2 was 29.97% (22.37% + 7.6%), which means that these two principal components can explain nearly one-third of the variability in the dataset.

[0049] To identify the differentially expressed genes between MN samples and MP samples, the "DESeq2" package was used to analyze the differentially expressed genes between MP and MN samples (MP group VS MN group) in the transcriptome dataset, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC|>0.5, p-value<0.05); subsequently, the R packages "ggplot2" and "ComplexHeatmap" were used to plot the volcano plot and heatmap of DEGs, as Figure 1A and Figure 1BAs shown below. The top 10 up-regulated genes with the largest change in |log2FC|: SOHLH1 (spermatogenesis- and oogenesis-specific basic helix-loop-helix protein 1), MUCL3 (mucin-like protein 3), PNMA5 (paraneoplastic Ma antigen 5), REG1A (regenerating gene 1A), FGA (fibrinogen alpha chain gene), PGC (transcriptional coactivator), DSCAM-AS1 (Down syndrome cell adhesion molecule antisense RNA 1), IGHV3-22 (immunoglobulin heavy chain variable region 3-22), LINC013200 (long non-coding RNA 013200), TEKT4 (Tektin 4); the top 10 down-regulated genes: FOLR1P1 (folate receptor 1 pseudogene 1), TLX1 (T-cell leukemia homeobox gene 1), MGAM2 (maltase-dextrinase 2), KRT1 (keratin 1), KRTDAP (keratin-related protein), LINC00167 (long non-coding RNA 00167), MUCL1 (breast cancer differentiation-related protein 1), LINC00457 (long non-coding RNA 00457), BPIFC (BPI fold domain family member C), PAK5 (p21-activated kinase 5).

[0050] Identification of differential microorganisms

[0051] To explore the main microbial flora causing the dynamic differences in the microbiota and identify the differential microorganisms at each level in different groups, we used linear discriminant analysis (LDA) and effect size (LEfSe) to analyze the differential microorganisms mainly present at the genus level between the MN and MP groups (p < 0.05, LDA score (log10) > 2). The role of LEfSe analysis is to find the microorganisms with significant differences between different groups. Through the LDA value distribution histogram (as Figure 2 shown), it intuitively shows the genera that are significantly enriched at different levels in the three groups and have a key impact on grouping. There are a total of 6: Fusobacterium, Sneathia, Dialister, Campylobacter, Streptophyta_Unknown_family_Unknown_genus123, Catonella, which are recorded as differential microorganisms.

[0052] Subsequently, based on the union of key genes and key microorganisms, the AUC values of the ROC curves of the nomogram models constructed with different combinations were calculated to select the gene-microorganism combination with the highest AUC value. It was found that the diagnostic model constructed with the combination of CARD9, MNX1_AS2, MRAS, OLFML2A, RPS28, Dialister, Catonella, and Campylobacter had the highest AUC value (0.892). The nomogram consists of "scores" and "total scores". The former represents the scores of each key gene, and the latter represents the sum of the scores of all key genes. The higher this value, the higher the probability of lymph node metastasis in CC (as Figure 3 shown), and the "rms" package was used to draw the calibration curve in the analysis (as Figure 4 shown), the "ggDCA" package was used to draw the DCA decision curve (as Figure 5 shown), and the "pROC" package was used to draw the ROC curve to evaluate the fitness of the model (as Figure 6 shown). The results showed that the nomogram had good predictive ability (HL test p.value > 0.05, the benefits of the DCA curve were basically above ALL and NONE, and the AUC value = 0.892).

[0053] In this example, the Logistic regression model was constructed using the rms package, and the nomogram of the regression model was drawn using the regplot package. The Logistic regression model was used, and the log2(x + 1) values of the CPM of gene expression and the log2(x + 1) values of the relative abundances of microbial genera were used. The expression formula of this model is as follows:

[0054]

[0055] The OR values of each factor in the diagnostic model were calculated as shown in the following table.

[0056] Table 1 OR values of each factor in the model

[0057]

[0058] Example 2: Establishment and verification of a calculation model for evaluating the risk of lymph node metastasis in cervical cancer

[0059] In this example, transcriptome sequencing data were collected from cervical tissue collected from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis, as described in Example 1. Differential expression analysis was performed using DESeq2 to identify candidate genes with significant differential expression between metastatic and non-metastatic samples. LASSO regression was further used for variable selection, ultimately identifying five gene markers with high predictive value: CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28; and three microbial genera: Dialister, Catonella, and Campylobacter. A random forest algorithm was used to train a binary classification model using the expression levels of these five genes and the relative abundance of the three microbial genera as feature variables. The data were divided into training and test sets at a ratio of 7:3. In the test set, the model achieved an AUC (area under the curve) of 0.89, demonstrating that the model has extremely high predictive efficacy in distinguishing patients with cervical cancer from those with lymph node metastasis. To ensure the stability and accuracy of model performance evaluation, 10-fold cross-validation was used to evaluate model performance. In this example, 10-fold cross-validation was performed within the original data consisting of the training and test sets, employing a nested cross-validation (CV) strategy. During each round of training, the original data was first divided into 10 subsets: nine serving as training sets and one serving as validation sets for evaluating model performance. Each of the ten subsets was used as the validation set in turn, while the remaining subsets served as training subsets for model training. Performance metrics were recorded for each round. The results from all 10 rounds were finally averaged to evaluate model performance. Furthermore, to further confirm the robustness of the model under different random partitions, a repeated 10-fold cross-validation strategy (5 repetitions) was employed. The calculated average AUC was 0.90, with a standard deviation of ±0.037 and a 95% confidence interval of [0.88, 0.95]. This demonstrates that the model's predictive performance across the training samples is stable and has a low risk of overfitting.

[0060] Example 3: External independent sample verification of model effectiveness and comparative analysis with traditional methods

[0061] To further verify the reliability and generalization ability of the calculation model for evaluating the risk of cervical cancer lymph node metastasis described in Example 2, this example introduces a group of external independent verification samples and compares and analyzes the model prediction results with traditional imaging assessment methods.

[0062] A total of 24 independent cervical cancer patient samples were included in this example, including 12 cases in the metastasis group and 12 cases in the non-metastasis group. All samples were not involved in the model training process and the source was the same as that in Example 1. The sample collection and pretreatment process included RNA extraction, preparation of dPCR premix with specific primers, probes, dNTPs, and polymerase, generation of droplets, thermal cycle amplification, total DNA extraction, droplet detection, quality control, and standardization processing.

[0063] For the above samples, the absolute quantification technology method of dPCR was used to detect the absolute expression levels of 5 key genes (CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28) and the absolute quantification of 3 key microbial genera (Dialister, Catonella, and Campylobacter). Since the training model was constructed based on RNA-seq counts, to ensure data consistency, a standardization conversion strategy (such as Z-score standardization) was adopted in the external validation, so that the dPCR expression values could be input into the model for prediction.

[0064] The standardized expression data was input into the random forest model trained in Example 2, and the model automatically output the lymph node metastasis prediction probability for each patient. The final results were as follows: AUC (ROC): 0.884; Accuracy: 87.5%; Sensitivity: 91.7%; Specificity: 83.3%; Positive Predictive Value (PPV): 84.6%; Negative Predictive Value (NPV): 90.9%. Compared with the traditional imaging evaluation method (CT / MRI combined judgment, with an accuracy of about 71.0%, a sensitivity of about 65.0%, and a specificity of about 76.0%), this model showed higher prediction accuracy and sensitivity in external independent samples. The above results indicated that the evaluation model described in Example 2 could still maintain good prediction performance in new independent samples, had strong generalization ability and stability, and was superior to the existing traditional imaging evaluation methods, with high clinical application potential.

[0065] Example 4: Implementation of a computer system based on the model of the present invention

[0066] In this example, a computer system was constructed to run the above risk prediction model, and the system included the following modules:

[0067] Data Import Module: Used to read the standardized gene expression and microbial abundance data of patient samples, supporting formats such as CSV and Excel. Preprocessing Module: Includes functions such as missing value handling and standardized transformation to ensure the consistency of input data. Model Calculation Module: Loads the pre-trained random forest model file, performs inference on the input data, and outputs the predicted risk scores and probabilities for each sample. Visualization Module: Displays the prediction results in the form of bar charts, heatmaps, etc., facilitating doctors to quickly understand the results. Interface Module: Can be docked with the hospital information system (HIS) or laboratory data system (LIMS) to achieve automatic data transfer and automatic triggering of risk assessment. This system is deployed on the hospital server or private cloud environment, using the Django framework to develop the front-end interface, and the back-end model is implemented using the scikit-learn library in Python (see Figure 7 )

[0068] Example 5: Integrated Lymph Node Metastasis Risk Assessment Device

[0069] The inventors of the present invention designed an intelligent device for evaluating the risk of cervical cancer lymph node metastasis, mainly including:

[0070] Detection Module: Construct a Panel by designing multiplex primers and probes for the above 5 genes and 3 genera of microorganisms. Based on the exclusive Panel, the dPCR detection system can integrally detect the expression levels of 5 genes and 3 genera of microorganisms in cervical cancer tissue samples. Data Acquisition Module: Connects to the detection module and transmits the expression data to the main control system in real time. Main Control Processing Module: The standardized transformation strategy (Z-score standardization) converts the dPCR expression values into model standardized data. The embedded processor runs the model program of the present invention, integrates the Linux system, has a memory of 8GB, and is equipped with a dedicated operation acceleration chip (such as the ARM Cortex-A72 series). Output Display Module: Built-in touch screen display interface, which can display the prediction score, metastasis probability, result suggestions, and visualization report in real time. Communication Module: Supports WiFi, Bluetooth, USB, and Ethernet interfaces, facilitating communication with the hospital database or the host computer. After the device is started, the operator only needs to import the sample, click the detection and prediction buttons, and then the full-process evaluation can be completed, providing a timely auxiliary basis for clinical surgical decisions (see Figure 8 )

[0071] The above are only the embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A computational model product for assessing the risk of cervical cancer lymph node metastasis, which takes the expression levels of multiple gene markers and the microbial abundances based on microbial genera as inputs and outputs the metastasis risk score of a patient; the gene markers are a group consisting of the following gene markers: CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28; the microbial genera are composed of Dialister, Catonella, and Campylobacter genera.

2. The computational model product according to claim 1, wherein the model is established based on a machine learning algorithm selected from random forest, logistic regression, LASSO regression, COX regression, or neural network algorithm.

3. A computer program product that, when executed by a processor, is used to implement the evaluation function of the computational model product according to claim 1 or 2, and the program includes the following steps: a) Receive the gene expression levels and microbial abundance data of the sample to be tested; b) Perform data preprocessing and normalization; c) Input the normalized data into the prediction model; d) Output the lymph node metastasis risk score and prediction probability corresponding to the sample.

4. A computer system for assessing the risk of cervical cancer lymph node metastasis, which includes a processor, a memory, and a program module stored in the memory and executed by the processor. The program module implements the steps according to claim 3 and is used to implement the risk assessment of the target sample.

5. An evaluation device for assessing the risk of cervical cancer lymph node metastasis, characterized in that, Including: A data acquisition device for acquiring the detection data of the evaluation object, and the detection data is the level value of the biomarker in the cervical cancer lymph node metastasis biomarker detected from the sample of the evaluation object suspected of having cervical cancer lymph node metastasis; A data processing device for calculating the index value of the biomarker in the cervical cancer lymph node metastasis biomarker according to the detection data of the evaluation object, and then determining whether the calculated index value of the biomarker is within the risk range of the corresponding biomarker; The biomarker includes gene markers and microbial markers, wherein the gene markers are a group consisting of the following gene markers: CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28; the microbial markers are microorganisms composed of Dialister, Catonella, and Campylobacter genera.

6. The evaluation device according to claim 5, wherein Wherein the device further includes a detection device for detecting the level value of the biomarker.

7. The evaluation device according to claim 5, characterized in that, The data acquisition device is an input device for inputting data or a communication device for reading data from an external data storage device or the storage device of the detection device through an interface.

8. The evaluation device according to claim 7, characterized in that When using a communication device, its corresponding external data storage device is the data memory on the device for measuring biomarkers.

Citation Information

Patent Citations

  • Method for checking cervical cancer using living cell technology

    CN101059505A

  • HNC (Head and Neck Cancer) prognosis biomarker based on lymph node microbial flora and application of HNC prognosis biomarker

    CN113684242A