Method and model based on transcriptome and intratumoral microorganisms
By constructing a computational model that integrates gene expression and microbial abundance characteristics, the problem of insufficient sensitivity and accuracy of cervical cancer lymph node metastasis diagnosis in the prior art is solved, and high-precision prediction of the risk of lymph node metastasis in cervical cancer patients is achieved.
Patent Information
- Application Number
- CN202510671977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art has low sensitivity and accuracy in diagnosing lymph node metastasis in cervical cancer, resulting in many patients with CC lymph node metastasis unrecognized.
A calculation model integrating gene expression characteristics and microbial abundance characteristics was adopted, and a classification prediction model was constructed through the mining and feature screening of multiple sets of clinical sample data, combined with machine learning algorithms, and the risk of lymph node metastasis in patients with cervical cancer was evaluated.
It improves the prediction accuracy of lymph node metastasis in cervical cancer, achieves early warning of high-risk metastasis patients, and enhances the accuracy and sensitivity of diagnosis.
Smart Images

Figure CN120199322A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical diagnosis, and in particular to a diagnostic model for cervical cancer lymph node metastasis and a construction method thereof, as well as an application of the model in the preparation of a diagnostic product for cervical cancer lymph node metastasis. Background Art
[0002] Cervical cancer (CC) ranks fourth among female malignant tumors. It is a common malignant tumor of the female reproductive system and seriously threatens women's life and physical and mental health. One of the important factors affecting the prognosis of CC patients is lymph node metastasis. Studies have confirmed that lymph node metastasis in early CC will seriously affect its 5-year survival rate. Whether the lymph nodes are metastatic not only affects the prognosis of cervical cancer patients, but also plays a decisive role in individualized treatment. For early CC patients, systematic lymph node resection is often used to determine the occurrence of lymph node metastasis and provide a basis for postoperative supplementary treatment. However, it also causes a certain number of CC patients without lymph node metastasis to undergo unnecessary surgery, which in turn leads to unnecessary surgical risks. Therefore, accurately judging the lymph node status of CC patients and accurately assessing whether lymph node metastasis occurs are crucial for clarifying the staging, judging the prognosis, and formulating treatment plans. At present, the clinical diagnosis of CC lymph node metastasis is mainly based on the morphological characteristics of lymph nodes examined by magnetic resonance imaging and CT or PET / CT imaging. The most commonly used standard is whether the short diameter of the lymph node exceeds 10 mm. Although this method has a reasonable specificity, it has poor sensitivity and accuracy. As a result, a large proportion of patients with CC lymph node metastasis have not been identified. Therefore, it is still necessary to find new biomarkers (or key genes) that can predict CC lymph node metastasis. The intratumoral microbiome refers to the genome of microorganisms (bacteria, archaea, fungi and viruses) present in the tumor parenchyma and the microenvironment around the tumor. The discovery of intratumoral microorganisms has greatly increased people's understanding of the complexity of the tumor microenvironment. At first, these microorganisms were regarded as bystanders in the occurrence and development of tumors. However, with the deepening of research, people began to realize that microorganisms and their metabolites in the tumor microenvironment play a key role in the pathogenesis of cancer. They not only promote or inhibit the occurrence and development of tumors, but also significantly affect the metabolism and therapeutic effects of related drugs in the body, and thus gradually become new tumor treatment targets. In addition, studies have found that CC intratumoral microorganisms are related to CC metastasis and affect disease prognosis, but its specific mechanism of action has not yet been fully elucidated.
[0003] In recent years, with the emergence and rapid development of high-throughput sequencing technology, more and more CC genes and epigenetic features have been discovered. However, few studies have been conducted on intratumoral microbiomics based on clinical CC lymph node metastasis and CC primary site patient samples. Multi-omics research is a method to explore the interactions between multiple substances in biological systems, including transcriptomics, microbiomics, proteomics, metabolomics, etc. These substances jointly affect the phenotypes, traits, etc. of the life system. Each type of omics data usually provides a list of differential factors that may be related to diseases, and these data can be used as disease markers. Multi-omics data helps to explore different levels of systems biology from an overall perspective. Integrating different types of omics data may help to clarify the potential pathogenic changes leading to diseases, or can be used to identify potential therapeutic targets for further molecular research. Summary of the Invention
[0004] The object of the present invention is to provide a computational model that fuses gene expression features and microbial abundance features for evaluating the lymph node metastasis risk of cervical cancer patients. This model is based on multi-group clinical sample data, mines and screens multi-dimensional bioinformatics features, and constructs a classification prediction model in combination with machine learning algorithms to achieve early warning of high-risk metastatic patients.
[0005] To achieve the above object, the present invention provides the following technical solutions: On the one hand, the present invention provides a computational model whose input is the expression levels of specific gene markers (CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28) in cervical cancer samples, and the relative abundances of specific microbial genera (Dialister, Catonella, and Campylobacter), and the output is the lymph node metastasis risk score of the patient.
[0006] The model is obtained by training on a large sample data set. The model algorithms include but are not limited to random forest, logistic regression, LASSO regression, COX regression, artificial neural network, decision tree, support vector machine, naive Bayes, K-nearest neighbor algorithm, gradient boosting tree, Adaboost algorithm, XGBoost algorithm, LightGBM algorithm, CatBoost algorithm, multi-layer perceptron, convolutional neural network, recurrent neural network, long short-term memory network, gated recurrent unit network, etc. The robustness and generalization ability of the model are ensured through cross-validation and feature importance analysis.
[0007] The model of the present invention can further be implemented by a set of computer programs. When executed by a processor, the program includes the following steps: receiving the original data of the sample to be tested, performing data preprocessing and standardization, inputting the trained model, and outputting the metastasis risk score and prediction probability.
[0008] The present invention also provides a method for constructing the model, and the steps include: collecting cervical cancer tissue samples with and without lymph node metastasis, and respectively performing transcriptome and microbiome detections; using statistical and machine learning methods for differential analysis and feature screening; constructing a prediction model and an optional nomogram based on the screened key features.
[0009] Furthermore, the present invention provides a software system and a hardware device integrating the model. Users can input the detection data of patients, and the transfer risk score can be automatically output for clinical auxiliary diagnosis.
[0010] The evaluation model proposed by the present invention has one of the following significant technical advantages and / or clinical values: Multi-dimensional feature integration: The present invention first incorporates gene expression information highly related to cervical cancer metastasis and cervical-related microecological microbial genera into the modeling input together, captures the comprehensive biological characteristics of cancer progression, and improves the prediction accuracy.
[0011] High-throughput technology support: The model is based on multi-omics data (transcriptome and microbiome) of a large sample of real-world patients, and has both clinical representativeness and wide applicability.
[0012] Artificial intelligence algorithm-driven: Multiple machine learning algorithms are used for comparison and optimization in model construction, significantly improving the model's ability to distinguish between metastasis and non-metastasis, and supporting the model's migration application on different data sets.
[0013] Early warning: The model can assist doctors in achieving risk stratification before treatment.
[0014] Standardized evaluation tool: The present invention also constructs a diagnostic nomogram tool, which can intuitively display the contribution of each feature to the score, and provides an interpretable prediction result for clinicians.
[0015] System equipment can be integrated: By integrating computer programs and evaluation equipment, the integration of risk assessment can be realized, and it has the potential for industrial transformation and popularization and application.
[0016] The present invention proposes a machine learning model integrating molecular biology information and microecological information for predicting and risk assessing cervical cancer lymph node metastasis.
[0017] In the present disclosure, a "gene marker" refers to a gene whose expression level in tumor tissue is highly related to the disease state (such as metastasis).
[0018] A "microbial genus" refers to a microbial population at the taxonomic level of "genus" identified by, for example, microbial 16S rRNA or metagenomic detection.
[0019] "Abundance" refers to the microbial abundance, which represents the relative occurrence frequency of a genus in the total sample and is often expressed as a percentage.
[0020] "Nomogram" refers to a visualization tool used to display a regression model that can map multiple predictors to a score interval to estimate the result probability.
[0021] In order to identify the differential genes between MN samples and MP samples, the present invention analyzed the differential genes between MP and MN samples (MP group VS MN group) in the transcriptome dataset through the "DESeq2" package, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC| > 0.5, p-value < 0.05); subsequently, the R packages "ggplot2" and "ComplexHeatmap" were used to draw the volcano plot and heatmap of DEGs (the top 10 up-regulated genes with the largest change in |log2FC|: SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, TEKT4; the top 10 down-regulated genes: FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, PAK5).
[0022] The gene numbers of the top 10 up-regulated genes with the largest change in |log2FC| are respectively: SOHLH1 (ID: 402381), MUCL3 (ID: 283232), PNMA5 (ID: 114836), REG1A (ID: 5967), FGA (ID: 2243), PGC (ID: 5225), DSCAM-AS1 (ID: 102723407), IGHV3-22 (ID: 28450), LINC01320 (ID: 100507487), TEKT4 (ID: 150465).
[0023] The gene numbers of the top 10 down-regulated genes with the largest change in |log2FC| are FOLR1P1 (ID: 107986793), TLX1 (ID: 3195), MGAM2 (ID: 9342), KRT1 (ID: 3848), KRTDAP (ID: 200172), LINC00167 (ID: 100507065), MUCL1 (ID: 118430), LINC00457 (ID: 285237), BPIFC (ID: 653145), PAK5 (ID: 57144).
[0024] Based on the above analysis, 1057 up-regulated genes were obtained in the samples of patients with cervical cancer lymph node metastasis. Further, the top 10 up-regulated genes with the largest |log2FC| change include at least one of SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, and TEKT4.
[0025] Based on the above analysis, 317 down-regulated genes were obtained in the samples of patients with cervical cancer lymph node metastasis. Further, the top 10 down-regulated genes with the largest |log2FC| change include at least one of FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5.
[0026] Preferably, in step 4, in order to obtain key genes, bioinformatics analysis methods are carried out, including systematically performing correlation analysis, machine learning, ROC analysis, etc. on the differential genes.
[0027] The present invention relates to a method for evaluating the risk of lymph node metastasis in cervical cancer patients, in particular to a nomogram prediction model constructed based on key genes and microbiome characteristics and its construction and validation methods. This method combines the transcriptome expression information of patients with the compositional characteristics of the intestinal or cervical local microbiome, and improves the accuracy and practicability of cervical cancer clinical risk stratification through statistical modeling and visualization tools.
[0028] In a preferred embodiment of the present invention, first, mRNA expression profile data of cervical cancer tissues are obtained through transcriptome sequencing technology, and the DESeq2 software is used to perform differential expression analysis on patient samples with and without lymph node metastasis, and candidate genes significantly related to lymph node metastasis are screened out therefrom. Subsequently, combined with variable selection methods such as LASSO, a key gene set with high predictive value is further identified.
[0029] In the process of constructing the prediction model, it is preferred to divide the original sample data set into a training set and a validation set. Based on the training set, a multi-factor regression analysis method (preferably a Logistic regression model) is used to evaluate the correlation between the expression levels of the identified key genes and lymph node metastasis, and a nomogram is drawn based on the regression results. Preferably, the rms package in R language is used to construct the nomogram model. The nomogram can map the expression level of each key gene to a risk score, and predict the probability (Pr(Y)) of each patient having lymph node metastasis by calculating the total score (Total Points).
[0030] To improve the interpretability and predictive performance of the model, the present invention further uses multiple sets of key gene combinations (including any two, three, four, or all five genes) for nomogram modeling respectively, and selects the best combination by calculating the area under the ROC curve (AUC) of each combined model. The performance evaluation of the nomogram model also includes C-index (Harrell's concordance index), calibration curve (Calibration curve, preferably drawn using the rms package), decision curve analysis (DCA, preferably drawn using the ggDCA package), and receiver operating characteristic curve (ROC, preferably drawn using the pROC package) to comprehensively reflect the goodness of fit, clinical applicability, and discriminative ability of the model.
[0031] The present invention further preferably combines microbiome information, and obtains the composition data of the local cervix or intestinal microbiota of each patient through sequencing analysis. Based on the key microbial genera that have been identified as being closely related to lymph node metastasis (such as Dialister, Catonella, and Campylobacter), their relative abundances are extracted as modeling features, and together with the key gene expression values, they form combined features and are input into the model to construct a more comprehensive nomogram prediction tool. This combined model significantly improves the AUC value of the model, demonstrating the important role of microbiome data in predicting cervical cancer progression.
[0032] During the model construction process, the top layer of the nomogram is used to display the scores (Points) of each key gene and microbial genus. The corresponding scores are obtained according to the specific values of each variable and summed to obtain "Total Points", and the corresponding probability axis below is the quantitative probability of predicting the risk of lymph node metastasis.
[0033] The key genes of the present invention preferably include at least one of CARD9 (Gene ID: 64170), MNX1_AS2 (Gene ID: 100873971), MRAS (Gene ID: 22808), OLFML2A (Gene ID: 79589), and RPS28 (Gene ID: 6234), or any combination of two, three, or four of them, more preferably the complete combination of the above five genes; the key microorganisms of the present invention preferably include at least one of the key microbial genera that are closely related to lymph node metastasis, such as Dialister, Catonella, and Campylobacter genera, or any combination of two or three of them, more preferably the complete combination of the above three microbial genera. Brief Description of the Drawings
[0034] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0035] Figure 1A and Figure 1B are the volcano plot ( Figure 1A ), and the expression heatmap ( Figure 1B ) of the differentially expressed genes between the MP group and the MN group in the preferred Embodiment 1 of the present invention, drawn using the R packages "ggplot2" and "ComplexHeatmap"; among them, the volcano plot shows the TOP10 genes with the highest up- and down-regulation fold changes, with red indicating high expression, green indicating low expression, and gray indicating genes with no significant difference. The heatmap consists of two parts: the upper part is the density heatmap of the expression levels of the differentially expressed genes, showing the lines of five quantiles and the average value; the lower part is the expression heatmap of the differentially expressed genes. Each row represents the expression level of each gene in different group samples, and each column represents the expression levels of the top 10 up- and down-regulated DEGs in each sample. The color represents the magnitude of the gene expression level, and the darker the color, the higher the expression level of the DEGs (red for high expression and cyan for low expression).
[0036] Figure 2 : The histogram of differential microorganisms between the MP group and the MN group, where the X-axis is the LDA score, indicating the contribution degree of different genera to the differences between groups; the Y-axis lists the names of each genus. The larger the LDA value, the more significant the difference of the genus between different groups.
[0037] Figure 3 : The diagnostic nomogram of the model constructed by key genes and microorganisms, where the top layer of the nomogram represents the scores (Points) of each key gene and key microorganism. The corresponding scores are obtained according to the values of the key genes and microorganisms, and then the scores are added together to get "Total points". The higher the total score, the higher the predicted risk. The risk probability (Pr(Y)) calculated based on the total score is shown at the bottom.
[0038] Figure 4 : The calibration curve of the model nomogram in the training set, where the X-axis represents the predicted probability and the Y-axis represents the actual probability. The blue curve is the actual prediction result (Apparent), the black curve is the result after bias correction (Bias-corrected), and the dashed line is the ideal situation (Ideal), that is, the reference line where the model prediction is completely consistent with the actual situation. The p-value of the Hosmer-Lemeshow test is 0.726, indicating that the calibration performance of the model is good and does not deviate significantly from the ideal situation.
[0039] Figure 5 : The DCA curve of the model nomogram in the training set, where the X-axis represents the risk threshold and the Y-axis represents the net benefit. Curves of different colors represent different prediction models, including the nomogram model, the control lines of "All" and "None". The higher the net benefit of the curve, the greater the decision-making value of the model at the corresponding risk threshold. The dashed line represents the control situation without any intervention in all samples.
[0040] Figure 6 : The ROC curve of the diagnostic model in the training set. The ROC (Receiver Operating Characteristic) curve is used to evaluate the performance of the diagnostic model, where the horizontal axis is 1 - Specificity (false positive rate) and the vertical axis is Sensitivity (true positive rate). The gray dashed line represents the baseline of random prediction. The red curve is the ROC curve of the diagnostic model, and the area under the curve (AUC, Area Under the Curve) is 0.8924, indicating that the model has good discrimination ability.
[0041] Figure 7 : The flowchart of the computer system operation in an embodiment of the present invention.
[0042] Figure 8 : The structural block diagram of the intelligent device in an embodiment of the present invention. Specific embodiments
[0043] Example 1: Screening and modeling of key genes and microbial genera In this study, transcriptome and intratumoral microbiome sequencing data of cervical tissues from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis were used. Through bioinformatics methods, the molecular changes and pathogenesis of CC lymph node metastasis were deeply explored in all aspects. Based on a variety of bioinformatics analysis methods, the diversity, community composition and differential microorganisms of intratumoral microorganisms in CC lymph node metastasis were systematically explored, and their correlation analysis, machine learning and ROC analysis were performed with differential genes to screen out key genes.
[0044] Total sample RNA was isolated and purified using TRIzol (Thermofisher, 15596018) according to the manufacturer's protocol. Then, the quantity and purity of total RNA were quality controlled using NanoDrop ND-1000 (NanoDrop, Wilmington, DE, USA), and the integrity of RNA was detected using Bioanalyzer 2100 (Agilent, CA, USA); a concentration >50 ng / μL, RIN value >7.0, and total RNA >1 μg were required to meet the downstream experiment requirements. Poly(A)-tailed mRNA was specifically captured using oligo(dT) magnetic beads (Dynabeads Oligo(dT), cat.25-61005, Thermo Fisher, USA) through two rounds of purification. The captured mRNA was fragmented at high temperature using a magnesium ion fragmentation kit (NEBNextR Magnesium RNA Fragmentation Module, cat.E6150S, USA) at 94°C for 5 - 7 minutes. The fragmented RNA was reverse transcribed into cDNA using reverse transcriptase (Invitrogen SuperScriptTM II Reverse Transcriptase, cat.1896649, CA, USA). Then, E. coli DNA polymerase I (NEB, cat.m0209, USA) and RNase H (NEB, cat.m0297, USA) were used for second-strand synthesis to convert the DNA-RNA hybrid double-strand into a DNA double-strand. At the same time, dUTP Solution (Thermo Fisher, cat.R0133, CA, USA) was incorporated into the second strand to blunt the ends of the double-stranded DNA, and then an A base was added to each end to enable ligation with an adapter with a T base at the end. Magnetic beads were used to screen and purify the fragment sizes. UDG enzyme (NEB, cat.m0280, MA, US) was used to digest the second strand, and then PCR was performed - pre-denaturation at 95°C for 3 minutes, 98°C denaturation for a total of 8 cycles, 15 seconds each, annealing to 60°C for 15 seconds, extension at 72°C for 30 seconds, and finally extension at 72°C for 5 minutes to form a library with a fragment size of 300 bp ± 50 bp (strand-specific library). Finally, illumina NovaseqTM 6000 was used to perform paired-end sequencing on it according to the standard operation, and the sequencing mode was PE150.Valid Data that can be aligned to the reference genome can be defined as being aligned to exons, introns, and intergenic regions according to the regional information of the reference genome. Generally, for species with relatively well-annotated genomes (such as model species like humans and Arabidopsis thaliana), the percentage content of the sequencing sequences mapped to the exon regions should be the highest. Reads mapped to intron and intergenic regions may be due to precursor mRNA splicing events, ncRNAs (non-coding RNAs), imperfect genome annotation, DNA contamination, and background noise, etc. It was found that the content of all samples mapped to the exon regions reached over 90%, indicating that the sequencing data quality was good and the sequencing results were reliable.
[0045] To determine whether there are clusters or outliers in the samples, the FactoMineR package was used to perform principal component analysis (PCA) for data dimensionality reduction on the transcriptome sequencing dataset, in order to identify the dispersion of samples in the metastasis-negative MN group and metastasis-positive MP group through patternization and visualization. It was found that PC1 explained 22.37% of the variance, which means that the first principal component can capture relatively high variability in the data. The combined variance explained by PC1 and PC2 was 29.97% (22.37% + 7.6%), which means that these two principal components can explain nearly one-third of the variability in the dataset.
[0046] To identify the differentially expressed genes between MN samples and MP samples, the "DESeq2" package was used to analyze the differentially expressed genes between MP and MN samples (MP group VS MN group) in the transcriptome dataset, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC|>0.5, p-value<0.05); subsequently, the R packages "ggplot2" and "ComplexHeatmap" were used to draw volcano plots and heatmaps of the DEGs, as Figure 1A and Figure 1BAs shown below. The top 10 upregulated genes with the largest change in |log2FC|: SOHLH1 (spermatogenesis- and oogenesis-specific basic helix-loop-helix protein 1), MUCL3 (mucin-like protein 3), PNMA5 (paraneoplastic Ma antigen 5), REG1A (regenerating gene 1A), FGA (fibrinogen alpha chain gene), PGC (transcriptional coactivator), DSCAM-AS1 (Down syndrome cell adhesion molecule antisense RNA 1), IGHV3-22 (immunoglobulin heavy chain variable region 3-22), LINC013200 (long non-coding RNA 013200), TEKT4 (Tektin 4); the top 10 downregulated genes: FOLR1P1 (folate receptor 1 pseudogene 1), TLX1 (T-cell leukemia homeobox gene 1), MGAM2 (maltase-dextrinase 2), KRT1 (keratin 1), KRTDAP (keratin-related protein), LINC00167 (long non-coding RNA 00167), MUCL1 (breast cancer differentiation-related protein 1), LINC00457 (long non-coding RNA 00457), BPIFC (BPI fold domain family member C), PAK5 (p21-activated kinase 5).
[0047] Identification of differential microorganisms To explore the main microbial flora causing dynamic differences in the microbiota and identify differential microorganisms at various levels in different groups, we used linear discriminant analysis (LDA) and effect size (LEfSe) to identify the main differential microorganisms at the genus level between the MN and MP groups for analysis (p < 0.05, LDA score (log10) > 2). The role of LEfSe analysis is to discover microorganisms with significant differences between different groups. Through the LDA value distribution bar chart (as Figure 2 shown), it visually shows the genera that are significantly enriched at different levels in the three groups and have a key impact on grouping. There are a total of 6: Fusobacterium, Sneathia, Dialister, Campylobacter, Streptophyta_Unknown_family_Unknown_genus123, Catonella, which are recorded as differential microorganisms.
[0048] Subsequently, based on the union of key genes and key microorganisms, the AUC values of the ROC curves of the nomogram models constructed with different combinations were calculated to select the gene-microorganism combination with the highest AUC value. It was found that the diagnostic model constructed with the combination of CARD9, MNX1_AS2, MRAS, OLFML2A, RPS28, Dialister, Catonella, and Campylobacter had the highest AUC value (0.892). The nomogram consists of "scores" and "total scores". The former represents the scores of each key gene, and the latter represents the sum of the scores of all key genes. The higher this value, the higher the probability of lymph node metastasis in CC (as Figure 3 shown), and the "rms" package was used to plot the calibration curve in the analysis (as Figure 4 shown), the "ggDCA" package was used to plot the DCA decision curve (as Figure 5 shown), and the "pROC" package was used to plot the ROC curve to evaluate the goodness of fit of the model (as Figure 6 shown). The results showed that the nomogram had good predictive ability (HL test p.value > 0.05, the benefits of the DCA curve were basically above ALL and NONE, and the AUC value = 0.892).
[0049] In this example, the Logistic regression model was constructed using the rms package, and the nomogram of the regression model was plotted using the regplot package. The Logistic regression model was used, with the log2(x + 1) values of the CPM of gene expression and the log2(x + 1) values of the relative abundances of microbial genera. The expression formula of this model is as follows:
[0050] The OR values of each factor in the diagnostic model were calculated as shown in the following table.
[0051] Table 1 OR values of each factor in the model
[0052] Example 2: Establishment and verification of a computational model for evaluating the risk of lymph node metastasis in cervical cancer In this example, cervical tissues of 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis in Example 1 were used for transcriptome sequencing data. Differential expression analysis was performed by DESeq2 to screen candidate genes with significant expression differences between metastatic and non-metastatic samples. The LASSO regression model was further used for variable selection, and 5 gene markers with high predictive value were finally determined: CARD9, MNX1_AS2, MRAS, OLFML2A and RPS28; 3 microbial genera: Dialister, Catonella and Campylobacter. The expression levels of the above 5 genes and the relative abundance of microorganisms of 3 microbial genera were used as feature variables, and the random forest algorithm was used to train the binary classification model. The data was divided into a training set and a test set in a ratio of 7:3. In the test set, the AUC (area under the curve) of the model was 0.89, indicating that the model has extremely high predictive efficiency in distinguishing whether cervical cancer patients have lymph node metastasis. To ensure the stability and accuracy of the model performance evaluation, the model performance is evaluated using 10-fold cross-validation. The 10-fold cross-validation described in this embodiment is performed within the original data consisting of the training set and the test set, that is, a nested cross-validation (nested CV) strategy is adopted. In each round of training, the original data is first divided into 10 subsets, nine as training sets, and one as a validation set for evaluating model performance. Each of the ten subsets is used as a validation set in turn, and the rest are used as sub-training sets for model training, and the performance indicators of each round are recorded. Finally, the results of 10 rounds are averaged to evaluate the model performance. At the same time, in order to further confirm the robustness of the model under different random partitions, a repeated 10-fold cross-validation strategy (repeated 10-fold CV, repeated 5 times) is adopted, and the calculated average AUC is 0.90, the standard deviation is ± 0.037, and the 95% confidence interval is [0.88, 0.95], showing that the prediction performance of the model in the training sample is stable and the risk of overfitting is low.
[0053] Example 3: External independent sample verification of model effectiveness and comparative analysis with traditional methods In order to further verify the reliability and generalization ability of the calculation model for evaluating the risk of cervical cancer lymph node metastasis described in Example 2, this example introduces a group of external independent verification samples, and compares and analyzes the model prediction results with traditional imaging evaluation methods.
[0054] A total of 24 independent cervical cancer patient samples were included in this example, including 12 cases in the metastasis group and 12 cases in the non-metastasis group. All samples were not involved in the model training process, and the source was the same as that in Example 1. The sample collection and pretreatment process included RNA extraction, preparation of dPCR premix with specific primers, probes, dNTPs, and polymerase, generation of droplets, thermal cycle amplification, total DNA extraction, droplet detection, quality control, and standardization processing.
[0055] For the above samples, the absolute expression levels of 5 key genes (CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28) and the absolute quantification of 3 key microbial genera (Dialister, Catonella, and Campylobacter) were detected using the dPCR absolute quantification technique. Since the training model was constructed based on RNA-seq counts, to ensure data consistency, a standardization conversion strategy (such as Z-score standardization) was adopted in the external validation, so that the dPCR expression values could be input into the model for prediction.
[0056] The standardized expression data were input into the random forest model trained in Example 2, and the model automatically output the lymph node metastasis prediction probability for each patient. The final results were as follows: AUC (ROC): 0.884; Accuracy: 87.5%; Sensitivity: 91.7%; Specificity: 83.3%; Positive Predictive Value (PPV): 84.6%; Negative Predictive Value (NPV): 90.9%. Compared with the traditional imaging evaluation method (CT / MRI combined judgment, with an accuracy of about 71.0%, a sensitivity of about 65.0%, and a specificity of about 76.0%), this model showed higher prediction accuracy and sensitivity in external independent samples. The above results indicate that the evaluation model described in Example 2 can still maintain good prediction performance in new independent samples, has strong generalization ability and stability, and is superior to the existing traditional imaging evaluation methods, with high clinical application potential.
[0057] Example 4: Implementation of a computer system based on the model of the present invention In this example, a computer system was constructed to run the above risk prediction model, and the system included the following modules: Data Import Module: Used to read the standardized gene expression and microbial abundance data of patient samples, supporting formats such as CSV and Excel. Preprocessing Module: Includes functions such as missing value handling and normalization transformation to ensure the consistency of input data. Model Calculation Module: Loads the pre-trained random forest model file, performs inference on the input data, and outputs the predicted risk scores and probabilities for each sample. Visualization Module: Displays the prediction results in the form of bar charts, heatmaps, etc., facilitating doctors to quickly understand the results. Interface Module: Can be docked with the hospital information system (HIS) or laboratory data system (LIMS) to achieve automatic data transfer and automatic triggering of risk assessment. This system is deployed on the hospital server or private cloud environment. The front-end interface is developed using the Django framework, and the back-end model is implemented using the scikit-learn library in Python (see Figure 7 ).
[0058] Example 5: Integrated Lymph Node Metastasis Risk Assessment Device The inventor of the present invention designed an intelligent device for assessing the risk of cervical cancer lymph node metastasis, mainly including: Detection Module: Construct a Panel by designing multiplex primers and probes for the above 5 genes and 3 genera of microorganisms. Based on the exclusive Panel, the dPCR detection system can integrally detect the expression levels of 5 genes and 3 genera of microorganisms in cervical cancer tissue samples. Data Acquisition Module: Connects to the detection module and transmits the expression data to the main control system in real time. Main Control Processing Module: The normalization transformation strategy (Z-score normalization) converts the dPCR expression values into model-standardized data. The embedded processor runs the model program of the present invention, integrates the Linux system, has a memory of 8GB, and is equipped with a dedicated operation acceleration chip (such as the ARM Cortex-A72 series). Output Display Module: Built-in touch screen display interface, which can display the prediction score, metastasis probability, result suggestions, and visualization reports in real time. Communication Module: Supports WiFi, Bluetooth, USB, and Ethernet interfaces, facilitating communication with the hospital database or the host computer. After the device is started, the operator only needs to import the sample, click the detection and prediction buttons, and the whole process assessment can be completed, providing a timely auxiliary basis for clinical surgical decisions (see Figure 8 ).
[0059] The above are only the embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A computational model for assessing the risk of lymph node metastasis in cervical cancer, the model taking the expression of multiple gene markers and the abundance of microorganisms based on microbial genera as input, and outputting a metastasis risk score for the patient; the gene markers are selected from the group consisting of the following gene markers: CARD9, MNX1_AS2, MRAS, OLFML2A and RPS28; the microbial genera include Dialister, Catonella and Campylobacter.
2. The computational model according to claim 1, wherein the model is established based on a machine learning algorithm, and the algorithm is selected from random forest, logistic regression, LASSO regression, COX regression or neural network algorithm.
3. A computer program, which, when executed by a processor, is used to implement the evaluation function of the computing model described in claim 1 or 2, the program comprising the following steps: a) receiving gene expression and microbial abundance data of the sample to be tested; b) Perform data preprocessing and standardization; c) input the standardized data into the prediction model; d) Output the lymph node metastasis risk score and predicted probability corresponding to the sample.
4. A computer system for assessing the risk of lymph node metastasis of cervical cancer, comprising a processor, a memory, and a program module stored in the memory and executed by the processor, wherein the program module implements the steps as described in claim 3 to achieve risk assessment of a target sample.
5. A method for constructing a diagnostic model for cervical cancer lymph node metastasis, characterized in that, The following steps are involved: Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 2: Perform gene expression detection and microbial identification on the tissue samples obtained in step 1; Step 3: Based on the gene expression detection and microbial identification data obtained in step 2, the gene expression differences and microbial presence differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis are compared; Step 4: Screen out key genes and key microorganisms based on bioinformatics analysis methods; Step 5: Construct a prediction model based on the selected key genes and key microorganisms, and optionally draw a diagnostic nomogram of key genes and key microorganisms, obtain corresponding scores according to the values of key genes and key microorganisms, and then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis based on the total score. Wherein, in step 4, the key genes include at least 1, 2, 3, 4 or 5 of CARD9, MNX1_AS2, MRAS, OLFML2A and RPS28; the microbial genera include Dialister, Catonella and Campylobacter.
6. The method according to claim 5, wherein in step 5, the risk probability of cervical cancer lymph node metastasis is estimated based on the total score obtained from the diagnostic nomogram of key genes and key microorganisms, and the higher the total score, the higher the risk of cervical cancer lymph node metastasis is predicted.
7. An evaluation device for evaluating the risk of cervical cancer lymph node metastasis, characterized in that, include: A data acquisition device for acquiring detection data of an evaluation object, where the detection data is the level value of a biomarker in a biomarker for cervical cancer lymph node metastasis detected from a sample of an evaluation object suspected of having cervical cancer lymph node metastasis; A data processing device for calculating an index value of a biomarker in a biomarker for cervical cancer lymph node metastasis based on the detection data of the evaluation object, and then determining whether the calculated index value of the biomarker is within the risk range of the corresponding biomarker; The biomarker includes a gene biomarker and a microbial biomarker, where the gene biomarker is CARD9, MNX1_AS2, MRAS, OLFML2A, and RPS28; the microbial biomarker includes microorganisms of the genera Dialister, Catonella, and Campylobacter.
8. The evaluation device according to claim 7, characterized in that Wherein the device further includes a detection device for detecting the level value of the biomarker.
9. The evaluation device according to claim 7, characterized in that, The data acquisition device is an input device for inputting data or a communication device for reading data from an external data storage device or a storage device of the detection device through an interface.
10. The evaluation device according to claim 9, characterized in that, When a communication device is adopted, its corresponding external data storage device is a data memory on a device for measuring biomarkers.
Citation Information
Patent Citations
Method for checking cervical cancer using living cell technology
CN101059505A
HNC (Head and Neck Cancer) prognosis biomarker based on lymph node microbial flora and application of HNC prognosis biomarker
CN113684242A
Group of markers for predicting lymph node metastasis risk of cervical cancer
CN117070625A
TCT image quality control method and system based on deep learning
CN117952933A
Equipment, method and system for diagnosing lymphatic metastasis of cervical cancer and application of equipment, method and system
CN118471335A
Cited By
Esophageal cancer immune data prediction device and equipment based on tumor markers and storage medium
CN120932727A