Methods and models for diagnosing early cervical cancer lymph node metastasis based on transcriptome sequencing
By constructing a risk assessment model for lymph node metastasis of cervical cancer based on machine learning, using transcriptome sequencing to screen key genes, and establishing a risk assessment model for lymph node metastasis of cervical cancer, the problem of low sensitivity and accuracy of imaging examinations is solved, and early and accurate diagnosis of lymph node metastasis and individualized treatment are achieved.
Patent Information
- Application Number
- CN202510671066.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the prior art, the diagnosis of lymph node metastasis of cervical cancer mainly relies on imaging examinations, with low sensitivity and accuracy, resulting in a large number of patients not being identified, and efficient, non-invasive and quantifiable diagnostic methods are urgently needed.
A cervical cancer lymph node metastasis risk assessment model is constructed based on machine learning algorithms. Using the expression data of multiple gene markers, key genes with significant differences are screened out through transcriptome sequencing, a diagnostic model is established and a nomogram is drawn, and a lymph node metastasis risk score is output.
It improves the accuracy and sensitivity of lymph node metastasis diagnosis, can predict lymph node metastasis early, reduce unnecessary surgical risks, and provide individualized treatment recommendations.
Smart Images

Figure CN120183512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical diagnosis, and in particular to a diagnostic model for cervical cancer lymph node metastasis, a construction method thereof, and an application thereof in preparing a diagnostic product for cervical cancer lymph node metastasis. Background Art
[0002] Cervical cancer (CC) ranks fourth among female malignancies and is a common malignancy of the female reproductive system, posing a serious threat to women's life and physical and mental health. Lymph node metastasis is a key factor influencing CC patients' prognosis. Studies have confirmed that the presence of lymph node metastasis in early-stage CC significantly impacts 5-year survival. The presence of lymph node metastasis not only influences the prognosis of cervical cancer patients but also plays a decisive role in individualized treatment. Systematic lymph node resection is often performed for early-stage CC to confirm the presence of lymph node metastasis and provide a basis for supplemental postoperative treatment. However, this also results in a certain number of CC patients without lymph node metastasis undergoing unnecessary surgery, leading to unnecessary surgical risks. Therefore, accurately determining the lymph node status of CC patients and accurately assessing the presence of lymph node metastasis are crucial for staging, prognosis, and treatment planning. Currently, the clinical diagnosis of CC lymph node metastasis is primarily based on morphological examination of lymph nodes using magnetic resonance imaging, CT, or PET / CT imaging. The most commonly used criterion is whether the short diameter of the lymph node exceeds 10 mm. While this method has reasonable specificity, it suffers from poor sensitivity and accuracy. This results in a large proportion of patients with CC lymph node metastasis remaining unidentified. Therefore, the identification of new biomarkers (or key genes) that can predict CC lymph node metastasis remains a challenge.
[0003] In recent years, with the emergence and rapid development of high-throughput sequencing technology, more and more CC genes and epigenetic features have been discovered. However, few studies have constructed accurate diagnostic models based on transcriptome sequencing of clinical CC lymph node metastases and CC primary lesion patient samples. Summary of the Invention
[0004] Cervical cancer is a common gynecological malignancy, and lymph node metastasis is a key factor influencing treatment options and prognosis. Traditional assessment methods rely on imaging examinations and surgical pathology analysis, but these methods suffer from low accuracy, invasiveness, and delayed diagnosis. Therefore, there is an urgent need for an efficient, non-invasive, and quantifiable method for early prediction and assessment of the risk of lymph node metastasis in patients with cervical cancer.
[0005] In the prior art, the diagnosis of pelvic lymph node metastasis in cervical cancer primarily relies on imaging, which has poor sensitivity and accuracy. To overcome these shortcomings and identify novel biomarkers (or key genes) that can predict CC lymph node metastasis, the present invention relates to the fields of bioinformatics and tumor-assisted diagnosis, and more specifically to a computational model, construction method, computer program, computer system, and related evaluation equipment for assessing the risk of cervical cancer lymph node metastasis.
[0006] The present invention provides a computational model for assessing the risk of lymph node metastasis in cervical cancer. The model uses the expression data of multiple gene markers as input variables and outputs a lymph node metastasis risk score for each subject. The gene markers are selected from the following gene panel: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. These genes were screened using bioinformatics analysis methods and exhibit significant differential expression between metastatic and non-metastatic populations, demonstrating strong predictive power.
[0007] The model is built based on a machine learning algorithm, and methods such as random forest, logistic regression, LASSO regression, COX regression or neural network can be used to perform feature modeling on the above genes, thereby improving the performance of discriminating the possibility of lymph node metastasis.
[0008] The present invention also provides a computer program, system, and evaluation method based on this model. The program can be deployed in an evaluation system to automate the entire process from sample input to risk score output. The steps include receiving gene expression data from raw samples, normalizing the data, inputting it into the constructed model for inference, and outputting a metastasis risk probability value.
[0009] The present invention also proposes a method for constructing the model, which collects tissue samples from a large number of cervical cancer patients (including metastatic and non-metastatic groups) and conducts high-throughput gene expression testing. Then, through differential analysis and model training steps, key genes related to metastasis prediction are systematically screened out, and finally a quantitative scoring model (such as a nomogram model) is established to estimate the individual's metastasis risk.
[0010] To achieve the engineering application of this technical solution, the present invention also provides an assessment device, comprising a data acquisition device, a data processing module, and an optional detection device, capable of completing the marker data collection, analysis, and risk assessment process. This assessment device can be used in hospitals, laboratories, or other biological sample testing scenarios, facilitating rapid assessment of lymph node metastasis risk, intelligent decision-making, and personalized treatment recommendations.
[0011] More specifically, the present invention provides a method for constructing a diagnostic model for cervical cancer lymph node metastasis, comprising the following steps:
[0012] Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis;
[0013] Step 2: Perform gene expression testing on the tissue samples obtained in step 1;
[0014] Step 3: Based on the gene expression detection data obtained in step 2, the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis are compared;
[0015] Step 4: Screen key genes based on bioinformatics analysis methods;
[0016] Step 5: Draw a diagnostic nomogram for key genes, obtain corresponding scores based on the values of key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis based on the total score.
[0017] The present invention is different from traditional imaging examinations. It performs molecular testing on cervical tissue rather than imaging testing on lymph nodes. Therefore, regardless of whether the lymph nodes are enlarged in the image, the present invention can diagnose and predict cervical cancer lymph node metastasis with higher accuracy.
[0018] The present invention establishes a diagnostic model for cervical cancer lymph node metastasis based on gene expression detection. Preferably, gene expression detection includes transcriptional level detection and protein level detection. Gene expression detection methods include, but are not limited to, molecular biology techniques such as PCR and in situ hybridization; immunological techniques such as ELISA and immunofluorescence analysis; sequencing techniques such as RNA sequencing, Sanger sequencing, and transcriptome sequencing; bioinformatics methods such as microarray technology and protein chips; or mass spectrometry analysis.
[0019] A more preferred gene expression detection scheme of the present invention is to detect gene transcriptomes. Detection of relevant gene transcriptomes more directly reflects the gene expression status and differences in cervical cancer patient samples. Since the transcriptome is a direct product of gene expression, transcriptome detection can accurately understand the expression levels of each gene in cervical cancer patient samples, helping to reveal the changing patterns of gene activity, thereby serving the purpose of auxiliary diagnosis. Detection of relevant gene transcriptomes can predict disease risk earlier. Compared with protein detection, transcriptome detection may be able to predict the occurrence of cervical cancer lymph node metastasis earlier, predict disease risk, capture earlier disease signals, and timely adjust prognosis and treatment plans. Detection of relevant gene transcriptomes has higher sensitivity and can detect trace changes in gene expression. It is particularly important for identifying low-abundance genes or rare transcriptomes, and helps to discover potential disease-related genes. Transcriptome detection provides more comprehensive information. Transcriptome detection can provide comprehensive information about gene expression patterns, including expression levels of different transcriptomes, splicing patterns, etc. This information helps to deeply understand the function and regulatory mechanism of genes.
[0020] Preferably, the method for constructing a diagnostic model for cervical cancer lymph node metastasis of the present invention is based on a gene sequencing method and comprises the following steps:
[0021] Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis;
[0022] Step 2: Perform transcriptome sequencing on the tissue sample obtained in step 1 to obtain transcriptome sequencing data;
[0023] Step 3: Based on the transcriptome sequencing data obtained in step 2, compare the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis;
[0024] Step 4: Screen key genes based on bioinformatics analysis methods;
[0025] Step 5: Draw a diagnostic nomogram for key genes, obtain corresponding scores based on the values of key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis based on the total score.
[0026] Any of the above is preferably that, in step 3, in order to compare the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis, the preferred embodiment of the present invention is:
[0027] Preferably, principal component analysis (PCA) for data dimensionality reduction is performed on the transcriptome sequencing dataset using the FactoMineR package.
[0028] To determine whether there are clusters or outliers in the samples, the present invention uses the FactoMineR package to perform principal component analysis (PCA) on the transcriptome sequencing dataset for data dimensionality reduction. This allows for the identification of discrete cases of metastasis-negative (MN) and metastasis-positive (MP) samples through patterning and visualization.
[0029] Principal component analysis (PCA) for data dimensionality reduction on transcriptome sequencing datasets using the FactoMineR package is a common practice in bioinformatics. It aims to simplify data analysis by reducing data dimensionality while preserving as much of the original data information as possible. PCA results include PC1 (the first principal component) and PC2 (the second principal component). PC1 is the axis with the largest variance in the PCA transformation and represents the direction of greatest variation in the data. In other words, it is the direction with the greatest variance in the original dataset and captures the most data information (variance). Because PC1 captures the greatest variation in the data, it is often the most important principal component, explaining the largest proportion of the total variance in the dataset. When visualizing PCA results, PC1 is often used as the horizontal axis to show the distribution of samples in the principal component space. PC2 is the principal component with the second largest eigenvalue, ensuring orthogonality (i.e., perpendicular orientation) to PC1. It represents the second-largest direction of variation in the data and captures the largest amount of variance remaining after PC1.
[0030] The present invention used the FactoMineR R package to load a transcriptome sequencing dataset of cervical tissue samples, perform PCA analysis, and view an overview of the PCA results, obtaining information such as the proportion of variance explained by the principal components. By patterning and visualizing the discrete distribution of samples with no metastasis (MN) and positive metastasis (MP), the results showed that PC1 explained 22.37% of the variance, indicating that the first principal component was able to capture a relatively high level of variability in the data. The variance explained by PC1 and PC2 combined was 29.97% (22.37% + 7.6%), indicating that these two principal components explained nearly one-third of the variability in the dataset.
[0031] Furthermore, the present invention uses the DESeq2 package to analyze differentially expressed genes between MP and MN samples (MP group vs. MN group) in a transcriptome dataset. The DESeq2 package is a common bioinformatics analysis process for analyzing differentially expressed genes between transcriptome datasets.
[0032] To identify differentially expressed genes (DEGs) between MN and MP samples, the DESeq2 package was used to analyze the differentially expressed genes (DEGs) between MP and MN samples (MP group vs MN group) in the transcriptome dataset (1374 in total: 1057 upregulated and 317 downregulated in the MP group) (threshold: |log2FC|>0.5, p-value<0.05). Subsequently, the R packages "ggplot2" and "ComplexHeatmap" were used to plot the volcano plot and heat map of the DEGs (the top 10 upregulated genes with the largest changes in |log2FC| were SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, and TEKT4; the top 10 downregulated genes were FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5).
[0033] The top 10 upregulated genes with the largest changes in |log2FC| were SOHLH1 (ID: 402381), MUCL3 (ID: 283232), PNMA5 (ID: 114836), REG1A (ID: 5967), FGA (ID: 2243), PGC (ID: 5225), DSCAM-AS1 (ID: 102723407), IGHV3-22 (ID: 28450), LINC01320 (ID: 100507487), and TEKT4 (ID: 150465).
[0034] The top 10 genes with the largest downregulated changes in |log2FC| were FOLR1P1 (ID: 107986793), TLX1 (ID: 3195), MGAM2 (ID: 9342), KRT1 (ID: 3848), KRTDAP (ID: 200172), LINC00167 (ID: 100507065), MUCL1 (ID: 118430), LINC00457 (ID: 285237), BPIFC (ID: 653145), and PAK5 (ID: 57144).
[0035] Based on the above analysis, 1057 upregulated genes in samples of patients with cervical cancer lymph node metastasis were obtained. Further, the upregulated TOP10 genes with the largest |log2FC| changes included at least one of SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, and TEKT4.
[0036] Based on the above analysis, 317 down-regulated genes in samples of patients with cervical cancer lymph node metastasis were obtained. Further, the down-regulated genes with the largest |log2FC| changes included at least one of FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5.
[0037] Any of the above is preferably that in step 4, in order to obtain key genes, a bioinformatics analysis method is performed, including systematically performing correlation analysis, machine learning, ROC analysis, etc. on the differential genes.
[0038] Nomograms constructed based on key genes are often used to assess prognosis in fields such as oncology and medicine.
[0039] Preferably, a nomogram model is constructed based on the transcriptome sequencing data set obtained in step 2 and the gene expression difference data obtained by the "DESeq2" analysis in step 3.
[0040] Preferably, the data set is divided into a training set and a validation set, and a multivariate regression analysis is used on the training set to analyze the relationship between the key gene transcriptome detection data and lymph node metastasis. The multivariate regression analysis is preferably a logistic regression model.
[0041] Preferably, a nomogram is drawn using statistical software based on the results of the regression analysis. Preferably, the nomogram function is used, such as the rms package in the R language. The nomogram visualizes the results of the regression analysis, facilitating the assessment of whether cervical cancer patients have lymph node metastasis.
[0042] Preferably, the nomogram model is evaluated using a validation set to assess the predictive ability and accuracy of the model. Preferably, the evaluation method includes calculating a C index, drawing a calibration curve or a receiver operating characteristic (ROC) curve, etc.
[0043] It is further preferred that, in the process of constructing the nomogram, the AUC value of the ROC curve of the nomogram model constructed by different key gene combinations is calculated to select the gene combination with the highest AUC value; in the analysis, "rms" is used to draw the calibration curve, "ggDCA" is used to draw the DCA decision curve, and "pROC" is used to draw the ROC curve to evaluate the fit of the model.
[0044] In a preferred embodiment of the present invention, the AUC values of the receiver operating characteristic (ROC) curves of different combinations of key genes were calculated to select the gene combination with the highest AUC. The results showed that the diagnostic model constructed with the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 had the highest AUC value (0.8704). The nomogram consists of a "score" representing the score for each key gene and a "total score" representing the sum of all key gene scores. A higher score indicates a higher probability of lymph node metastasis in CC. The model fit was assessed using the "rms" calibration curve, the "ggDCA" decision curve, and the "pROC" ROC curve. Results showed that the nomogram had strong predictive ability (HL test p-value > 0.05, the DCA curve yielded a higher return than the ALL and NONE groups, and an AUC value of 0.8704).
[0045] Furthermore, the top layer of the nomogram represents the score (Points) for each key gene. The corresponding score is obtained based on the key gene's numerical value (e.g., expression data obtained from transcriptome sequencing of key genes), and then the scores are summed to obtain the "Total Points." A higher total score indicates a higher predicted risk. The bottom layer displays the risk probability (Pr(Y)) calculated based on the total score.
[0046] Furthermore, the key genes include at least any one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0047] The gene numbers of the key genes are: CARD9 (ID: 64170), CFL1P1 (ID: 100129361), GRASLND (ID: 100507642), MNX1_AS2 (ID: 100873971), MRAS (ID: 22808), OLFML2A (ID: 79589), and RPS28 (ID: 6234).
[0048] Furthermore, the key genes include a combination of any two of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0049] Furthermore, the key genes include a combination of any three of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0050] Furthermore, the key genes include a combination of any four of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0051] Furthermore, the key genes include a combination of any five of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0052] Furthermore, the key genes include a combination of any six of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0053] Furthermore, the key genes are a combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0054] The most preferred diagnostic model of the present invention is a detection model constructed from multiple genes. In a preferred embodiment of the present invention, a nomogram function is used to generate a nomogram. The transcriptome expression of seven key genes (CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28) is used as predictor variables. The transcriptomes of these seven key genes are sequenced to obtain their expression values, and a scoring scale is created based on these seven key genes. At the top level of the nomogram, each key gene has a corresponding scoring scale. Based on the key gene's expression value (expression data obtained from transcriptome sequencing), a corresponding score is assigned on the corresponding scale. These scores are then summed to obtain a "Total Points" score. Furthermore, the total score is used to estimate the risk of cervical cancer lymph node metastasis. A higher total score indicates a higher predicted risk of cervical cancer lymph node metastasis. Patients with a total score > -0.011 have a higher risk of cervical cancer lymph node metastasis and should subsequently undergo more aggressive clinical treatment.
[0055] In a preferred embodiment of the present invention, transcriptome sequencing was performed on cervical tissue obtained through cervical biopsy. A diagnostic model was constructed using a marker combination consisting of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. A diagnostic nomogram was generated, and the model fit was assessed using the "rms" calibration curve, the "ggDCA" decision curve, and the "pROC" receiver operating characteristic (ROC) curve. The results showed that the nomogram had good predictive ability (HL test p-value > 0.05, the DCA curve yielded returns generally above those of ALL and NONE, and the AUC value was 0.8704). Specifically, the Hosmer-Lemeshow test p-value for the model predictions versus the actual results was 0.315, indicating good calibration performance and no significant deviation from the ideal. The receiver operating characteristic (ROC) curve was used to evaluate the performance of the diagnostic model. The area under the curve (AUC) was 0.8704, indicating that the model had good discriminatory ability.
[0056] This invention also provides a diagnostic nomogram for diagnosing cervical cancer lymph node metastasis. Transcriptome gene sequencing was performed on cervical cancer tissue samples, and 115 candidate genes were subjected to LASSO regression analysis with 10-fold cross-validation using the R language glmnet package. Lasso regression constructs a penalty function to generate a more refined model, compressing some regression coefficients, reducing data dimensionality, and avoiding multicollinearity and overfitting in the multivariate regression model. Results showed that nine candidate genes (CARD9, CFL1P1, GRASLND, LHX3, MNX1-AS2, MRAS, OLFML2A, PAX3, and RPS28) passed the LASSO regression, designated as candidate key genes. Based on these key genes, a nomogram model was constructed using the R language "rms" package, version 6.5.0, to predict the probability of cervical cancer lymph node metastasis. The nomogram model was composed of seven markers, namely CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, to form a diagnostic model. The "regplot" package was used to draw the nomogram of the retrospective model. The calibration curve was drawn using "rms", the DCA decision curve was drawn using "ggDCA", and the ROC curve was drawn using "pROC" to evaluate the goodness of fit of the model.
[0057] Preferably, any of the above items is that the cervical tissue sample is a cervical biopsy sample, and the transcriptome sequencing result includes a cervical tissue transcriptome sequencing result.
[0058] The present invention also provides a cervical cancer lymph node metastasis biomarker combination, wherein the cervical cancer lymph node metastasis biomarker includes at least one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0059] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarker combination is used to diagnose cervical cancer lymph node metastasis.
[0060] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarkers include a combination of any two of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0061] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarkers include a combination of any three of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0062] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarkers include a combination of any four of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0063] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarkers include a combination of any five of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0064] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarkers include a combination of any six of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0065] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarker is a combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0066] Any of the above items is preferably, through the expression value of the cervical cancer lymph node metastasis biomarker gene transcriptome, using the diagnostic nomogram obtained by the present invention, obtaining the corresponding score according to the value of each cervical cancer lymph node metastasis biomarker in the diagnostic nomogram, and then adding up the scores to obtain the total score, and calculating the risk probability of cervical cancer lymph node metastasis based on the total score. The higher the total score, the higher the predicted risk of cervical cancer lymph node metastasis.
[0067] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarker is a transcriptome of the cervical cancer lymph node metastasis biomarker gene, and expression data of the cervical cancer lymph node metastasis biomarker gene is obtained by sequencing the transcriptome of the cervical cancer lymph node metastasis biomarker gene.
[0068] The present invention also provides a detection product for diagnosing cervical cancer lymph node metastasis, which is a detection reagent for any of the cervical cancer lymph node metastasis biomarkers described above.
[0069] Preferably, any of the above items is that the cervical cancer lymph node metastasis biomarker is any one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, or any two, or any three, or any four, or any five, or any six, or a combination of seven.
[0070] Preferably, any of the above items is that the detection reagent for the cervical cancer lymph node metastasis biomarker can be a detection reagent for its gene level or a detection reagent for its protein level.
[0071] Any of the above is preferably that the detection reagents for the cervical cancer lymph node metastasis biomarkers include but are not limited to molecular biology detection reagents, such as PCR, in situ hybridization, etc.; or immunological technology detection reagents, such as ELISA, immunofluorescence analysis, etc.; or sequencing technology detection reagents such as RNA sequencing, Sanger sequencing, transcriptome sequencing, etc.; or bioinformatics method-related detection reagents, such as microarray technology, protein chip, etc., or mass spectrometry analysis technology.
[0072] Any of the above items is preferably a method for constructing a diagnostic model according to the present invention, and the detection product of the present invention is more preferably a detection reagent for the cervical cancer lymph node metastasis biomarker transcriptome; further preferably a transcriptome sequencing-related reagent.
[0073] The present invention also provides a diagnostic model for cervical cancer lymph node metastasis. The model determines the risk of cervical cancer by detecting the gene expression of cervical cancer lymph node metastasis biomarkers in cervical cancer patient samples. The cervical cancer lymph node metastasis biomarkers include at least one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. A diagnostic nomogram for the cervical cancer lymph node metastasis biomarkers is created. A corresponding score is obtained based on the numerical values of each cervical cancer lymph node metastasis biomarker in the diagnostic nomogram. The scores are then summed to obtain a total score. The total score is used to calculate the probability of cervical cancer lymph node metastasis. A higher total score indicates a higher predicted risk of cervical cancer lymph node metastasis. Patients with a total score greater than -0.011 have a higher risk of cervical cancer lymph node metastasis and should subsequently undergo more aggressive clinical treatment. The diagnostic sensitivity is 81.5% and the specificity is 75%. The numerical value of the cervical cancer lymph node metastasis biomarker is preferably a numerical value obtained by transcriptome sequencing of the cervical cancer lymph node metastasis biomarker. In the diagnostic model for cervical cancer lymph node metastasis provided by the present invention, higher expression of the four markers CARD9, CFL1P1, GRASLND, and MNX1_AS2 indicates a higher model score and a higher risk of cervical cancer lymph node metastasis. Higher expression of the three markers MRAS, OLFML2A, and RPS28 indicates a lower model score and a lower risk of cervical cancer lymph node metastasis.
[0074] In the diagnostic model for cervical cancer lymph node metastasis provided by the present invention, a diagnostic nomogram is constructed using a combination of seven markers: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. Based on these seven key genes, a nomogram model was constructed using the R language "rms" package, version 6.5.0, to predict the probability of cervical cancer lymph node metastasis. The nomogram model, which combines the seven markers: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, was constructed. The "regplot" package was used to draw the nomogram for the retrospective model. The calibration curve was drawn using "rms," the DCA decision curve was drawn using "ggDCA," and the ROC curve was drawn using "pROC" to evaluate the model's goodness of fit.
[0075] The corresponding score is obtained according to the numerical value of each cervical cancer lymph node metastasis biomarker in the diagnostic nomogram. The numerical value of the cervical cancer lymph node metastasis biomarker refers to the gene expression level of each marker in the patient sample. Preferably, the present invention obtains the gene expression level of each marker by transcriptome sequencing. The gene expression level of each marker is substituted into the LASSO regression model to obtain the score corresponding to the numerical value of each cervical cancer lymph node metastasis biomarker. The scores corresponding to all markers are added to obtain a total score. Those with a total score >-0.011 have a higher risk of cervical cancer lymph node metastasis and should adopt more active clinical treatment methods in the future.
[0076] This study used transcriptome sequencing data from cervical tissue from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis. Using bioinformatics methods, the study comprehensively explored the molecular changes and pathogenesis of CC lymph node metastasis. Using multiple bioinformatics analysis methods, the study systematically explored differentially expressed genes in CC lymph node metastasis, conducting correlation analysis, machine learning, and receiver operating characteristic (ROC) analysis to identify key genes. Subsequently, the AUC values of the ROC curves for different combinations of key genes were calculated to select the gene combination with the highest AUC value. The results showed that the diagnostic model constructed with the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 had the highest AUC value (0.8704). The nomogram consists of a "score" representing the score of each key gene and a "total score" representing the sum of all key gene scores. A higher score indicates a higher probability of lymph node metastasis in CC. The analysis used "rms" to plot the calibration curve, "ggDCA" to plot the DCA decision curve, and "pROC" to plot the ROC curve to assess model fit. Results showed that the nomogram had good predictive ability (HL test p-value > 0.05, DCA curve performance was generally higher than that of ALL and NONE, and AUC value = 0.8704).
[0077] The cervical cancer lymph node metastasis diagnostic model, diagnostic nomogram, cervical cancer lymph node metastasis biomarkers, and cervical cancer lymph node metastasis diagnostic products provided by the present invention are applicable to various clinical stages of cervical cancer, such as Stage I, Stage II, Stage III, and Stage IV. The present invention is particularly significant for the prediction and diagnosis of early-stage cervical cancer lymph node metastasis. Existing diagnostic techniques for cervical cancer pelvic lymph node metastasis primarily rely on imaging diagnostics—CT / MRI / PET-CT—but the diagnostic accuracy is relatively low. CT or MRI has high specificity for lymph nodes with a short diameter greater than 1 cm, but has relatively low specificity and sensitivity for lymph nodes smaller than 1 cm. Research results by Yang et al. showed that CT had a sensitivity of 51.2% and a specificity of 70.3% for diagnosing preoperative lymph node metastasis in cervical cancer patients. MRI had a sensitivity of 48.8% and a specificity of 29.7%. PET-CT uses the tumor's sugar uptake to more accurately diagnose tumors, but it still has significant limitations in diagnosing cervical cancer lymph node metastasis, with a sensitivity of 53.8%, a specificity of 95%, a positive predictive value of 75.9%, and a negative predictive value of 96.7%. Furthermore, all imaging tests have no diagnostic value for whether non-enlarged lymph nodes have metastases. The prediction model provided by the present invention uses genetic testing of cervical tissue. Traditional imaging tests have no diagnostic value for determining whether patients with no enlarged lymph nodes have lymph node metastasis. However, the method provided by the present invention, regardless of whether the lymph nodes are enlarged on imaging, the results of the molecular detection of the present invention have high diagnostic and predictive value for cervical cancer lymph node metastasis. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0079] Figure 1 The volcano map of DEGs was drawn using the R packages “ggplot2” and “ComplexHeatmap” in the preferred embodiment 1 of the present invention.
[0080] Figure 2 This is a diagnostic nomogram of seven preferred biomarkers of cervical cancer lymph node metastasis in preferred embodiment 1 of the present invention.
[0081] Figure 3 It is the calibration curve of the seven preferred nomograms of cervical cancer lymph node metastasis biomarkers in the preferred embodiment 1 of the present invention.
[0082] Figure 4 : is the ROC curve of the diagnostic model in the preferred embodiment 1 of the present invention.
[0083] Figure 5 It is an operation flow chart of a computer system constructed by the present invention.
[0084] Figure 6 This is a flowchart of the operation of a smart device for assessing the risk of lymph node metastasis in cervical cancer. DETAILED DESCRIPTION
[0085] Example 1: Screening and modeling of key genes
[0086] By analyzing transcriptome sequencing data of cervical tissues from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis, we comprehensively and deeply explored the molecular changes and pathogenesis of CC lymph node metastasis using bioinformatics methods. Based on multiple bioinformatics analysis methods, we systematically performed correlation analysis, machine learning, and ROC analysis on differentially expressed genes to screen out key genes.
[0087] Total RNA was isolated and purified using TRIzol (thermofisher, 15596018) according to the manufacturer's protocol. Total RNA quantity and purity were controlled using a NanoDrop ND-1000 (NanoDrop, Wilmington, DE, USA) and RNA integrity was assessed using a Bioanalyzer 2100 (Agilent, CA, USA). Concentrations >50 ng / μL, RIN values >7.0, and total RNA >1 μg were sufficient for downstream experiments.
[0088] Poly(A)-containing mRNA was specifically captured using oligo(dT) magnetic beads (Dynabeads Oligo (dT), cat. 25-61005, Thermo Fisher, USA) through two rounds of purification. The captured mRNA was fragmented using a magnesium fragmentation kit (NEBNextR Magnesium RNA Fragmentation Module, cat. E6150S, USA) at 94°C for 5-7 minutes. The fragmented RNA was then synthesized into cDNA using reverse transcriptase (Invitrogen SuperScript™ II Reverse Transcriptase, cat. 1896649, CA, USA). The DNA / RNA duplexes were then converted to DNA duplexes using E. coli DNA polymerase I (NEB, cat. m0209, USA) and RNase H (NEB, cat. m0297, USA). dUTP solution (Thermo Fisher, cat. R0133, CA, USA) was incorporated into the duplexes to blunt-end the duplexes. A bases were then added to each end to facilitate ligation with T-terminal adapters. The fragments were size-selected and purified using magnetic beads. The duplexes were digested with UDG enzyme (NEB, cat. m0280, MA, US) and then subjected to PCR with eight cycles of initial denaturation at 95°C for 3 minutes, denaturation at 98°C for 15 seconds each, annealing at 60°C for 15 seconds, extension at 72°C for 30 seconds, and a final extension at 72°C for 5 minutes to generate a library with a fragment size of 300 bp ± 50 bp (strand-specific library).
[0089] Finally, paired-end sequencing was performed using the Illumina Novaseq™ 6000 using standard protocols in the PE150 sequencing mode. Valid data that mapped to the reference genome was defined as mapping to exons, introns, and intergenic regions, depending on the region of interest. Generally, well-annotated species (such as human and Arabidopsis thaliana) tend to have the highest percentage of sequencing reads mapping to exon regions. Reads mapping to intron and intergenic regions may be due to pre-mRNA splicing events, ncRNA (non-coding RNA), incomplete genome annotation, DNA contamination, and background noise. Results showed that the percentage of reads mapping to exon regions exceeded 90% for all samples, indicating high sequencing data quality and reliable sequencing results.
[0090] To determine whether there were clusters or outliers in the samples, principal component analysis (PCA) was performed on the transcriptome sequencing dataset using the FactoMineR package to perform data dimensionality reduction. This allowed for the identification of discrete patterns in the MN and MP samples through patterning and visualization. The results showed that PC1 explained 22.37% of the variance, indicating that the first principal component captured a relatively high degree of variability in the data. The combined variance explained by PC1 and PC2 was 29.97% (22.37% + 7.6%), meaning that these two principal components accounted for nearly one-third of the variability in the dataset.
[0091] To identify differentially expressed genes (DEGs) between MN and MP samples, the “DESeq2” package was used to analyze the differentially expressed genes (DEGs) between MP and MN samples (MP group vs. MN group) in the transcriptome dataset (a total of 1374 DEGs: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC|>0.5, p-value<0.05). Subsequently, the R packages “ggplot2” and “ComplexHeatmap” were used to plot the volcano plot and heat map of the DEGs (the top 10 up-regulated genes with the largest changes in |log2FC| were SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, and TEKT4; the top 10 down-regulated genes were FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5).
[0092] Figure 1As shown, the volcano plot of DEGs was drawn using the R packages "ggplot2" and "ComplexHeatmap". The volcano plot shows the TOP10 genes with the highest up- and down-regulation fold differences, with high expression in red, low expression in green, and genes with no significant differences in gray.
[0093] Subsequently, the AUC values of the receiver operating characteristic (ROC) curves for different combinations of key genes were calculated to select the gene combination with the highest AUC. The diagnostic model constructed using the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 had the highest AUC value (0.8704). The nomogram consists of a "score" representing the score for each key gene and a "total score" representing the sum of all key gene scores. A higher score indicates a higher probability of lymph node metastasis in CC. The model fit was assessed using the "rms" calibration curve, the "ggDCA" decision curve, and the "pROC" ROC curve. Results showed that the nomogram had good predictive ability (HL test p-value > 0.05, the DCA curve yielded a higher return than the ALL and NONE groups, and the AUC value was 0.8704).
[0094] Figure 2 Shown is a diagnostic nomogram for seven key genes identified through screening: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. The top layer of the nomogram represents the score (points) for each key gene. The corresponding score is calculated based on the key gene's value, and the total points are then summed to form the "Total Points." A higher total score indicates a higher predicted risk. The bottom layer displays the risk probability (Pr(Y)) calculated based on the total score.
[0095] Figure 3 Shown are the calibration curves for the nomogram of seven key genes identified, namely, CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, for cervical cancer lymph node metastasis. The X-axis represents predicted probability, and the Y-axis represents actual probability. The blue curve represents the actual prediction (Apparent), the black curve represents the bias-corrected result (Bias-corrected), and the dashed line represents the ideal case (Ideal), the reference line where the model prediction is completely consistent with the actual result. The p-value of the Hosmer-Lemeshow test was 0.315, indicating that the model calibration performance was good and did not deviate significantly from the ideal case.
[0096] This example uses the rms package to construct a logistic regression model and the regplot package to draw a nomogram of the regression model. Using the logistic regression model and the log2 value of the CPM of gene expression, the following diagnostic model expression is obtained:
[0097]
[0098] Subsequently, the OR value of each factor in the diagnostic model was calculated, as shown in the following table.
[0099]
[0100] Figure 4 This is the ROC curve of the diagnostic model. The ROC (Receiver Operating Characteristic) curve is used to evaluate the performance of the diagnostic model. The horizontal axis represents 1-Specificity (false positive rate), and the vertical axis represents Sensitivity (true positive rate). The gray dashed line in the figure represents the baseline for random prediction, and the black dot marks the optimal cutoff point, which corresponds to a logit (p) value of -0.011, a probability of approximately 0.497, a sensitivity of 0.815, and a specificity of 0.750. The red curve is the ROC curve of the diagnostic model. The area under the curve (AUC) is approximately 0.8704, indicating that the model has good discriminatory ability. A score greater than the cutoff is defined as high-risk, while a score less than the cutoff indicates low-risk. Conversely, a score greater than the cutoff indicates high risk, and patients with high risk should receive active treatment.
[0101] Example 2: Development and validation of a computational model for assessing the risk of lymph node metastasis in cervical cancer
[0102] In this example, transcriptome sequencing data were collected from cervical tissue collected from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis, as described in Example 1. Differential expression analysis was performed using DESeq2 to identify candidate genes with significantly differential expression between metastatic and non-metastatic samples. LASSO regression was further used for variable selection, ultimately identifying seven gene markers with high predictive value: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. A random forest algorithm was used to train a binary classification model using the expression levels of these seven genes as feature variables. Model performance was evaluated using 10-fold cross-validation with a training set to test set ratio of 7:3. In the independent test set, the model achieved an AUC (area under the curve) of 0.93, demonstrating that the model has extremely high predictive efficacy in distinguishing patients with cervical cancer who have lymph node metastasis.
[0103] Example 3: External independent sample verification of model effectiveness and comparative analysis with traditional methods
[0104] To further verify the reliability and generalization ability of the calculation model for evaluating the risk of cervical cancer lymph node metastasis described in Example 2, this example introduces a group of external independent verification samples and compares and analyzes the model prediction results with traditional imaging assessment methods.
[0105] 1. External Validation Sample Collection
[0106] This example includes a total of 20 independent cervical cancer patient samples, including 10 cases in the metastasis group and 10 cases in the non-metastasis group. All samples did not participate in the model training process, and the sample processing method was consistent with Example 1.
[0107] 2. Gene Expression Data Detection
[0108] Referring to the process described in Example 1, sample RNA was extracted and qRT-PCR was performed to obtain the expression levels of seven key model genes (CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28).
[0109] 3. Model External Prediction Results
[0110] The normalized expression data were input into the model trained in Example 2, and the model automatically output the predicted probability of lymph node metastasis for each patient. The results were as follows: AUC (ROC): 0.912; accuracy: 90.0%; sensitivity: 93.3%; specificity: 86.7%.
[0111] The above results show that the model can still maintain excellent predictive ability in new independent samples and has good stability and generalization ability.
[0112] Example 4: Computer system implementation based on the model of the present invention
[0113] In this embodiment, a computer system is constructed to run the above risk prediction model, which includes the following modules (see Figure 5 ):
[0114] Data import module: used to read standardized gene expression data of patient samples, supporting CSV, Excel and other formats.
[0115] Preprocessing module: includes functions such as missing value processing and standardization conversion to ensure the consistency of input data.
[0116] Model calculation module: loads the pre-trained random forest model file, performs inference on the input data, and outputs the predicted risk score and probability of each sample.
[0117] Visualization module: Displays the prediction results through bar charts, heat maps, etc., so that doctors can quickly understand the results.
[0118] Interface module: can be connected to the hospital information system (HIS) or laboratory data system (LIMS) to realize automatic data flow and automatic triggering of risk assessment.
[0119] The system is deployed on a hospital server or private cloud environment. The front-end interface is developed using the Django framework, and the back-end model is implemented using the scikit-learn library in Python.
[0120] Example 5: Integrated Lymph Node Metastasis Risk Assessment Device
[0121] The inventors of this invention have designed a smart device for assessing the risk of lymph node metastasis in cervical cancer (see Figure 6 ), mainly including:
[0122] Detection module: Built-in fluorescence quantitative PCR system, used to detect the expression levels of the above 7 genes in cervical cancer tissue samples, and optional automated RNA extraction unit.
[0123] Data acquisition module: connects to the detection module and transmits expression data to the main control system in real time.
[0124] Main control processing module: An embedded processor runs the model program described in the present invention, integrates the Linux system, has 8GB of memory, and is equipped with a dedicated computing acceleration chip (such as the ARM Cortex-A72 series).
[0125] Output display module: Built-in touch screen display interface, which can display the prediction score, transfer probability, result suggestions and visual reports in real time.
[0126] Communication module: supports WiFi, Bluetooth, USB and Ethernet interfaces, facilitating communication with hospital databases or host computers.
[0127] After the device is started, the operator only needs to import the sample and click the detection and prediction button to complete the full process evaluation, providing timely auxiliary basis for clinical surgical decisions.
[0128] The above descriptions are merely embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. A computational method for assessing the risk of lymph node metastasis in cervical cancer, the method taking the expression levels of multiple gene markers as input and outputting a metastasis risk score for the patient; the gene markers are a group consisting of the following gene markers: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
2. The calculation method according to claim 1, wherein the calculation method is established based on a machine learning algorithm, and the algorithm is selected from random forest, logistic regression, LASSO regression, COX regression or neural network algorithm.
3. A computer program product, which, when executed by a processor, is used to implement the evaluation function of the calculation method according to claim 1 or 2, the program comprising the following steps: a) receiving gene expression data of a sample to be tested; b) Perform data preprocessing and standardization; c) inputting standardized data into calculation methods; d) Output the lymph node metastasis risk score and predicted probability corresponding to the sample.
4. A computer system for assessing the risk of lymph node metastasis of cervical cancer, comprising a processor, a memory, and a program module stored in the memory and executed by the processor, wherein the program module implements the steps described in claim 3 to achieve risk assessment of a target sample.
5. A method for constructing a diagnostic model for cervical cancer lymph node metastasis, characterized in that: The following steps are involved: Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 2: Perform gene expression testing on the tissue samples obtained in step 1; Step 3: Based on the gene expression detection data obtained in step 2, the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis are compared; Step 4: Screen key genes based on bioinformatics analysis methods; Step 5: Draw a diagnostic nomogram of key genes, obtain corresponding scores based on the values of key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis based on the total score. Wherein, in step 4, the key genes are CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A and RPS28.
6. The method according to claim 5, wherein in step 5, the risk probability of cervical cancer lymph node metastasis is estimated based on the total score obtained from the diagnostic nomogram of key genes, and a higher total score indicates a higher risk of cervical cancer lymph node metastasis.
7. An assessment device for assessing the risk of lymph node metastasis in cervical cancer, characterized in that: include: a data acquisition device for acquiring detection data of an evaluation subject, wherein the detection data is a level value of a marker in a cervical cancer lymph node metastasis marker detected from a sample of an evaluation subject suspected of having cervical cancer lymph node metastasis; a data processing device for calculating an index value of a marker in a cervical cancer lymph node metastasis marker based on the test data of the evaluation subject, and then calculating whether the obtained index value of the marker is within a risk range of the corresponding marker; The cervical cancer lymph node metastasis markers are CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
8. The evaluation device according to claim 7, characterized in that The device further comprises a detection device to detect the level of the marker.
9. The evaluation device according to claim 7, characterized in that The data acquisition device is an input device for inputting data or a communication device for reading data from an external data storage device or a storage device of the detection device through an interface.
10. The evaluation device according to claim 9, characterized in that When a communication device is used, the corresponding external data storage device is a data storage device on a device for measuring biomarkers.
Citation Information
Patent Citations
Micronucleus DNA of peripheral red blood cells and application of micronucleus DNA
CN112094907A
Equipment, method and system for diagnosing lymphatic metastasis of cervical cancer and application of equipment, method and system
CN118471335A