Method and model constructed based on transcriptome sequencing and used for diagnosing early cervical cancer lymph node metastasis
By constructing a calculation model based on multiple gene markers, predicting the risk of cervical cancer lymph node metastasis, solving the problem of low diagnostic accuracy of cervical cancer lymph node metastasis in the prior art, and achieving more efficient and accurate prediction and diagnosis.
Patent Information
- Application Number
- CN202510671066.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the prior art, the diagnosis of lymph node metastasis of cervical cancer depends on imaging examinations, with poor sensitivity and accuracy, resulting in many patients with lymph node metastasis of CC are not recognized.
A computational model based on multiple gene markers was constructed, and the gene expression data was characterized by machine learning algorithms to predict the risk of lymph node metastasis in cervical cancer. Specifically, genes such as CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28 were selected as input variables to output the lymph node metastasis risk score for each subject.
It improves the accuracy of predicting lymph node metastasis in cervical cancer, and provides an efficient, non-invasive, quantifiable method for early prediction and auxiliary judgment of lymph node metastasis risk in cervical cancer patients.
Smart Images

Figure CN120183512A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical diagnosis, and in particular to a diagnostic model for cervical cancer lymph node metastasis and a construction method thereof, as well as an application of the model in the preparation of a diagnostic product for cervical cancer lymph node metastasis. Background Art
[0002] Cervical cancer (CC) ranks fourth among female malignant tumors. It is a common malignant tumor of the female reproductive system and seriously threatens women's life and physical and mental health. One of the important factors affecting the prognosis of CC patients is lymph node metastasis. Studies have confirmed that lymph node metastasis in early CC will seriously affect its 5-year survival rate. Whether lymph nodes are metastatic not only affects the prognosis of cervical cancer patients, but also plays a decisive role in individualized treatment. For early CC patients, systematic lymph node resection is often used to determine the occurrence of lymph node metastasis and provide a basis for postoperative supplementary treatment. However, it also causes a certain number of CC patients without lymph node metastasis to undergo unnecessary surgery, which in turn leads to unnecessary surgical risks. Therefore, accurately judging the lymph node status of CC patients and accurately assessing whether lymph node metastasis occurs are crucial for clarifying the staging, judging the prognosis, and formulating treatment plans. At present, the diagnosis of CC lymph node metastasis in clinical practice is mainly based on the morphological characteristics of lymph nodes examined by magnetic resonance imaging, CT or PET / CT imaging. The most commonly used standard is whether the short diameter of the lymph node exceeds 10 mm. Although this method has a reasonable specificity, its sensitivity and accuracy are poor. As a result, a large proportion of CC lymph node metastasis patients are not identified. Therefore, it is still necessary to find new biomarkers (or key genes) that can predict CC lymph node metastasis.
[0003] In recent years, with the emergence and rapid development of high-throughput sequencing technology, more and more CC genes and epigenetic features have been discovered. However, few studies have constructed accurate diagnostic models based on transcriptome sequencing of clinical CC lymph node metastasis and CC primary lesion patient samples. Summary of the invention
[0004] Cervical cancer is one of the common gynecological malignancies, and lymph node metastasis is a key factor affecting its treatment plan and prognosis assessment. Traditional evaluation methods mostly rely on imaging examinations and surgical pathology analysis, but they have problems such as low accuracy, strong invasiveness, and delayed diagnosis. Therefore, there is an urgent need for an efficient, non-invasive, and quantifiable method for early prediction and auxiliary judgment of the risk of lymph node metastasis in patients with cervical cancer.
[0005] In the existing technology, the diagnosis of pelvic lymph node metastasis in cervical cancer mainly relies on imaging diagnosis, with poor sensitivity and accuracy. To overcome the defects in the existing technology and find new biomarkers (or key genes) that can predict CC lymph node metastasis, the present invention relates to the fields of bioinformatics and tumor-assisted diagnosis technology, and specifically relates to a computational model, a construction method, a computer program, a computer system and related evaluation devices for evaluating the risk of lymph node metastasis in cervical cancer.
[0006] The present invention provides a computational model for evaluating the risk of lymph node metastasis in cervical cancer. The model takes the expression data of multiple gene markers as input variables and outputs the lymph node metastasis risk score for each subject. The gene markers are selected from the gene set consisting of: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28. These genes are screened out by bioinformatics analysis methods in the present invention and have the characteristics of significant expression differences between the metastatic and non-metastatic groups and strong predictive ability.
[0007] The model is constructed based on machine learning algorithms. Methods such as random forest, logistic regression, LASSO regression, COX regression or neural network can be used to perform feature modeling on the above genes, so as to improve the discrimination performance of the possibility of lymph node metastasis.
[0008] In addition, the present invention also provides a computer program, system and evaluation method based on this model. The program can be deployed in an evaluation system to realize the full-process automated processing from inputting sample data to outputting risk scores. The steps include receiving the gene expression data of the original sample, performing standardization processing on the data, inputting it into the constructed model for inference, and outputting the probability value of the metastasis risk.
[0009] The present invention also proposes a method for constructing the model. This method collects tissue samples of a large number of cervical cancer patients (including metastatic and non-metastatic groups), conducts high-throughput gene expression detection, and then systematically screens out key genes related to metastasis prediction through differential analysis and model training steps, and finally establishes a quantitative scoring model (such as a nomogram model) for calculating the metastasis risk of an individual.
[0010] To realize the engineering application of this technical solution, the present invention also provides an evaluation device, which includes a data acquisition device, a data processing module and an optional detection device, and can complete the processes of biomarker data collection, analysis and risk determination. The evaluation device can be used in hospitals, laboratories or other biological sample detection scenarios, which helps to realize the rapid evaluation of lymph node metastasis risk, intelligent auxiliary decision-making and the provision of individualized treatment suggestions.
[0011] More specifically, the present invention provides a method for constructing a diagnostic model for cervical cancer lymph node metastasis, comprising the following steps: Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 2: Perform gene expression detection on the tissue samples obtained in step 1; Step 3: Based on the gene expression detection data obtained in step 2, the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis are compared; Step 4: Screen out key genes based on bioinformatics analysis methods; Step 5: Draw a diagnostic nomogram for key genes, obtain corresponding scores based on the values of key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis based on the total score.
[0012] The present invention is different from traditional imaging examinations. The present invention performs molecular testing on cervical tissue rather than imaging testing on lymph nodes. Therefore, regardless of whether the lymph nodes are enlarged in the image, the present invention can diagnose and predict cervical cancer lymph node metastasis with higher accuracy.
[0013] The present invention establishes a diagnostic model for cervical cancer lymph node metastasis based on the detection of gene expression. Preferably, the detection of gene expression includes transcription level detection and protein level detection. The detection method of gene expression includes but is not limited to molecular biology techniques, such as PCR, in situ hybridization, etc.; or immunological techniques, such as ELISA, immunofluorescence analysis, etc.; or sequencing techniques such as RNA sequencing, Sanger sequencing, transcriptome sequencing, etc.; or bioinformatics methods, such as microarray technology, protein chips, etc., or mass spectrometry analysis technology.
[0014] A more preferred gene expression detection scheme of the present invention is to detect the gene transcriptome. The detection of the relevant gene transcriptome can more directly reflect the gene expression conditions and differences in cervical cancer patient samples. Since the transcriptome is the direct product of gene expression, by detecting the transcriptome, the expression levels of each gene in cervical cancer patient samples can be accurately understood, which helps to reveal the variation law of gene activity, thereby achieving the purpose of auxiliary diagnosis. The detection of the relevant gene transcriptome can predict disease risks earlier. Compared with protein detection, the detection of the transcriptome may be able to predict the occurrence of lymph node metastasis in cervical cancer earlier, predict disease risks, capture earlier disease signals, and timely adjust the prognosis treatment plan. The detection of the relevant gene transcriptome has higher sensitivity and can detect trace gene expression changes, which is particularly important for identifying low-abundance genes or rare transcriptomes and helps to discover potential disease-related genes. Transcriptome detection provides more comprehensive information. Transcriptome detection can provide comprehensive information about gene expression patterns, including the expression levels and splicing methods of different transcriptomes, and this information helps to deeply understand the functions and regulatory mechanisms of genes.
[0015] Preferably, the method for constructing a diagnostic model for cervical cancer lymph node metastasis of the present invention is based on a gene sequencing method, and includes the following steps: Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 2: Perform transcriptome sequencing on the tissue samples obtained in Step 1 to obtain transcriptome sequencing data; Step 3: According to the transcriptome sequencing data obtained in Step 2, compare the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 4: Screen out key genes based on bioinformatics analysis methods; Step 5: Draw a diagnostic nomogram of the key genes, obtain the corresponding scores according to the values of the key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis according to the total score.
[0016] Preferably for any of the above, in Step 3, in order to compare the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis, the preferred solution of the present invention is: Preferably, use the FactoMineR package to perform principal component analysis (PCA) for data dimensionality reduction on the transcriptome sequencing data set.
[0017] To determine whether there are clusters or outliers in a sample, the present invention uses the FactoMineR package to perform principal component analysis (PCA) for data dimensionality reduction on a transcriptome sequencing dataset, so as to identify the discreteness of metastasis-negative (MN) and metastasis-positive (MP) samples through patternization and visualization.
[0018] Performing principal component analysis (PCA) for data dimensionality reduction on a transcriptome sequencing dataset using the FactoMineR package is a common practice in bioinformatics, aiming to simplify data analysis by reducing the dimensionality of the data while retaining as much of the original data information as possible. In the results of PCA (principal component analysis), PC1 (the first principal component) and PC2 (the second principal component) are included. PC1 is the axis that contains the largest variance in the PCA transformation, and it represents the direction of the largest change in the data. In other words, it is the direction with the largest variance in the original dataset, capable of capturing as much data information (variance) as possible. Since PC1 captures the largest variation in the data, it is usually the most important principal component, capable of explaining the largest proportion of the total variance in the dataset. When visualizing the PCA results, PC1 is often used as the horizontal axis to show the distribution of samples in the principal component space. PC2 is the principal component with the second largest eigenvalue on the premise of ensuring orthogonality (i.e., perpendicular directions) to PC1. It represents the second largest variation direction in the data, capturing the largest remaining variance after PC1.
[0019] The present invention loads the R package for the transcriptome sequencing dataset of cervical tissue samples through FactoMineR, performs PCA analysis, views the overview of the PCA results, and obtains information such as the proportion of the explained variance of the principal components. By patternization and visualization, the discreteness of metastasis-negative (MN) and metastasis-positive (MP) samples is identified. The results show that PC1 explains 22.37% of the variance, which means that the first principal component can capture relatively high variability in the data. The combined variance explanation of PC1 and PC2 is 29.97% (22.37% + 7.6%), which means that these two principal components can explain nearly one-third of the variability in the dataset.
[0020] Further preferably, the present invention analyzes the differential genes between MP and MN samples (MP group vs MN group) in the transcriptome dataset through the "DESeq2" package. Analyzing differential genes between transcriptome data using the DESeq2 package is a common bioinformatics analysis process.
[0021] In order to identify the differential genes between MN samples and MP samples, the differential genes between MP and MN samples (MP group VS MN group) in the transcriptome dataset were analyzed through the "DESeq2" package, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC| > 0.5, p-value < 0.05); Subsequently, the volcano plot and heatmap of DEGs were drawn using the R packages "ggplot2" and "ComplexHeatmap" (Top 10 up-regulated genes with the largest change in |log2FC|: SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, TEKT4; Top 10 down-regulated genes: FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, PAK5).
[0022] The gene numbers of the top 10 up-regulated genes with the largest change in |log2FC| are respectively: SOHLH1 (ID: 402381), MUCL3 (ID: 283232), PNMA5 (ID: 114836), REG1A (ID: 5967), FGA (ID: 2243), PGC (ID: 5225), DSCAM-AS1 (ID: 102723407), IGHV3-22 (ID: 28450), LINC01320 (ID: 100507487), TEKT4 (ID: 150465).
[0023] The top 10 down-regulated genes with the largest change in |log2FC| are FOLR1P1 (ID: 107986793), TLX1 (ID: 3195), MGAM2 (ID: 9342), KRT1 (ID: 3848), KRTDAP (ID: 200172), LINC00167 (ID: 100507065), MUCL1 (ID: 118430), LINC00457 (ID: 285237), BPIFC (ID: 653145), PAK5 (ID: 57144).
[0024] Based on the above analysis, 1057 up-regulated genes were obtained in the samples of patients with cervical cancer lymph node metastasis. Further, at least one of SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, TEKT4 is included in the top 10 up-regulated genes with the largest change in |log2FC|.
[0025] Based on the above analysis, 317 down-regulated genes were obtained from the samples of patients with cervical cancer lymph node metastasis. The down-regulated genes with the largest further |log2FC| changes include at least one of FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, and PAK5.
[0026] Preferably, in step 4, in order to obtain key genes, bioinformatics analysis methods are performed, including systematically performing correlation analysis, machine learning, ROC analysis, etc. on the differential genes.
[0027] Constructing a nomogram based on key genes is usually used to evaluate the prognosis in the fields of oncology, medicine, etc.
[0028] Preferably, based on the transcriptome sequencing data set obtained in step 2 and the gene expression difference data obtained by the "DESeq2" analysis in step 3, a nomogram model is constructed.
[0029] Preferably, the data set is divided into a training set and a validation set, and the relationship between the key gene transcriptome detection data and the lymph node metastasis situation is analyzed using multivariate regression analysis on the training set. The multivariate regression analysis is preferably a Logistic regression model, etc.
[0030] Preferably, according to the results of the regression analysis, a nomogram is drawn using statistical software. Preferably, the nomogram function is used, such as preferably using the rms package in R language to draw the nomogram. The nomogram visualizes the results of the regression analysis, facilitating the evaluation of whether cervical cancer patients have lymph node metastasis.
[0031] Preferably, the validation set is used to evaluate the nomogram model to evaluate the prediction ability and accuracy of the model. Preferably, the evaluation methods include calculating the C-index, drawing a calibration curve or an ROC curve, etc.
[0032] More preferably, in the present invention, during the construction of the nomogram: calculate the AUC value of the ROC curve of the nomogram model constructed by different key gene combinations to select the gene combination with the highest AUC value; use "rms" to draw the calibration curve, "ggDCA" to draw the DCA decision curve, and "pROC" to draw the ROC curve in the analysis to evaluate the goodness of fit of the model.
[0033] In a preferred embodiment of the present invention, based on the key genes, the AUC values of the ROC curves of the nomogram models constructed with different combinations are calculated to select the gene combination with the highest AUC value. It is found that the diagnostic model constructed with the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 has the highest AUC value (0.8704). The nomogram consists of "score" and "total score". The former represents the score of each key gene, and the latter represents the sum of the scores of all key genes. The higher this value, the higher the probability of lymph node metastasis in CC. In the analysis, the "rms" is used to draw the calibration curve, the "ggDCA" is used to draw the DCA decision curve, and the "pROC" is used to draw the ROC curve to evaluate the fitness of the model. The results show that the nomogram has good predictive ability (HL test p.value > 0.05, the benefit of the DCA curve is basically above ALL and NONE, AUC value = 0.8704).
[0034] Furthermore, the top layer of the nomogram represents the score (Points) of each key gene. According to the value of the key gene (such as the expression data obtained from the transcriptome sequencing of the key gene), the corresponding score is obtained, and then the scores are added together to get the "Total points". The higher the total score, the higher the predicted risk. The bottom shows the risk probability (Pr(Y)) calculated based on the total score.
[0035] Furthermore, the key gene includes at least any one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0036] The gene numbers of the key genes are respectively: CARD9 (ID: 64170), CFL1P1 (ID: 100129361), GRASLND (ID: 100507642), MNX1_AS2 (ID: 100873971), MRAS (ID: 22808), OLFML2A (ID: 79589), and RPS28 (ID: 6234).
[0037] Furthermore, the key gene includes any combination of two of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0038] Furthermore, the key gene includes any combination of three of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0039] Furthermore, the key genes include any combination of four of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0040] Furthermore, the key genes include any combination of five of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0041] Furthermore, the key genes include any combination of six of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0042] Furthermore, the key genes are the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0043] The most preferred diagnostic model of the present invention is a detection model constructed for multiple genes. In a preferred embodiment of the present invention, the nomogram function is used to generate a nomogram. Using the transcriptome expression levels of seven key genes, CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, as predictive variables, sequence the transcriptomes of the seven key genes to obtain their expression values; create a scoring scale based on the seven key genes. At the top layer of the nomogram, each key gene has a corresponding scoring scale. Find the corresponding score on the corresponding scale according to the value of the key gene (the expression data obtained from the transcriptome sequencing of the key gene). Then, add these scores to get the "Total points". Furthermore, based on the total score obtained, estimate the risk probability of cervical cancer lymph node metastasis. The higher the total score, the higher the risk of predicting cervical cancer lymph node metastasis. Those with a total score > -0.011 have a higher risk of cervical cancer lymph node metastasis, and more aggressive clinical treatment methods should be adopted subsequently.
[0044] In a preferred embodiment of the present invention, cervical tissue is obtained through cervical biopsy for transcriptome sequencing. A diagnostic model is constructed using a biomarker combination consisting of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, and a diagnostic nomogram is drawn. The "rms" package is used to draw the calibration curve, the "ggDCA" package is used to draw the DCA decision curve, and the "pROC" package is used to draw the ROC curve to evaluate the goodness of fit of the model. The results show that the nomogram has good predictive ability (HL test p.value > 0.05, the benefit of the DCA curve is basically above ALL and NONE, and the AUC value = 0.8704). That is, the p value of the Hosmer-Lemeshow test for the model prediction and the actual situation is 0.315, indicating that the calibration performance of the model is good and does not deviate significantly from the ideal situation. The ROC (Receiver Operating Characteristic) curve is used to evaluate the performance of the diagnostic model, and the area under the curve (AUC, Area Under the Curve) is 0.8704, indicating that the model has good discriminatory ability.
[0045] The present invention also provides a diagnostic nomogram for diagnosing cervical cancer lymph node metastasis. Through transcriptome gene sequencing of cervical cancer tissue samples, 10-fold cross-validation LASSO regression analysis is performed on 115 candidate genes using the glmnet package in R language. Lasso regression constructs a penalty function to obtain a more refined model, which compresses some regression coefficients, reduces the dimension of the data, and avoids multicollinearity and overfitting in the multiple regression model. The results show that 9 candidate genes passed the LASSO regression (CARD9, CFL1P1, GRASLND, LHX3, MNX1-AS2, MRAS, OLFML2A, PAX3, RPS28), which are denoted as candidate key genes. Based on the key genes, the "rms" package in R language, version: 6.5.0 is used to construct a nomogram model to predict the probability of cervical cancer lymph node metastasis. The nomogram model consists of a diagnostic model combined with seven biomarkers, namely CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. The "regplot" package is used to draw the nomogram of the retrospective model, the "rms" package is used to draw the calibration curve, the "ggDCA" package is used to draw the DCA decision curve, and the "pROC" package is used to draw the ROC curve to evaluate the goodness of fit of the model.
[0046] Preferably, in any of the above, the cervical tissue sample is a cervical biopsy sample, and the transcriptome sequencing results include the transcriptome sequencing results of the cervical tissue.
[0047] The present invention also provides a biomarker combination for cervical cancer lymph node metastasis, and the biomarker for cervical cancer lymph node metastasis includes at least one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0048] Preferably, any one of the above is used for diagnosing cervical cancer lymph node metastasis.
[0049] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis includes a combination of any two of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0050] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis includes a combination of any three of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0051] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis includes a combination of any four of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0052] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis includes a combination of any five of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0053] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis includes a combination of any six of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0054] Preferably, any one of the above, the biomarker for cervical cancer lymph node metastasis is a combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
[0055] Preferably, any one of the above, through the gene transcriptome expression values of the biomarkers for cervical cancer lymph node metastasis, using the diagnostic nomogram obtained by the present invention, obtaining the corresponding scores according to the values of each biomarker for cervical cancer lymph node metastasis in the diagnostic nomogram, and then adding up the scores to obtain the total score, and calculating the risk probability of cervical cancer lymph node metastasis according to the total score. The higher the total score, the higher the risk of predicting cervical cancer lymph node metastasis.
[0056] Preferably, in any of the above, the cervical cancer lymph node metastasis biomarker is the transcriptome of the cervical cancer lymph node metastasis biomarker gene. By sequencing the transcriptome of the cervical cancer lymph node metastasis biomarker gene, the expression data of the cervical cancer lymph node metastasis biomarker gene is obtained.
[0057] The present invention also provides a detection product for diagnosing cervical cancer lymph node metastasis, a detection reagent for the cervical cancer lymph node metastasis biomarker described in any of the above.
[0058] Preferably, in any of the above, the cervical cancer lymph node metastasis biomarker is any one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28, or any two, or any three, or any four, or any five, or any six, or a combination of all seven.
[0059] Preferably, in any of the above, the detection reagent for the cervical cancer lymph node metastasis biomarker can be a detection reagent for its gene level or a detection reagent for the protein level.
[0060] Preferably, in any of the above, the detection reagent for the cervical cancer lymph node metastasis biomarker includes but is not limited to molecular biology detection reagents such as PCR, in situ hybridization, etc.; or immunological technology detection reagents such as ELISA, immunofluorescence analysis, etc.; or sequencing technology detection reagents such as RNA sequencing, Sanger sequencing, transcriptome sequencing, etc.; or detection reagents related to bioinformatics methods such as microarray technology, protein chip, etc., or mass spectrometry analysis technology.
[0061] Preferably, in any of the above, according to the method for constructing the diagnostic model of the present invention, the detection product of the present invention is more preferably a detection reagent for the transcriptome of the cervical cancer lymph node metastasis biomarker; further preferably a reagent related to transcriptome sequencing.
[0062] The present invention also provides a diagnostic model for cervical cancer lymph node metastasis. By detecting the gene expression of biomarkers for cervical cancer lymph node metastasis in samples of cervical cancer patients, the occurrence risk of cervical cancer is judged; the biomarkers for cervical cancer lymph node metastasis include at least one of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28; a diagnostic nomogram of the biomarkers for cervical cancer lymph node metastasis is drawn, corresponding scores are obtained according to the values of each biomarker for cervical cancer lymph node metastasis in the diagnostic nomogram, and then the scores are added together to obtain a total score. The risk probability of cervical cancer lymph node metastasis is deduced according to the total score. The higher the total score, the higher the risk of predicting cervical cancer lymph node metastasis. Those with a total score > -0.011 have a higher risk of cervical cancer lymph node metastasis, and more aggressive clinical treatment methods should be adopted subsequently. The diagnostic sensitivity is 81.5% and the specificity is 75%. The values of the biomarkers for cervical cancer lymph node metastasis are preferably the values obtained by transcriptome sequencing of the biomarkers for cervical cancer lymph node metastasis. In the diagnostic model for cervical cancer lymph node metastasis provided by the present invention, the higher the expression of the four biomarkers CARD9, CFL1P1, GRASLND, and MNX1_AS2, the higher the model score and the higher the risk of cervical cancer lymph node metastasis. The higher the expression of the three biomarkers MRAS, OLFML2A, and RPS28, the lower the model score and the lower the risk of cervical cancer lymph node metastasis.
[0063] In the diagnostic model for cervical cancer lymph node metastasis provided by the present invention, a diagnostic nomogram is constructed by combining seven biomarkers, namely CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. Based on the seven key genes, the "rms" package in R language, version: 6.5.0 is used to construct a nomogram model to predict the probability of cervical cancer lymph node metastasis. The nomogram model is composed of seven biomarkers, namely CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28, to form a diagnostic model. The "regplot" package is used to draw the nomogram of the retrospective model, the "rms" is used to draw the calibration curve, the "ggDCA" is used to draw the DCA decision curve, and the "pROC" is used to draw the ROC curve to evaluate the fitness of the model.
[0064] Obtain the corresponding score according to the values of various biomarkers for cervical cancer lymph node metastasis in the diagnostic nomogram. The values of the biomarkers for cervical cancer lymph node metastasis refer to the gene expression levels of each biomarker in the patient sample. Preferably, in the present invention, the gene expression levels of each biomarker are obtained by transcriptome sequencing. Substitute the gene expression levels of each biomarker into the LASSO regression model to obtain the scores corresponding to the values of each biomarker for cervical cancer lymph node metastasis. The scores corresponding to all biomarkers are added together to obtain the total score. Those with a total score > -0.011 have a higher risk of cervical cancer lymph node metastasis, and more aggressive clinical treatment measures should be adopted subsequently.
[0065] In the present invention, transcriptome sequencing data of cervical tissues from 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis were used. Through bioinformatics methods, the molecular changes and pathogenesis of CC lymph node metastasis were deeply explored in all aspects. Based on a variety of bioinformatics analysis methods, the differential genes of CC lymph node metastasis were systematically explored for correlation analysis, machine learning, and ROC analysis to screen out key genes. Subsequently, based on the key genes, the AUC values of the ROC curves of the nomogram models constructed with different combinations were calculated to select the gene combination with the highest AUC value. It was found that the diagnostic model constructed with the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 had the highest AUC value (0.8704). The nomogram consists of "score" and "total score". The former represents the score of each key gene, and the latter represents the sum of the scores of all key genes. The higher this value, the higher the probability of CC developing lymph node metastasis. In the analysis, the "rms" was used to draw the calibration curve, the "ggDCA" was used to draw the DCA decision curve, and the "pROC" was used to draw the ROC curve to evaluate the fitness of the model. The results showed that the nomogram had good predictive ability (HL test p.value > 0.05, the benefit of the DCA curve was basically above ALL and NONE, and the AUC value = 0.8704).
[0066] The cervical cancer lymph node metastasis diagnosis model, diagnostic nomogram, cervical cancer lymph node metastasis biomarker, and cervical cancer lymph node metastasis diagnostic product provided by the present invention are applicable to all clinical stages of cervical cancer, such as stage I, stage II, stage III, and stage IV. The present invention has particularly prominent significance for the prediction and diagnosis of early cervical cancer lymph node metastasis. In the prior art, the diagnosis of pelvic lymph node metastasis in cervical cancer mainly relies on imaging diagnosis - CT / MRI / PET-CT, but the diagnostic accuracy is relatively low. CT or MRI has high specificity for lymph nodes with a short diameter greater than 1 cm, but the specificity and sensitivity for lymph nodes less than 1 cm are relatively low. The research results of Yang et al. show that the sensitivity of CT for diagnosing whether the lymph nodes of cervical cancer patients metastasize before surgery is 51.2%, and the specificity is 70.3%. The sensitivity of MRI is 48.8%, and the specificity is 29.7%. PET-CT can more accurately diagnose tumors by the uptake of sugar by tumors, but there are still great limitations in the diagnosis of cervical cancer lymph node metastasis. The sensitivity is 53.8%, the specificity is 95%, the positive predictive value is 75.9%, and the negative predictive value is 96.7%. And all images have no diagnostic value for whether the unenlarged lymph nodes metastasize, while the prediction model provided by the present invention detects genes in cervical tissues. Traditional imaging examinations have no diagnostic value for whether patients with non-enlarged lymph nodes have lymph node metastasis. However, for the method provided by the present invention, regardless of whether the lymph nodes are enlarged on imaging, the detection results of the molecular detection of the present invention have high diagnostic and predictive value for cervical cancer lymph node metastasis. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 It is a volcano plot of DEGs drawn using the R packages "ggplot2" and "ComplexHeatmap" in the preferred Embodiment 1 of the present invention.
[0069] Figure 2 It is a diagnostic nomogram of seven preferred cervical cancer lymph node metastasis biomarkers in the preferred Embodiment 1 of the present invention.
[0070] Figure 3 It is a calibration curve of the nomogram of seven preferred cervical cancer lymph node metastasis biomarkers in the preferred Embodiment 1 of the present invention.
[0071] Figure 4 It is an ROC curve of the diagnostic model in the preferred Embodiment 1 of the present invention.
[0072] Figure 5 It is the operation flowchart of a computer system constructed by the present invention.
[0073] Figure 6 It is the operation flowchart of an intelligent device for evaluating the risk of cervical cancer lymph node metastasis. Detailed implementation manners
[0074] Example 1: Screening and modeling of key genes By performing transcriptome sequencing on the cervical tissues of 27 patients with cervical cancer lymph node metastasis and 32 patients without lymph node metastasis, the molecular changes and pathogenesis of CC lymph node metastasis were deeply explored comprehensively through bioinformatics methods. Based on a variety of bioinformatics analysis methods, differential genes were systematically subjected to correlation analysis, machine learning, and ROC analysis to screen out key genes.
[0075] Use TRIzol (thermofisher, 15596018) to isolate and purify the RNA of the total sample according to the operation protocol provided by the manufacturer. Then use NanoDrop ND-1000 (NanoDrop, Wilmington, DE, USA) to perform quality control on the quantity and purity of the total RNA and detect the integrity of the RNA through Bioanalyzer 2100 (Agilent, CA, USA); a concentration > 50 ng / μL, RIN value > 7.0, and total RNA > 1 μg meet the requirements of downstream experiments.
[0076] Specific capture of PolyA (polyadenylic acid)-containing mRNA therein was performed by two rounds of purification using oligo(dT) magnetic beads (Dynabeads Oligo (dT), cat.25-61005, Thermo Fisher, USA). The captured mRNA was fragmented at high temperature using a magnesium ion fragmentation kit (NEBNextRMagnesium RNA FragmentationModule, cat.E6150S, USA) at 94 °C for 5 - 7 minutes. The fragmented RNA was synthesized into cDNA by the action of reverse transcriptase (Invitrogen SuperScriptTM II ReverseTranscriptase, cat.1896649, CA, USA). Then, E. coli DNA polymerase I (NEB, cat.m0209, USA) and RNaseH (NEB, cat.m0297, USA) were used for second-strand synthesis to convert these DNA-RNA hybrid double-strands into DNA double-strands. Meanwhile, dUTP Solution (Thermo Fisher, cat.R0133, CA, USA) was incorporated into the second strand to fill in the ends of the double-stranded DNA to blunt ends, and then an A base was added to each end to enable ligation with an adapter with a T base at the end, and magnetic beads were used to screen and purify the fragment sizes. The second strand was digested with UDG enzyme (NEB, cat.m0280, MA, US), and then PCR was performed - pre-denaturation at 95 °C for 3 minutes, denaturation at 98 °C for a total of 8 cycles, 15 seconds each, annealing to 60 °C for 15 seconds, extension at 72 °C for 30 seconds, and finally extension at 72 °C for 5 minutes to form a library (strand-specific library) with a fragment size of 300 bp ± 50 bp.
[0077] Finally, paired-end sequencing was performed on it using the Illumina NovaseqTM 6000 according to the standard operation, and the sequencing mode was PE150. The Valid Data that could be aligned to the reference genome could be defined as being aligned to exon, intron, and intergenic regions according to the regional information of the reference genome. Generally, for species with relatively well-annotated genomes (such as model species like humans and Arabidopsis thaliana), the percentage content of the sequencing sequences located in the exon region should be the highest, while the reads aligned to the intron and intergenic regions may be due to precursor mRNA splicing events, ncRNA (non-coding RNA), imperfect genome annotation, DNA contamination, and background noise, etc. It was found that the content of all samples aligned to the exon region reached more than 90%, indicating that the sequencing data quality was good and the sequencing results were reliable.
[0078] To determine whether there were clusters or outliers in the samples, the FactoMineR package was used to perform principal component analysis (PCA) for data dimensionality reduction on the transcriptome sequencing dataset to identify the discreteness of MN and MP samples through patternization and visualization. It was found that PC1 explained 22.37% of the variance, which means that the first principal component could capture relatively high variability in the data. The combined variance explanation of PC1 and PC2 was 29.97% (22.37% + 7.6%), which means that these two principal components could explain nearly one-third of the variability in the dataset.
[0079] To identify the differentially expressed genes between MN samples and MP samples, the "DESeq2" package was used to analyze the differentially expressed genes between MP and MN samples (MP group VS MN group) in the transcriptome dataset, denoted as DEGs (a total of 1374: 1057 up-regulated and 317 down-regulated in the MP group) (threshold: |log2FC| > 0.5, p-value < 0.05); subsequently, the R packages "ggplot2" and "ComplexHeatmap" were used to draw volcano plots and heatmaps of the DEGs (the top 10 up-regulated genes with the largest change in |log2FC|: SOHLH1, MUCL3, PNMA5, REG1A, FGA, PGC, DSCAM-AS1, IGHV3-22, LINC01320, TEKT4; the top 10 down-regulated genes: FOLR1P1, TLX1, MGAM2, KRT1, KRTDAP, LINC00167, MUCL1, LINC00457, BPIFC, PAK5).
[0080] Figure 1As shown, a volcano plot of DEGs was drawn using the R packages "ggplot2" and "ComplexHeatmap". The top 10 genes with the highest up- and down-regulation fold differences are shown in the volcano plot, with red indicating high expression, green indicating low expression, and gray indicating genes with no significant difference.
[0081] Subsequently, based on the key genes, the AUC values of the ROC curves of the nomogram models constructed with different combinations were calculated to select the gene combination with the highest AUC value. It was found that the diagnostic model constructed with the combination of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28 had the highest AUC value (0.8704). The nomogram consists of "Score" and "Total Score". The former represents the score of each key gene, and the latter represents the sum of the scores of all key genes. The higher this value, the higher the probability of lymph node metastasis in CC. In the analysis, the "rms" package was used to draw the calibration curve, the "ggDCA" package was used to draw the DCA decision curve, and the "pROC" package was used to draw the ROC curve to evaluate the goodness of fit of the model. The results showed that the nomogram had good predictive ability (HL test p.value > 0.05, the benefit of the DCA curve was basically above ALL and NONE, and the AUC value = 0.8704).
[0082] Figure 2 Shown is the diagnostic nomogram of the key genes obtained by screening, that is, the diagnostic nomogram of seven biomarkers for lymph node metastasis in cervical cancer: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. The top layer of the nomogram represents the score (Points) of each key gene. The corresponding score is obtained according to the value of the key gene, and then the scores are added together to get the "Total points". The higher the total score, the higher the predicted risk. The bottom shows the risk probability (Pr(Y)) calculated based on the total score.
[0083] Figure 3 Shown is the calibration curve of the nomogram of the key genes obtained by screening, that is, the calibration curve of the nomogram of seven biomarkers for lymph node metastasis in cervical cancer: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28. The X-axis represents the predicted probability, and the Y-axis represents the actual probability. The blue curve is the actual prediction result (Apparent), the black curve is the result after bias correction (Bias-corrected), and the dashed line is the ideal situation (Ideal), that is, the reference line where the model prediction is completely consistent with the actual situation. The p value of the Hosmer-Lemeshow test was 0.315, indicating that the calibration performance of the model was good and did not deviate significantly from the ideal situation.
[0084] In this embodiment, the Logistic regression model is constructed using the rms package, and the nomogram of the regression model is plotted using the regplot package. Using the Logistic regression model and the log2 value of the CPM of gene expression, the expression of the following diagnostic model is obtained:
[0085] Subsequently, the OR value of each factor in the diagnostic model is calculated as shown in the following table.
[0086]
[0087] Figure 4 is the ROC curve of the diagnostic model. The ROC (Receiver Operating Characteristic) curve is used to evaluate the performance of the diagnostic model, where the horizontal axis is 1 - Specificity (false positive rate) and the vertical axis is Sensitivity (true positive rate); the gray dashed line in the figure represents the baseline of random prediction, and the black dot marks the optimal cutoff point, whose corresponding logit(p) value is -0.011, the corresponding probability is about 0.497, the Sensitivity is 0.815, and the Specificity is 0.750; the red curve is the ROC curve of the diagnostic model, and the area under the curve (AUC, Area Under the Curve) is about 0.8704, indicating that the model has good discrimination ability. The score calculated by the model > cutoff is defined as high - risk, and less than is low - risk. Conversely, it is high - risk, and high - risk cases are treated actively.
[0088] Example 2: Establishment and verification of a computational model for evaluating the risk of cervical cancer lymph node metastasis In this embodiment, the cervical tissue transcriptome sequencing data of 27 cervical cancer patients with lymph node metastasis and 32 patients without lymph node metastasis in Example 1 are used. Differential expression analysis is performed through DESeq2 to screen out candidate genes with significant differential expression in metastatic and non - metastatic samples. Further, the LASSO regression model is used for variable selection, and finally 7 gene markers with high predictive value are determined: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28. Using the expression levels of the above 7 genes as feature variables, a binary classification model is trained using the random forest algorithm. The performance of the model is evaluated using 10 - fold cross - validation, and the ratio of the training set to the test set is 7:3. In the independent test set, the AUC (area under the curve) of the model is 0.93, indicating that the model has extremely high predictive efficacy in differentiating whether cervical cancer patients have lymph node metastasis.
[0089] Example 3: Comparative Analysis of External Independent Sample Validation of Model Effectiveness and Traditional Methods To further verify the reliability and generalization ability of the risk calculation model for cervical cancer lymph node metastasis described in Example 2, this example introduces a set of external independent validation samples and compares and analyzes the model prediction results with traditional imaging evaluation methods.
[0090] I. Collection of External Validation Samples In this example, a total of 20 independent cervical cancer patient samples were included, including 10 cases in the metastasis group and 10 cases in the non - metastasis group. All samples did not participate in the model training process, and the sample processing method was the same as that in Example 1.
[0091] II. Detection of Gene Expression Data Referring to the process described in Example 1, sample RNA was extracted and qRT - PCR was performed to obtain the expression levels of 7 key genes of the model (CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28).
[0092] III. External Prediction Results of the Model The standardized expression data was input into the model trained in Example 2, and the model automatically output the predicted probability of lymph node metastasis for each patient. The results are as follows: AUC (ROC): 0.912; accuracy: 90.0%; sensitivity: 93.3%; specificity: 86.7%.
[0093] The above results indicate that the model can still maintain excellent prediction ability in new independent samples, with good stability and generalization ability.
[0094] Example 4: Implementation of a Computer System Based on the Model of the Present Invention In this example, a computer system was constructed to run the above - mentioned risk prediction model. The system includes the following modules (see Figure 5 ): Data Import Module: Used to read the standardized gene expression data of patient samples, supporting formats such as CSV and Excel.
[0095] Pre - processing Module: Includes functions such as missing value processing and standardized conversion to ensure the consistency of input data.
[0096] Model Calculation Module: Loads the pre - trained random forest model file, infers the input data, and outputs the predicted risk score and probability for each sample.
[0097] Visualization Module: Displays the prediction results in the form of bar charts, heat maps, etc., to facilitate doctors to quickly understand the results.
[0098] Interface module: can be connected to the hospital information system (HIS) or laboratory data system (LIMS) to realize automatic data flow and automatic triggering of risk assessment.
[0099] The system is deployed on a hospital server or private cloud environment. The front-end interface is developed using the Django framework, and the back-end model is implemented using the scikit-learn library in Python.
[0100] Example 5: Integrated lymph node metastasis risk assessment device The inventor of the present invention has designed a smart device for assessing the risk of lymph node metastasis in cervical cancer (see Figure 6 ), mainly including: Detection module: Built-in fluorescence quantitative PCR system, used to detect the expression levels of the above 7 genes in cervical cancer tissue samples, with optional automated RNA extraction unit.
[0101] Data acquisition module: connects to the detection module and transmits the expression data to the main control system in real time.
[0102] Main control processing module: an embedded processor running the model program of the present invention, integrating the Linux system, having 8GB of memory, and equipped with a dedicated computing acceleration chip (such as the ARM Cortex-A72 series).
[0103] Output display module: Built-in touch screen display interface, which can display the prediction score, transfer probability, result suggestion and visual report in real time.
[0104] Communication module: supports WiFi, Bluetooth, USB and Ethernet interfaces to facilitate communication with hospital database or host computer.
[0105] After the device is started, the operator only needs to import the sample and click the detection and prediction button to complete the full process evaluation and provide timely auxiliary basis for clinical surgical decisions.
[0106] The above descriptions are merely embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A computational model for evaluating the risk of cervical cancer lymph node metastasis, which takes the expression levels of multiple gene markers as input and outputs the metastasis risk score of the patient; the gene markers are selected from the group consisting of the following gene markers: CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, RPS28.
2. The computational model according to claim 1, wherein the model is established based on a machine learning algorithm, and the algorithm is selected from random forest, logistic regression, LASSO regression, COX regression or neural network algorithm.
3. A computer program, which is used to implement the evaluation function of the computational model according to claim 1 or 2 when executed by a processor, and the program includes the following steps: a) Receiving the gene expression data of the sample to be tested; b) Performing data preprocessing and standardization; c) Inputting the standardized data into the prediction model; d) Outputting the lymph node metastasis risk score and prediction probability corresponding to the sample.
4. A computer system for evaluating the risk of cervical cancer lymph node metastasis, which includes a processor, a memory, and a program module stored in the memory and executed by the processor. The program module implements the steps according to claim 3 and is used to implement the risk assessment of the target sample.
5. A method for constructing a diagnostic model for cervical cancer lymph node metastasis, characterized in that It includes the following steps: Step 1: Select cervical tissue samples from patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis; Step 2: Perform gene expression detection on the tissue samples obtained in Step 1; Step 3: Compare the gene expression differences between patients with cervical cancer lymph node metastasis and patients without cervical cancer lymph node metastasis according to the gene expression detection data obtained in Step 2; Step 4: Screen out key genes based on bioinformatics analysis methods; Step 5: Draw a diagnostic nomogram of the key genes, obtain the corresponding scores according to the values of the key genes, then add up the scores to get the total score, and calculate the risk probability of cervical cancer lymph node metastasis according to the total score. Among them, in Step 4, the key genes include at least 1, 2, 3, 4, 5, 6, or 7 of CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
6. According to the method of claim 5, in step 5, the risk probability of cervical cancer lymph node metastasis is deduced from the total score obtained from the diagnostic nomogram of key genes. The higher the total score, the higher the risk of predicting cervical cancer lymph node metastasis.
7. An evaluation device for evaluating the risk of cervical cancer lymph node metastasis, characterized in that It includes: A data acquisition device for acquiring detection data of an evaluation object, where the detection data is the level value of a biomarker in a sample of an evaluation object suspected of having cervical cancer lymph node metastasis; A data processing device for calculating the index value of the biomarker in the cervical cancer lymph node metastasis biomarker according to the detection data of the evaluation object, and then determining whether the calculated index value of the biomarker is within the risk range of the corresponding biomarker; The cervical cancer lymph node metastasis biomarkers are CARD9, CFL1P1, GRASLND, MNX1_AS2, MRAS, OLFML2A, and RPS28.
8. According to the evaluation device of claim 7, characterized in that Among them, the device further includes a detection device for detecting the level value of the biomarker.
9. The evaluation device according to claim 7, wherein The data acquisition device is an input device for inputting data or a communication device for reading data from an external data storage device or the storage device of the detection device through an interface.
10. The evaluation device according to claim 9, wherein When using a communication device, its corresponding external data storage device is the data memory on the device for measuring biomarkers.
Citation Information
Patent Citations
Micronucleus DNA of peripheral red blood cells and application of micronucleus DNA
CN112094907A
Group of markers for predicting lymph node metastasis risk of cervical cancer
CN117070625A
Equipment, method and system for diagnosing lymphatic metastasis of cervical cancer and application of equipment, method and system
CN118471335A
Method of determining lymph node metastasis in cervical cancer, device for determining the same, and computer program
US20120095301A1
Novel peptides, combination of peptides and scaffolds for use in immunotherapeutic treatment of various cancers
US20180250373A1