A method, system and device for evaluating recurrence risk and prognosis of a patient with resectable non-small cell lung cancer based on plasma cfDNA fragmentomics
By performing high-throughput sequencing and machine learning on plasma cfDNA from patients with operable NSCLC, CsFMA features of the GATA3 and CEP57 genes were screened out, and a Cox proportional hazards regression model was constructed. This solved the problem that the TNM staging system could not reflect molecular heterogeneity, and enabled accurate assessment of recurrence risk and prognosis in NSCLC patients, providing important clinical guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GENESEEQ TECH INC
- Filing Date
- 2026-02-27
- Publication Date
- 2026-07-03
AI Technical Summary
The existing TNM staging system cannot reflect the heterogeneity of tumors at the molecular level, resulting in inaccurate assessment of postoperative recurrence and metastasis risk in patients with operable non-small cell lung cancer. Some patients miss the opportunity for a cure or suffer unnecessary treatment side effects.
High-throughput sequencing was performed on preoperative plasma from operable NSCLC patients to obtain DNA fragment omics features. Features were screened using the random survival forest algorithm, and key gene features were selected by combining machine learning algorithms. Fragment end pattern analysis (CsFMA) features based on GATA3 and CEP57 genes were established, and a Cox proportional hazards regression model was constructed to assess recurrence risk and prognosis.
It has achieved effective prediction of recurrence risk and prognosis in patients with operable non-small cell lung cancer, providing important guidance for clinical treatment. The model has shown good discriminative performance in the training set, internal validation set and external validation set.
Smart Images

Figure CN122337337A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, system, and device for assessing recurrence risk and prognosis in patients with operable non-small cell lung cancer based on plasma circulating cell-free DNA (cfDNA) fragment omics, belonging to the field of molecular biomedical technology. Background Technology
[0002] Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related death worldwide[1]. For patients with operable stage I-III NSCLC, radical surgical resection is the primary treatment. However, 30%-60% of patients experience recurrence and metastasis after surgery, which is the main reason for treatment failure[2]. Currently, the core basis for assessing postoperative recurrence risk and guiding adjuvant therapy decisions in clinical practice is the AJCC TNM staging system[3]. However, TNM staging is a macroscopic classification based on anatomy and cannot reflect the heterogeneity of tumors at the molecular level, resulting in huge differences in prognosis among patients within the same stage[4,5]. This creates a clinical dilemma: some low-stage patients may miss the opportunity for a cure due to insufficient treatment, while some high-stage patients may suffer unnecessary toxic side effects from adjuvant therapy.
[0003] In recent years, liquid biopsy technology, especially plasma cfDNA-based detection, has brought revolutionary changes to non-invasive tumor diagnosis and prognostic assessment [6-8]. Fragmentomics of cfDNA is an emerging research direction that focuses on the size distribution, terminal sequence, methylation, nucleosome protection patterns, and other characteristics of cfDNA fragments. These characteristics are closely related to gene expression and chromatin open state, and can reflect tumor biological activities without relying on specific mutations [9-12]. Currently, some studies have suggested that fragmentation features can be used for early tumor screening and tissue origin determination [8], but interpretable, stable, and engineerable feature models for postoperative recurrence risk and prognostic assessment are still lacking.
[0004] References
[0005] [1] Bray, F., Laversanne, M., Sung, H., et al. (2024). Global cancerstatistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians 74, 229-263.
[0006] [2] Chansky, K., Detterbeck, F.C., Nicholson, A.G., et al. (2017).The IASLC Lung Cancer Staging Project: External Validation of the Revision ofthe TNM Stage Groupings in the Eighth Edition of the TNM Classification ofLung Cancer. Journal of thoracic oncology : official publication of theInternational Association for the Study of Lung Cancer 12, 1109-1121.
[0007] [3] Kim, I.H., Lee, G.D., Choi, S., et al. (2024). Validation Studyfor the N Descriptor of the Newly Proposed Ninth Edition of the TNM StagingSystem Proposed by the International Association for the Study of LungCancer. Journal of thoracic oncology : official publication of theInternational Association for the Study of Lung Cancer 19, 1218-1227.
[0008] [4] Wang, C., Shao, J., Song, L., et al. (2023). Persistent increaseand improved survival of stage I lung cancer based on a large-scale real-world sample of 26,226 cases. Chinese medical journal 136, 1937-1948.
[0009] [5] Liu, S.Y., Bao, H., Wang, Q., et al. (2021). Genomic signaturesdefine three subtypes of EGFR-mutant stage II-III non-small-cell lung cancerwith distinct adjuvant therapy outcomes. Nature communications 12, 6450.
[0010] [6] Xia, L., Mei, J., Kang, R., et al. (2022). Perioperative ctDNA-Based Molecular Residual Disease Detection for Non-Small Cell Lung Cancer: AProspective Multicenter Cohort Study (LUNGCA-1). Clinical cancer research :an official journal of the American Association for Cancer Research 28, 3308-3317.
[0011] [7] Chen, K., Yang, F., Shen, H., et al. (2023). Individualizedtumor-informed circulating tumor DNA analysis for postoperative monitoring ofnon-small cell lung cancer. Cancer cell 41, 1749-1762.e1746.
[0012] [8] Bao, H., Yang, S., Chen, X., et al. (2025). Early detection ofmultiple cancer types using multidimensional cell-free DNA fragmentomics.Nature medicine 31, 2737-2745.
[0013] [9] Annapragada, A.V., Niknafs, N., White, J.R., et al. (2024). Genome-wide repeat landscapes in cancer and cell-free DNA. Science translational medicine 16, eadj9283.
[0014]
[10] Doebley, A.L., Ko, M., Liao, H., et al. (2022). A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nature communications 13, 7475.
[0015]
[11] Zhou, Q., Kang, G., Jiang, P., et al. (2022). Epigenetic analysis of cell-free DNA by fragmentomic profiling. Proceedings of the National Academy of Sciences of the United States of America 119, e2209852119.
[0016]
[12] Mouliere, F., Chandrananda, D., Piskorz, A.M., et al. (2018). Enhanced detection of circulating tumor DNA by fragment size analysis. Science translational medicine 10. Summary of the Invention
[0017] This invention provides a method, system, and device for assessing recurrence risk and prognosis in patients with operable non-small cell lung cancer (NSCLC). High-throughput sequencing is performed on preoperative plasma from operable NSCLC patients to obtain DNA fragment omics characteristics. A random survival forest algorithm is used for preliminary screening of genome-wide CsFMA features. Subsequently, machine learning algorithms are employed for variable selection of candidate features, ultimately determining the fragment end pattern features (CsFMA) of the GATA3 and CEP57 genes. A risk scoring model is then established using Cox proportional hazards regression. Based on confirmed thresholds, patients are divided into high-risk and low-risk groups, achieving effective prediction and stratification of patient recurrence risk and prognosis, providing important guidance for clinical treatment.
[0018] A method for constructing a model for assessing recurrence risk and prognosis in patients with operable non-small cell lung cancer, comprising the following steps:
[0019] S1. Obtain high-throughput sequencing data of preoperative plasma cfDNA samples from multiple operable non-small cell lung cancer patients, as well as the corresponding follow-up clinical information of the patients, wherein the clinical information includes at least disease-free survival (DFS) or overall survival (OS).
[0020] S2. Align the sequencing data to the human reference genome to determine the location and sequence information of the cfDNA fragment in the genome;
[0021] S3. For each candidate gene on the reference genome, perform the following feature calculation procedure: (1) Identify all CpG sites within the gene range;
[0022] (2) For each CpG site, count the number of fragments with a 5' break sequence of "CGN" in the cfDNA fragments covering that site as the numerator, and count the number of fragments with a 5' break sequence of "NCG" in the trinucleotide pattern as the denominator. Calculate the ratio between the two to obtain the initial characteristic value of that CpG site; N represents A / T / C / G.
[0023] (3) Calculate the arithmetic mean of the initial feature values of all CpG sites within the gene range, and use it as the original CsFMA feature of the gene;
[0024] (4) The original CsFMA features of the gene were Z-score normalized based on the sample set used for model construction to obtain the final CsFMA feature values of the gene.
[0025] S4. The genome-wide CsFMA features obtained in step S3 are initially screened, and variable selection is performed using machine learning algorithms to ultimately screen out key feature genes that are significantly related to patient prognosis.
[0026] Step S5. Using the CsFMA eigenvalues of key characteristic genes significantly associated with patient prognosis as input variables and the patient's clinical prognostic information as dependent variables, a Cox proportional hazards regression model is used for fitting to determine the regression coefficients of each characteristic; the final evaluation model formula contains the CsFMA eigenvalues of the key characteristic genes significantly associated with patient prognosis that have been selected.
[0027] The key characteristic genes were identified as GATA3 and CEP57.
[0028] The final assessment model formula is: Risk value = m × CsFMA.GATA3 eigenvalue - n × CsFMA.CEP57 eigenvalue; m and n are coefficients; CsFMA.GATA3 eigenvalue and CsFMA.CEP57 eigenvalue are the values calculated and standardized as described in step S3.
[0029] The value of m ranges from 1.9 to 3.1, and the value of n ranges from 0.9 to 2.1.
[0030] A system for assessing recurrence risk and prognosis in patients with operable non-small cell lung cancer includes:
[0031] The sequencing data acquisition module is used to acquire high-throughput sequencing data of plasma cfDNA from patients to be evaluated.
[0032] The alignment module is used to align the sequencing data to a reference genome to determine the chromosomal location, start and end sites, and sequence information of the cfDNA fragment;
[0033] The feature calculation module is used to calculate the CsFMA feature values of the GATA3 and CEP57 genes. Specifically, the feature calculation module includes:
[0034] The site identification unit is used to identify all CpG sites within the GATA3 and CEP57 gene range of the reference genome; the end sequence counting unit is used to traverse the cfDNA fragments aligned to each identified CpG site, and count the number of first fragments with a 5' end break sequence of "CGN" and the number of second fragments with a 5' end break sequence of "NCG".
[0035] The site ratio calculation unit is used to calculate the ratio of the number of the first fragment to the number of the second fragment, and obtain the site feature value of each CpG site;
[0036] The gene feature summary unit is used to calculate the arithmetic mean of the site feature values of all CpG sites within the GATA3 and CEP57 gene ranges, respectively.
[0037] The standardization processing unit is used to perform Z-score standardization on the arithmetic mean, using Z-score parameters consistent with the model construction, and outputs the final CsFMA.GATA3 eigenvalues and CsFMA.CEP57 eigenvalues.
[0038] The risk assessment module internally stores a risk scoring formula, which includes CsFMA feature values of GATA3 and CEP57; the risk assessment module is used to substitute the feature values output by the feature calculation module into the formula to calculate the risk value.
[0039] The results output module is used to compare the calculated risk value with a preset threshold. When the risk value is higher than the preset threshold, it outputs an assessment result of high recurrence risk or poor prognosis.
[0040] The beneficial effects of this invention are:
[0041] Preoperative blood samples were collected from 138 NSCLC patients, and cfDNA was extracted for high-throughput sequencing to obtain fragmented features. The random survival forest (RSF) algorithm was used for initial screening of genome-wide CsFMA features. Subsequently, machine learning algorithms (Lasso, Enet, plsRcox, CoxBoost, or Ridge) were used for variable selection of candidate features. Finally, a risk scoring model was established using Cox proportional hazards regression. This invention, based on plasma cfDNA fragmentation results, can effectively predict the recurrence risk and prognosis of operable non-small cell lung cancer. Attached Figure Description
[0042] Figure 1 Study the roadmap.
[0043] Figure 2 The consistency index (C-index) of 110 algorithm combinations on the training set and internal validation set.
[0044] Figure 3 Area under the curve (AUC) of the predicted 1-year and 2-year DFS and OS on the training set and internal validation set. AUC curves of the predicted 1-year and 2-year DFS on the training set (A); AUC curves of the predicted 1-year and 2-year DFS on the internal validation set (B); AUC curves of the predicted 1-year and 2-year OS on the training set (C); AUC curves of the predicted 1-year and 2-year OS on the internal validation set (D).
[0045] Figure 4Survival curves for DFS or OS in patients with different risk stratifications in each dataset. In the training set, patients in the high-risk group had shorter DFS (A) and OS (B); in the internal validation set, patients in the high-risk group had shorter DFS (C) and OS (D); in the external independent validation set, AUC curves for predicting 1-year and 2-year OS (E); in the external independent validation set, patients in the high-risk group had shorter OS (F). Detailed Implementation
[0046] This invention first requires the extraction, library construction, and sequencing of cfDNA from blood samples. The invention uses the QIAamp Circulating Nucleic Acid Kit (Qiagen) to extract genomic DNA from plasma samples, then uses a Qubit 3.0 fluorometer and a dsDNA HS Assay Kit (Thermo Fisher Scientific) to measure the amount of extracted DNA, and finally uses the KAPA Hyper Prep Kit (KAPA Biosystems) for library construction.
[0047] Dataset situation during model building process
[0048] The research route is as follows: Figure 1 As shown, a total of 138 patients with operable non-small cell lung cancer (NSCLC) were included in the model development, all of whom had plasma samples collected preoperatively. They were divided into a training set of 88 patients and an internal validation set of 50 patients, based on their enrollment time, for model construction and initial validation. To further evaluate the model's generalization and robustness in prognostic prediction, an independent cohort of NSCLC patients was introduced as an external validation set. Patient information for the study is shown in the table below:
[0049] Table 1. Basic clinical characteristics of patients
[0050]
[0051] The training and internal validation sets both consisted of patients with operable stage I-III NSCLC, with a median age of 62 years (range 33-79 and 40-79 years), and the proportions under 60 years old were 35.2% and 38.0%, respectively. The male-to-female ratio was similar (50.0% and 52.0% female), and a high percentage of non-smokers (76.1% and 78.0%). Adenocarcinoma was the predominant disease in both cohorts (81.8% and 88.0%), and approximately half of the patients had received adjuvant therapy (55.7% and 56.0%). The external independent validation set was a stage IV adenocarcinoma cohort, with a median age of 61 years (range 31-86 years), and 44.9% under 60 years old, with a high proportion of smokers.
[0052] Patient plasma cfDNA extraction, sequencing, and analysis:
[0053] Preoperatively, liquid biopsies were performed on patients, and 10 ml of whole blood samples were collected using EDTA anticoagulant tubes. Plasma was immediately separated by centrifugation (within 2 hours) and frozen at -80°C before being transferred to the laboratory for analysis. Upon arrival at the laboratory, ctDNA was extracted from the plasma samples using the QIAGEN plasma DNA extraction kit according to the instructions. After library construction, high-throughput sequencing (paired-end sequencing, deduplicated, fragments formed using properly paired reads, with the fragment origin at the smaller coordinate end and the endpoint at the larger coordinate end) was performed on the HiSeq 4000 (Illumina) platform. After obtaining the raw fastq data, quality control (removal of low-quality bases (<20) or N bases) was performed using Trimmomatic software to obtain clean fastq data. After quality control, the data were aligned to the human reference genome (hg19 version) using Burrows-Wheller Aligner (BWA-mem, v0.7.12) software with default parameters to generate BAM files. The BAM files were then sorted according to genomic coordinates and duplicates were removed using Picard software (https: / / broadinstitute.github.io / picard / ) to obtain the base data of the corresponding reads. The average sequencing depth of all samples after deduplication was 7500 X.
[0054] The marker data of this invention mainly utilizes the CsFMA information of cfDNA as the model input feature: CsFMA feature value is an indicator that quantitatively describes the fragmentation pattern of cfDNA fragments by analyzing the cleavage preference of cfDNA fragments near specific genomic sites. The calculation process is as follows: First, all CpG sites are identified within the target gene range; then, the number of cfDNA fragments that can align to the 5' break sequence of each CpG site with "CGN" and "NCG" (where N represents A, T, C, or G) is counted. Specifically, the three base sequences at the 5' break are counted. If the 3bp break sequence satisfies the "CGN" pattern (where N represents A, T, C, or G), it is considered a "CGN" sequence. If the 3bp break sequence satisfies the "NCG" pattern (where N represents A, T, C, or G), it is considered an "NCG" sequence. Finally, the initial CsFMA value for each CpG locus is obtained by calculating the ratio of the "CGN" fragment count to the "NCG" fragment count (when the NCG count is 0, CGN / NCG is directly set to NA, and downstream operations and the final feature value are also treated as missing values of NA. Points with NA are not included when calculating the mean). After obtaining the initial feature values for all CpG loci, the feature values of all CpG loci covered by each gene are further averaged to generate gene-level CsFMA features. Subsequently, the CsFMA features of each gene are Z-score standardized so that the standardized feature values can be used for subsequent model construction. After the model is built, the standardized parameters are stored as part of the model parameters, and the same parameters are called when performing risk assessment on new samples to ensure consistency in the standardization method.
[0055] Model Feature Filtering
[0056] Using the R package 'Mime1' (version 0.0.0.9000) on the training set, Cox univariate analysis was first performed on the CsFMA features of candidate genes, revealing 11 features significantly associated with disease-free survival (DFS). Higher feature values for GSTP1, CEP57, PPP2R1A, GRM3, and MPL genes indicated longer DFS, while the remaining features were associated with shorter DFS (Table 2). Subsequently, using these 11 significant features as candidates, nine algorithms built into the R package 'Mime1' (version 0.0.0.9000)—RFS, Enet, Lasso, Ridge, Stepwise Cox, CoxBoost, plsRcox, SuperPC, and Survival SVM—were used to construct 110 model combinations with default parameters. Their C-index performance on the training and internal validation sets was compared to select the optimal solution. The results are as follows: Figure 2As shown, the five combinations of RSF+Lasso, RSF+Enet, RSF+plsRcox, RSF+CoxBoost, and RSF+Ridge all performed optimally on both the training and internal validation sets, and the selection results were consistent. More importantly, the key gene features selected by these five combinations and the final performance were highly consistent. This cross-algorithm consistency evidence strongly confirms the stability and predictive value of the obtained gene signatures.
[0057] Table 2. Genetic characteristics that are significant under single-factor Cox influence.
[0058]
[0059] Model building and performance verification
[0060] Based on consensus screening using multiple algorithms (RSF+Lasso, RSF+Enet, RSF+plsRcox, RSF+CoxBoost, RSF+Ridge), the fragmentomics biomarker with stable prognostic value was ultimately determined to be CsFMA.GATA3 (GATA3 gene chromosomal location: Chr10: 8,095,567-8,117,161 ) and CsFMA.CEP57 (CEP57 gene chromosomal location: Chr 11: 95, 523,129-95,565,857 Based on the above two characteristics, a risk scoring model was constructed using Cox proportional hazards regression. The evaluation was performed using the model score as a continuous variable, with the risk value calculated as 3.02 × CsFMA.GATA3 - 1.02 × CsFMA.CEP57. The area under the curve (AUC) for predicting 1-year and 2-year DFS in the training set was both >0.85. Figure 3 A), and higher than the prediction performance of single features of CsFMA.GATA3 or CsFMA.CEP57 (Table 3); AUC > 0.70 in the internal validation set. Figure 3 B). Furthermore, the model's risk score was used to predict overall survival (OS) of patients. In the training set, the AUC for predicting 1-year and 2-year OS was both >0.80 (B). Figure 3 C), and its predictive performance is higher than that of single features of CsFMA.GATA3 or CsFMA.CEP57 (Table 3); AUC is >0.70 in the internal validation set. Figure 3 D). The above results indicate that the model has good discriminative performance for both recurrence risk and prognosis.
[0061] Table 3. 1-year and 2-year DFS and OS AUC data for CsFMA.GATA3 and CsFMA.CEP57 feature predictions
[0062]
[0063] In the training set, patients were divided into high-risk and low-risk groups using a ternary threshold (4.09), with values above the threshold defined as high-risk. Kaplan-Meier analysis showed that the high-risk group had significantly shorter disease-free survival (DFS), with a median DFS of 42.7 months vs. not reached (HR=8.03, P<0.001); overall survival (OS) was also worse, with a median OS of 42.7 months vs. not reached (HR=13.29, P=0.008). Figure 4 A and B). Using the same threshold in the internal validation set, the results were consistent: the high-risk group had significantly shorter DFS and OS. Figure 4 (C and D). Given the limited maturity of OS data in the training and internal validation sets, we further included an external validation set for stage IV NSCLC (n=107). The AUC of this risk score against 1-year and 2-year OS in the external validation set was 0.777 and 0.844, respectively. Figure 4 E). Grouped by the same threshold, the high-risk group had significantly worse OS: median OS was 22.46 months vs 65.45 months (HR=4.73, P<0.001). Figure 4 F). The above results indicate that the model has stable recurrence and prognostic capabilities in both internal and external cohorts.
[0064] The constructed risk scoring model was further incorporated into a Cox proportional hazards regression multivariate analysis along with clinical characteristics (such as age, gender, and tumor stage). The results showed that the risk scoring model maintained statistical significance in the multivariate analyses related to DFS and OS (P values less than 0.05), suggesting that it can serve as an independent predictor of patient DFS and OS.
[0065] Table 4. Results of multivariate analysis of DFS and OS Cox
[0066]
[0067] The above is an example of the implementation process of this patent and does not constitute a limitation on the scope of protection of this patent.
Claims
1. A method for constructing a model for evaluating recurrence risk and prognosis of a patient with operable non-small cell lung cancer, characterized in that, Includes the following steps: S1. Obtain high-throughput sequencing data of preoperative plasma cfDNA samples from multiple operable non-small cell lung cancer patients, as well as the corresponding follow-up clinical information of the patients, wherein the clinical information includes at least disease-free survival (DFS) or overall survival (OS). S2. Align the sequencing data to the human reference genome to determine the location and sequence information of the cfDNA fragment in the genome; S3. For each candidate gene on the reference genome, perform the following feature calculation procedure: (1) Identify all CpG sites within the gene range; (2) For each CpG site, count the number of fragments with the 5' end break sequence "CGN" in the cfDNA fragments covering the site as the numerator, count the number of fragments with the 5' end break sequence "NCG" as the denominator, calculate the ratio between the two, and obtain the initial characteristic value of the CpG site. (3) Calculate the arithmetic mean of the initial feature values of all CpG sites within the gene range, and use it as the original CsFMA feature of the gene; (4) The original CsFMA features of the gene were Z-score normalized to obtain the final CsFMA feature values of the gene. S4. The genome-wide CsFMA features obtained in step S3 are initially screened, and variable selection is performed using machine learning algorithms to ultimately screen out key feature genes that are significantly related to patient prognosis. Step S5. Using the CsFMA eigenvalues of key characteristic genes significantly associated with patient prognosis as input variables and the patient's clinical prognostic information as dependent variables, a Cox proportional hazards regression model is used for fitting to determine the regression coefficients of each characteristic; the final evaluation model formula contains the CsFMA eigenvalues of the key characteristic genes significantly associated with patient prognosis that have been selected.
2. The construction method according to claim 1, characterized in that, The key characteristic genes were identified as GATA3 and CEP57.
3. The construction method of claim 2, wherein, The final assessment model formula is: Risk value = m × CsFMA.GATA3 eigenvalue - n × CsFMA.CEP57 eigenvalue; m and n are coefficients; CsFMA.GATA3 eigenvalue and CsFMA.CEP57 eigenvalue are the values after calculation and standardization as described in step S3.
4. The construction method of claim 2, wherein, The value of m ranges from 1.9 to 3.1, and the value of n ranges from 0.9 to 2.
1.
5. A system for assessing recurrence risk and prognosis in patients with operable non-small cell lung cancer, characterized in that, include: The sequencing data acquisition module is used to acquire high-throughput sequencing data of plasma cfDNA from patients to be evaluated. The alignment module is used to align the sequencing data to a reference genome to determine the chromosomal location, start and end sites, and sequence information of the cfDNA fragment; The feature calculation module is used to calculate the CsFMA feature values of the GATA3 and CEP57 genes. Specifically, the feature calculation module includes: The site identification unit is used to identify all CpG sites within the GATA3 and CEP57 gene range of the reference genome; the end sequence counting unit is used to traverse the cfDNA fragments aligned to each identified CpG site, and count the number of first fragments with a 5' end break sequence of "CGN" and the number of second fragments with a 5' end break sequence of "NCG". The site ratio calculation unit is used to calculate the ratio of the number of the first fragment to the number of the second fragment, and obtain the site feature value of each CpG site; The gene feature summary unit is used to calculate the arithmetic mean of the site feature values of all CpG sites within the GATA3 and CEP57 gene ranges, respectively. The standardization processing unit is used to perform Z-score standardization on the arithmetic mean and output the final CsFMA.GATA3 eigenvalue and CsFMA.CEP57 eigenvalue; The risk assessment module internally stores a risk scoring formula, which includes CsFMA feature values of GATA3 and CEP57; the risk assessment module is used to substitute the feature values output by the feature calculation module into the formula to calculate the risk value. The results output module is used to compare the calculated risk value with a preset threshold. When the risk value is higher than the preset threshold, it outputs an assessment result of high recurrence risk or poor prognosis.
6. The recurrence risk and prognostic assessment system for operable non-small cell lung cancer patients according to claim 5, characterized in that, Risk value = m × CsFMA.GATA3 eigenvalue - n × CsFMA.CEP57 eigenvalue; m and n are coefficients.
7. The system for evaluating the risk of recurrence and the prognosis of operable non-small cell lung cancer patients according to claim 6, characterized in that, The value of m ranges from 1.9 to 3.1, and the value of n ranges from 0.9 to 2.1.