Lung adenocarcinoma osimertinib drug resistance prediction model construction method and system based on machine learning

By constructing a machine learning-based prediction model for osimertinib resistance in lung adenocarcinoma, using the TYMS and UAP1L1 gene expression levels, combined with residual correction and stacking integration methods, the problems of traumatic injury and high cost in existing technologies for monitoring osimertinib resistance in lung adenocarcinoma were solved, achieving a non-invasive, economical and accurate prediction effect.

CN120600101AActive Publication Date: 2025-09-05THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)

Patent Information

Application Number
CN202510760210.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-05
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In existing technologies, monitoring of osimertinib resistance in lung adenocarcinoma relies on imaging examinations and tissue biopsies, which are invasive and costly, and cannot accurately predict resistance in the early stages.

Method used

A machine learning-based prediction model for osimertinib resistance in lung adenocarcinoma was constructed. By integrating clinical data of lung adenocarcinoma patients and mRNA data from public databases, and using preprocessing such as the Mann-Whitney U test and FPKM normalization, the key genes TYMS and UAP1L1 were screened. The model performance was optimized by combining residual correction and stacking ensemble methods, and non-invasive serum testing was performed.

Benefits of technology

It achieves non-invasive, economical and accurate prediction of osimertinib resistance, with the model AUC increased to 0.924, providing early warning and personalized treatment plans, reducing testing costs and patient trauma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600101A_ABST
    Figure CN120600101A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical data analysis, in particular to a lung adenocarcinoma osimertinib drug resistance prediction model construction method and system based on machine learning. Clinical data of lung adenocarcinoma patients and mRNA (messenger ribonucleic acid) data of a public database are integrated, after Mann-Wh itney U inspection, FPKM standardization and other preprocessing are carried out, XGBoost is utilized to construct a basic model driven by clinical indexes, a prognosis model driven by gene characteristics is constructed through R packet Mime, and TYMS and UAP1L1 are screened out to serve as key genes. The performance of the model is optimized through residual correction and stacking integration, and finally the AUC is improved to 0.924. Serum RNA (Ribonucleic Acid) detection and verification show that TYMS and UAP1L1 are remarkably and highly expressed in a drug-resistant group. According to the method, traditional biopsy is replaced with noninvasive serum detection, the detection cost is reduced, the model generalization ability is verified through an independent queue, and an efficient and accurate solution is provided for early warning and personalized treatment of osimertinib drug resistance of the lung adenocarcinoma patient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data analysis, and in particular to a model construction method and system for predicting osimertinib resistance in lung adenocarcinoma based on machine learning. Background Art

[0002] Lung cancer is a common malignancy in the Chinese population, with lung adenocarcinoma being the most common type. Activating mutations in the epidermal growth factor receptor (EGFR) gene are common in lung adenocarcinoma. Osimertinib, a third-generation EGFR tyrosine kinase inhibitor, is widely used in frontline clinical settings, but drug resistance affects long-term patient outcomes. Current methods for monitoring drug resistance rely on imaging and tissue biopsies, which are invasive and expensive. An efficient, cost-effective, and non-invasive predictive method is urgently needed. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a model construction method and system for predicting osimertinib resistance in lung adenocarcinoma based on machine learning, so as to solve the problems in the existing technology that osimertinib resistance monitoring relies on imaging examinations and tissue biopsies, which are invasive, costly and cannot be accurately predicted in the early stage.

[0004] According to a first aspect of an embodiment of the present invention, a method for constructing a model for predicting osimertinib resistance in lung adenocarcinoma based on machine learning is provided, wherein the method comprises:

[0005] Clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases were collected;

[0006] performing data preprocessing on the clinical data collected from lung adenocarcinoma patients, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm;

[0007] Using the results of screening for indicators significantly associated with drug resistance, a basic prediction model based on clinical indicators was constructed. Using the preset univariate COX regression analysis to screen for statistically significant gene data, a prognostic model based on gene expression characteristics was constructed.

[0008] The gene features screened in the preset prognostic model based on the gene expression features are combined with the basic prediction model based on clinical indicators to optimize the model performance through a preset residual correction and stacking integration method; the gene features screened in the preset prognostic model based on the gene expression features include: TYMS and UAP1L1 gene expression levels;

[0009] The optimized model is used to predict the patient's risk of osimertinib resistance and output the predicted probability value.

[0010] Furthermore, the clinical data of the lung adenocarcinoma patient include gender, age, EGFR mutation site, blood routine, tumor markers, tumor stage, and osimertinib medication time information.

[0011] Furthermore, the osimertinib resistance-related mRNA data in the public database include mRNA data of cell lines before and after osimertinib resistance and mRNA data of lung adenocarcinoma patients.

[0012] Furthermore, the data preprocessing of the collected clinical data of lung adenocarcinoma patients, screening of indicators significantly associated with drug resistance, and screening of differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm include:

[0013] The clinical data collected from the lung adenocarcinoma patients were used to screen for indicators significantly associated with drug resistance using the Mann-Whitney U test;

[0014] The osimertinib resistance-related mRNA data in the public database were sequentially subjected to FPKM normalization, log2 conversion, and pyComBat de-batch processing, and differentially expressed genes were screened using the pyDEG algorithm.

[0015] Furthermore, the method utilizes the results of screening for indicators significantly associated with drug resistance to construct a basic prediction model based on clinical indicators, and utilizes the preset univariate COX regression analysis to screen for statistically significant gene data to construct a prognostic model based on gene expression characteristics, including:

[0016] Using the results of screening indicators significantly related to drug resistance, the data was processed and modeled using the Xgboost algorithm;

[0017] The pre-set univariate COX regression analysis was used to screen the statistically significant gene data, and the prognostic model was established using the R software package Mime.

[0018] Furthermore, the gene features screened in the preset prognostic model based on the gene expression features are combined with the basic prediction model based on clinical indicators, and the model performance is optimized by a preset residual correction and stacking integration method, including:

[0019] Using gene signatures screened in a preset prognostic model based on the gene expression signature; the gene signatures include: TYMS and UAP1 L1;

[0020] The RNA expression levels of TYMS and UAP1 L1 were examined using pre-specified serum samples;

[0021] Based on the detection results, the preset TYMS and UAP1 L1 gene expression levels were used as new features and merged with the residual data to train the secondary correction model to obtain the first processing results;

[0022] Utilizing the first processing result, the outputs of multiple base models are used as input, and the final prediction result is generated through meta-learning integration to optimize its model performance.

[0023] Furthermore, the preset method for measuring the expression levels of the TYMS and UAP1 L1 genes includes:

[0024] Total RNA was extracted using serum or plasma RNA extraction kit;

[0025] cDNA was synthesized by reverse transcription, and gene expression levels were detected using SYBR Green fluorescence quantitative PCR;

[0026] GAPDH was used as the internal reference gene, and the relative expression levels were calculated using the ΔΔCt method.

[0027] Furthermore, the construction of the basic prediction model based on clinical indicators also includes:

[0028] 5-fold cross validation was used to optimize the Xgboost model parameters, including learning rate, number of trees, and maximum depth;

[0029] Key clinical indicators, including NLR, lymphocyte percentage, and CEA, were determined by feature importance analysis.

[0030] Furthermore, the prognostic model of gene expression characteristics further includes:

[0031] A prognostic model was constructed using various algorithms in the Mime package;

[0032] Evaluate model performance using C-index and AUC to select the optimal algorithm combination;

[0033] The various algorithms in the Mime package include: random forest, elastic network, stepwise Cox regression, CoxBoost, partial least squares regression Cox, generalized boosted regression model, survival support vector machine, Ridge regression, and Lasso regression algorithm.

[0034] According to a second aspect of an embodiment of the present invention, a model construction system for predicting osimertinib resistance in lung adenocarcinoma based on machine learning is provided, which is applied to any of the above-mentioned model construction methods for predicting osimertinib resistance in lung adenocarcinoma based on machine learning, characterized in that the method comprises:

[0035] A data acquisition module is used to collect clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases;

[0036] a data preprocessing module for performing data preprocessing on the clinical data collected from patients with lung adenocarcinoma, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm;

[0037] The model building module is used to use the results of screening indicators significantly associated with drug resistance to build a basic prediction model based on clinical indicators, and to use the preset univariate COX regression analysis to screen statistically significant gene data to build a prognostic model based on gene expression characteristics;

[0038] A model fusion module is used to utilize gene features selected from a preset prognostic model based on the gene expression signature, combine it with the basic prediction model based on clinical indicators, and optimize its model performance through a preset residual correction and stacking integration method; the gene features selected from the preset prognostic model based on the gene expression signature include: TYMS and UAP1L1 gene expression levels;

[0039] The predictive analysis module is used to predict the patient's risk of osimertinib resistance using the optimized model and output the predicted probability value.

[0040] The technical solutions provided by the embodiments of the present invention may have the following beneficial effects:

[0041] By integrating the clinical data of lung adenocarcinoma patients with the mRNA data of public databases, and after preprocessing such as Mann-Whitney U test and FPKM normalization, XGBoost was used to build a basic model driven by clinical indicators, and the R package Mime was used to build a prognostic model driven by gene features. TYMS (Thymidylate Synthetase) and UAP1L1 were screened.

[0042] UDP-N-Acetylglucosamine Pyrophosphorylase 1Like 1 was selected as a key gene. Residual correction and stacking ensemble optimization were used to optimize model performance, ultimately increasing the AUC to 0.924, achieving accurate prediction. Serum RNA testing demonstrated significantly elevated expression of TYMS and UAP1L1 in the drug-resistant group. This method replaces traditional biopsy with noninvasive serum testing, and the model's generalizability has been validated in an independent cohort. This approach provides an efficient and precise solution for early warning and personalized treatment of osimertinib resistance in lung adenocarcinoma patients.

[0043] Specifically, non-invasive: based on routine blood tests and serum RNA testing, it avoids the trauma of tissue biopsy and reduces patient pain.

[0044] Efficiency: Through machine learning algorithms, large amounts of data can be quickly processed to achieve accurate prediction of drug resistance, providing a basis for timely clinical adjustments to treatment plans.

[0045] Accuracy: The prediction performance of the optimized model is significantly improved, with the residual correction model AUC reaching 0.873 and the stacked ensemble model AUC reaching 0.924, which are superior to traditional monitoring methods.

[0046] Economical: Utilize routine laboratory examination indicators to reduce testing costs and save medical resources.

[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0049] Figure 1 1 is a flow chart of a model construction method for predicting osimertinib resistance in lung adenocarcinoma based on machine learning according to an exemplary embodiment;

[0050] Figure 2 1 is a schematic diagram of an ROC curve for distinguishing osimertinib-sensitive and -resistant patients in a test set according to an exemplary embodiment;

[0051] Figure 3 1 is a schematic diagram showing the identification of common DEGs associated with osimertinib resistance from GEO data according to an exemplary embodiment;

[0052] Figure 4 This is a schematic diagram showing the ordering of the C-index of each model in different cohorts according to the average value of the C-index in the cohorts for constructing a prognostic model based on osimertinib resistance-related genes according to an exemplary embodiment;

[0053] Figure 5 is a schematic diagram showing the relationship between the risk score calculated by the optimal combination model StepCox[forward]+RSF and the patient outcomes in different cohorts according to an exemplary embodiment;

[0054] Figure 6 (A) is a schematic diagram of the top ten results with statistical significance (P<0.05) after GO analysis according to an exemplary embodiment;

[0055] Figure 71 is a schematic diagram showing a heat map of cluster analysis of target genes in a GEO dataset according to an exemplary embodiment of target gene expression verification (A);

[0056] Figure 8 1 is a schematic diagram of ROC curves of a basic model, a residual correction model, and a stacking model according to an exemplary embodiment;

[0057] Figure 9 3 is a schematic diagram of the composition of a model system for predicting osimertinib resistance in lung adenocarcinoma based on machine learning according to an exemplary embodiment. DETAILED DESCRIPTION

[0058] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0059] Example 1

[0060] See also Figure 1 , Figure 1 FIG1 is a flow chart of a model construction method for predicting osimertinib resistance in lung adenocarcinoma based on machine learning according to an exemplary embodiment, the method comprising:

[0061] S1. Collect clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases;

[0062] S2. performing data preprocessing on the clinical data collected from the lung adenocarcinoma patients, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm;

[0063] S3. Using the results of screening for indicators significantly associated with drug resistance, a basic prediction model based on clinical indicators was constructed. Furthermore, a prognostic model based on gene expression signatures was constructed using pre-defined univariate Cox regression analysis to screen for statistically significant gene data.

[0064] S4. Utilizing gene signatures selected from a pre-defined prognostic model based on the gene expression signature, combined with the basic prediction model based on clinical indicators, and optimizing the model performance using a pre-defined residual correction and stacking ensemble method; the gene signatures selected from the pre-defined prognostic model based on the gene expression signature include: TYMS and UAP1L1 gene expression levels;

[0065] S5. Use the optimized model to predict the patient's risk of osimertinib resistance and output the predicted probability value.

[0066] In one embodiment, as described in step S1, a comprehensive search was conducted on clinical data from a hospital for patients diagnosed with lung adenocarcinoma and whose medical records clearly mentioned "osimertinib" related information. After strict screening and rigorous identification, a total of 245 patient information was retrieved. In order to further explore the drug resistance of osimertinib in the treatment of lung adenocarcinoma, 90 patients who met the drug resistance criteria were selected as research subjects in the drug-resistant group based on established scientific and rigorous inclusion and exclusion criteria. Given that the research design needs to ensure the balance and scientificity of the comparison between the two groups, 90 cases were selected from the remaining eligible non-resistant patients as a sensitive group to ensure consistency in sample size between the two groups, thereby laying a solid foundation for subsequent in-depth and scientifically valuable comparative analysis. Therefore, a total of 180 patients who were pathologically confirmed to have lung adenocarcinoma by surgical resection or lung puncture biopsy in a hospital from March 2020 to December 2024, and whose lung cancer driver gene test results were EGFR positive, were included. After numbering, the following information was collected through the hospital information management system (HIS): gender, age, EGFR mutation site, blood routine, tumor markers such as carcinoembryonic antigen (CEA), cancer antigen 125 (Cancer Antigen 125), cytokeratin 19 fragment (CYFRA21-1) and neuron-specific enolase (NSE) serum values, tumor stage, osimertinib medication time, etc. At the same time, routine blood count and its derived indicators were calculated and recorded, such as the neutrophil-to-lymphocyte ratio (NLR), platelet-to-lymphocyte ratio (PLR), lymphocyte-to-monocyte ratio (LMR), systemic inflammatory response index (SIRI): absolute neutrophil count × absolute monocyte count / absolute lymphocyte count, and systemic immune response index (SII): absolute neutrophil count × absolute platelet count / absolute lymphocyte count.Data entry nodes: (1) Baseline data recording: Before patients start taking osimertinib, baseline clinical information, including imaging, laboratory tests, and related molecular markers, is collected; (2) Efficacy confirmation: After patients take osimertinib, relevant data are recorded when they achieve objective response (PR) or stable disease (SD) (according to RECIST 1.1 standards) during follow-up assessment (follow-up time points are based on clinical protocols, such as 6 weeks or 12 weeks); (3) Drug resistance confirmation: When patients are clearly resistant based on imaging assessment (PD) or genetic testing (such as disappearance of EGFR T790M or appearance of C797S mutation), relevant data are recorded.

[0067] Inclusion criteria: (1) All patients were diagnosed with lung adenocarcinoma by pathological examination, had only EGFR mutations, and received first-line or second-line osimertinib treatment. (2) The patients had signed the "Informed Consent Form for Sample Collection" and agreed that the remaining samples after testing would be used for medical research.

[0068] Exclusion criteria: (1) patients who are participating in other clinical studies; (2) patients with incomplete data.

[0069] Grouping criteria: patients are divided into resistant group and sensitive group according to whether they are resistant to osimertinib. The basis for judging drug resistance is:

[0070] Worsening of symptoms: such as worsening cough, chest tightness, and shortness of breath, and even new onset of distant metastasis symptoms such as bone pain and headache.

[0071] Imaging progression: CT, MRI, or PET-CT reveals tumor enlargement, new lesions, or progression of existing lesions.

[0072] Gene mutations, gene amplification, and fusions detected through tumor tissue biopsy or liquid biopsy (ctDNA): such as EGFR-related mutations (mutations such as C797S, L718Q, and L792F, EGFR amplification), bypass activation (MET gene amplification, HER2 amplification, PI 3KCA mutation, BRAF mutation, and KRAS mutation), histological transformation (transformation to small cell lung cancer and transformation to squamous cell carcinoma), other drug resistance mechanisms (EMT (epithelial-mesenchymal transition)), and new gene fusions (such as RET, ALK, and NTRK). Item 1, as a basis for suspected diagnosis, requires evidence from items 2 and 3 before a judgment can be made.

[0073] Furthermore, for public data: three groups of HCC827 cell mRNA data before and after osimertinib resistance were downloaded from the GEO database, namely GSE223006 (platform GPL24676), GSE243565 (platform GPL16791), and GSE249721 (platform GPL21697); clinical data and mRNA data of 442 patients with confirmed lung adenocarcinoma were downloaded from GSE72094 (platform GPL15048); mRNA data of biopsy tissues of four paired confirmed lung adenocarcinoma patients before and after osimertinib resistance were downloaded from GSE253742 (platform GPL24676); data including mRNA data, phenotypic characteristics, clinical data, etc. were downloaded from the LUAD cohort of the TCGA database, and preliminarily standardized preprocessed data were downloaded from UCSC Xena (https: / / xena.ucsc.edu / ), a comprehensive cancer genomics data analysis platform developed by the University of California, Santa Cruz (UCSC).

[0074] In specific implementation, as described in steps S2-S3, exploratory data analysis (EDA) was first performed using the clinical data of patients in the resistant and sensitive groups as features to understand the data distribution, preliminarily identify data patterns and outliers, and prepare for subsequent data cleaning. A Mann-Whitney U test was performed, and those with significant differences between the two groups (P < 0.05) were included as model features. The model was constructed using the Python software libraries Pandas, Numpy, matplotlib, Scikit-Learn, and Xgboost. Scikit-Learn provides simple and efficient tools for data mining and analysis, offering a rich set of algorithms and functions that support a variety of machine learning tasks, including classification, regression, clustering, and dimensionality reduction, and is suitable for both supervised and unsupervised learning scenarios. Xgboost is an efficient gradient boosting algorithm. Its core concept is to combine many "weak learners" (usually simple decision trees) into a powerful "strong learner," thereby improving the model's predictive ability. It has the advantages of high efficiency, high accuracy, and flexibility. This study used the Xgboost algorithm to process and model data to predict drug resistance in patients, thereby providing better support for precision medicine. After completing model development and preliminary evaluation, this study further introduced new features to achieve this goal by establishing a residual correction model and model stacking. The residual correction model aims to improve overall prediction accuracy by learning and correcting the errors (residuals) of the initial model's predictions. Specifically, a base model is first trained to make preliminary predictions of the target. Then, a secondary model is trained based on the residuals of this model to capture information that the base model failed to learn. The final prediction is the sum of the predictions of the base model and the residuals of the secondary model. Model stacking is an ensemble learning method that combines the predictions of multiple different base learners to construct a secondary learner (meta-model) to improve the model's generalization and predictive performance. Unlike other ensemble methods (such as bagging and boosting), stacking allows for heterogeneity in the base learners, meaning that different types of models can be used. In this study, to comprehensively evaluate the model's performance and ensure its generalization ability on unseen data, the dataset was divided into training and test sets (in a ratio of 7:3). K-fold cross-validation (K=5) was used, dividing the dataset into K subsets. K-1 subsets were used for training each time, and the remaining subset was used for validation. This was repeated K times, and the final performance indicators were averaged.After the model is built, evaluation metrics include accuracy, precision, recall, F1 score, ROC curve (Receiver Operating Characteristic curve), and AUC value. The optimal model is ultimately selected by comparing multiple models.

[0075] Furthermore, read count data for each group obtained from the database were first normalized. FPKM (Fragments Per Kilobase of exon per Million reads mapped) values ​​were calculated consistently based on gene length and sequencing depth, and log2 transformation was performed. Genes with low or no expression in both drug-resistant and sensitive genes were filtered out. To unify gene IDs across data sets, the R package biomaRt was used to uniformly identify gene tags for each group as Ensembl IDs. The data were debatch processed using the pyComBat method within the Bluk module of the Python library Omicverse. Differentially expressed genes (DEGs) were analyzed using the built-in pyDEG method within the library. All parameters were maintained at the default values ​​given in the official documentation (https: / / omicverse.readthedocs.io / en / latest / ). The differential analysis compared groups before and after osimertinib resistance, and the screening threshold for the results was set at an absolute logarithm of the fold change > 1.5 and a P < 0.05. After obtaining the DEGs, RRA (RobustRank Aggregation) analysis was performed using the R software package RobustRankAggreg. Genes with scores less than 0.005 were screened, and the results were finally presented in the form of a heat map.

[0076] In the specific implementation, lung cancer patient data from the TCGA and GEO databases were used as training data, and the R software package Mime was used to establish a prognostic model. Mime integrates many practical algorithms from data standardization, feature screening, model training, etc. according to the establishment scenarios of common clinical machine learning prognostic models, such as the random forest (Random Forest), elastic network (Elastic Net), stepwise Cox regression (Stepwise Cox), CoxBoost, partial least squares regression Cox model (Partial Least Squares Regression for Cox), generalized boosted regression model (Generalized Boosted Regression Models), survival support vector machine (Survival Support Vector Machine), Ridge regression, Lasso regression and other algorithms required for use in this article. K-fold cross-validation was performed on a combination of multiple algorithms to select the optimal model (including 10 algorithms, and a total of 101 models can be established after combination). First, we screened statistically significant genes by performing univariate COX regression analysis on the input genes. Then, we input the screened genes into various models for training. At the same time, we used the performance verification methods provided by this integration package to verify the model performance, such as C-index, AUC, etc., to select the optimal model and use the features used in the optimal model as the target genes for the next step of research.

[0077] As described in steps S4-S5, in one embodiment:

[0078] This study included 180 patients, 90 in the drug-resistant group and 90 in the sensitive group. The clinical characteristics of the two groups are shown in Table 1 (Comparison of clinical characteristics of sensitive and resistant patients). There was no statistically significant difference in clinical characteristics between the two groups (P>0.05).

[0079] Table 1

[0080]

[0081] The patient test results showed that the following values ​​were associated with the drug resistance status of the patients: neutrophil percentage, lymphocyte percentage, absolute lymphocyte count, CEA, CA125, CA153, NLR, PLR, LMR, SIRI, and SII, with an FDR of less than 0.01 by the Mann-Whitney U test. Other comparison results are also presented in Table 2 (Comparison of test results between sensitive and resistant patients).

[0082] Table 2

[0083]

[0084]

[0085] In the specific implementation, the indicators that were statistically significant in the previous U test were considered to be effective features for identifying drug resistance and were included as model construction indicators. The established model had an accuracy of 0.65 and a recall rate of 0.64 in distinguishing osimertinib-sensitive and -resistant patients. The area under the cross-validation receiver operating characteristic curve (AUC) was 0.716±0.051, and the validation set AUC was 0.64 (see Figure 2 . (Receiver Operating Characteristic curve).

[0086] Furthermore, differential analysis was performed on three groups of drug-resistant cell line data from different laboratories. The DEGs confirmed by the dataset GSE223006 included 253 up-regulated genes and 199 down-regulated genes ( Figure 3 A~C); DEGs confirmed by dataset GSE243565 include 1205 up-regulated genes and 1278 down-regulated genes; DEGs confirmed by dataset GSE249721 include 2162 up-regulated genes and 2089 down-regulated genes. Specifically, the volcano plot of differentially expressed genes in each GEO database is shown. A total of 1001 genes were found to be commonly expressed DEGs in the three groups of cell lines before and after drug resistance through RRA analysis, of which 473 up-regulated genes and 528 down-regulated genes were used as target genes. The top 20 genes ranked by RRA analysis are shown in the figure ( Figure 3 D).

[0087] Specifically, Figure 3 Figure 3 shows the common DEGs associated with osimertinib resistance identified from GEO data. (A) to (C) are volcano plots of differentially expressed genes from multiple data sets in the GEO database. (D) The top 20 genes with the highest common expression scores in the three groups of differentially expressed gene data were obtained by RRA analysis.

[0088] Furthermore, in order to find indicators that can predict drug resistance, co-expressed DEGs were used as features, and the prognosis of patients was used as the outcome to construct a prognostic model. In two large lung adenocarcinoma data sets (TCG A lung adenocarcinoma cohort 542 cases, GEO lung adenocarcinoma cohort 372 cases), the increasingly popular and scalable algorithm architecture was used to complete the feature screening of the single-factor Cox model. Then, 101 models were constructed using the screened genes. The best model was StepCox[forward]+GBM, which had a C-index of 0.72 in the TCGA training data set, 0.64 in the GEO validation data set, an average C-index of 0.68 in each cohort, and an average C-index of 0.64 in the validation set. The remaining model parameters are shown in Figure 4 ,The model finally selected features (78 genes in total) as target genes.

[0089] Based on the risk score calculated by the optimal model, patients were divided into high-risk and low-risk groups. Survival curves were drawn with survival time as the outcome and the log-rank test was performed. In the TCGA dataset, the hazard ratio (HR) was 13.52, with a 95% confidence interval of 10.02-18.23; in the GEO dataset, the HR was 2.21, with a 95% confidence interval of 1.49-3.28. The survival analysis of the two groups was statistically significant in both datasets (P < 0.05). The results are shown in Figure 5. Figure 5 .in, Figure 5 Represents the relationship between the risk score calculated by the optimal combination model StepCox[forward]+RSF and the outcomes of patients in different cohorts.

[0090] See also Figure 6 Differential gene expression analysis was performed on the high-risk group within the risk group, and enrichment analysis of the DEGs obtained from this analysis was performed. The results showed that the high-risk group in the prognostic model for drug resistance-related genes was significantly enriched in biological processes such as the cell cycle, mitotic cell cycle, and chromosome segregation; cellular component aggregation areas such as chromatin and cell membranes; and molecular functions such as nucleotide binding and drug binding. Furthermore, KEGG enrichment analysis suggested that the emergence of drug resistance is mainly concentrated in the cell cycle, cell senescence, p53 signaling pathway, FoxO signaling pathway, pyrimidine metabolism, and drug metabolism.

[0091] against Figure 6 , (A) The top ten results with statistical significance (P<0.05) after GO analysis; (B) The top five results with statistical significance (P<0.05) after KEGG pathway enrichment analysis.

[0092] Furthermore, the features in the previous prognostic model were used as target genes to verify whether these genes have the function of drug resistance and prognosis prediction in patient tissues; the tissue dataset before and after osimertinib resistance downloaded from the GEO database was used to verify the relevant genes, among which TYMS, UAP1 L1, OLFM2A, and CL EC7A showed differences in the database, and the comparison results of the remaining genes were presented in the form of heat maps. Figure 7 A. Serum RNA detection was performed on 73 serum samples from our hospital to identify genes with significant differences in the database. The results showed that TYMS and UAP1 L1 were elevated in the drug-resistant group (P<0.05), while the expression of the remaining genes in serum and body fluids showed no significant differences ( Figure 7 BC).

[0093] Figure 7 Indicates target gene expression verification; (A) Cluster analysis heat map of target genes in GEO dataset; (B) (C) Serum RNA detection verification of TYMS and UAP1 L1 gene expression in sensitive and resistant groups.

[0094] After identifying the target gene, measurements were completed on the remaining samples. New features were added to the original drug resistance prediction model to enhance model performance. The residual error correction model achieved an AUC of 0.924, while the stacked ensemble model achieved an AUC of 0.873. The receiver operating characteristic (ROC) curves for each model are shown in the figure. Figure 8 Represents the ROC curves of the basic model, residual correction model, and stacked model.

[0095] The examples in this application retrospectively collect clinical data of 180 patients with EGFR mutation-positive lung adenocarcinoma from a certain hospital between March 2020 and December 2024, and group them according to drug resistance status. Feature screening and model construction are performed using machine learning algorithms such as XGBoost. The preliminary model uses blood routine, biochemical tests and tumor marker results as input features, and then combines public databases (GEO, TCGA) for differential gene analysis to screen resistance-related genes. The prognosis model is established in combination with clinical data and features with high clinical value are screened out as target genes. They are further verified in public databases to see if they are related to drug resistance, and serum RNA detection is performed on drug-resistant samples to clarify the target gene expression level. Residual correction, stacking and other methods are used to optimize and update the target gene for the basic model, and the model prediction ability is finally evaluated by cross-validation, AUC, accuracy and other indicators.

[0096] In practice, the initially constructed prediction model performed poorly, with an area under the receiver operating characteristic curve (AUC) of 0.64 and a recall rate of 0.64, suggesting that the model's predictive performance needs improvement. Through bioinformatics analysis, we identified 1001 differentially expressed genes, which were primarily enriched in cell cycle regulation, drug metabolism, and pyrimidine metabolism pathways. Using machine learning methods, we screened 73 prognostic-related genes from these differentially expressed genes and constructed a prognostic prediction model. Results from the GEO dataset and serum RNA analysis showed that thymidylate synthase (TYMS) and UDP-N-acetylglucosamine pyrophosphorylase 1-like 1 (UAP1 L1) were significantly overexpressed in osimertinib-resistant patients, suggesting their potential role in resistance mechanisms. The AUC of the residual correction model after model optimization using TYMS and UAP1 L1 was improved to 0.873, and the AUC of the stacked ensemble model reached 0.924.

[0097] Furthermore, this application constructs a machine learning-based prediction model for osimertinib resistance and screens for potential resistance markers (TYMS, UAP1 L1), providing new insights for non-invasive resistance monitoring. In the future, the model can be optimized by combining multicenter data to enhance its clinical application value.

[0098] See also Figure 9 , Figure 9 FIG1 is a schematic diagram showing the composition of a model system for predicting osimertinib resistance in lung adenocarcinoma based on machine learning according to an exemplary embodiment. The system includes:

[0099] A data acquisition module 10 is used to collect clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases;

[0100] a data preprocessing module 20 for performing data preprocessing on the clinical data collected from lung adenocarcinoma patients, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm;

[0101] Model construction module 30, for constructing a basic prediction model based on clinical indicators using the results of screening indicators significantly associated with drug resistance, and for constructing a prognostic model based on gene expression characteristics using the preset univariate COX regression analysis to screen statistically significant gene data;

[0102] A model fusion module 40 is configured to utilize gene features selected from a preset prognostic model based on the gene expression signature, combine the gene features with the basic prediction model based on clinical indicators, and optimize the model performance through a preset residual correction and stacking integration method; the gene features selected from the preset prognostic model based on the gene expression signature include: TYMS and UAP1 L1 gene expression levels;

[0103] The prediction analysis module 50 is used to use the optimized model to predict the patient's osimertinib resistance risk and output a predicted probability value.

[0104] More specifically, data acquisition module 10: By integrating clinical data of lung adenocarcinoma patients (such as blood routine tests and tumor markers) and genetic data from public databases (such as mRNA before and after osimertinib resistance), covering phenotypic and molecular level characteristics, it provides a multi-dimensional data source for the model to ensure the comprehensiveness and accuracy of the prediction.

[0105] Data preprocessing module 20: Through statistical tests (such as Mann-Whitney U test) and algorithm screening (such as pyDEG), redundant information is eliminated, key features (such as NLR and TYMS genes) are retained, data quality is improved, model training complexity is reduced, and the reliability of subsequent analysis is ensured.

[0106] Model building module 30:

[0107] Basic prediction model: Rapidly build a preliminary screening model based on routine clinical indicators to achieve non-invasive preliminary assessment of drug resistance risk, suitable for primary medical scenarios.

[0108] Prognostic model: Through gene expression data, we can mine the deep mechanisms of drug resistance (such as cell cycle regulatory pathways) and screen key biomarkers (such as TYMS and UAP1L1) to provide direction for mechanism research and targeted therapy.

[0109] Model fusion module 40: Combines genetic features (TYMS, UAP1 L1) with clinical models, improves prediction accuracy through residual correction and stacking integration (AUC increased from 0.64 to 0.924), achieves early and accurate warning of drug resistance risks, and reduces patient trauma and costs through non-invasive serum testing.

[0110] Prediction and analysis module 50: Outputs quantitative prediction results (such as the probability value of drug resistance) to assist doctors in dynamically adjusting treatment strategies (such as combination therapy), promote the implementation of personalized medicine, and improve patients' quality of life and treatment efficiency.

[0111] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0112] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.

[0113] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0114] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0115] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0116] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0117] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0118] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0119] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A model construction method for predicting osimertinib resistance in lung adenocarcinoma based on machine learning, characterized in that: The method comprises: Clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases were collected; performing data preprocessing on the clinical data collected from lung adenocarcinoma patients, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm; Using the results of screening for indicators significantly associated with drug resistance, a basic prediction model based on clinical indicators was constructed. Using the preset univariate COX regression analysis to screen for statistically significant gene data, a prognostic model based on gene expression characteristics was constructed. The gene features screened in the preset prognostic model based on the gene expression features are combined with the basic prediction model based on clinical indicators to optimize the model performance through a preset residual correction and stacking integration method; the gene features screened in the preset prognostic model based on the gene expression features include: TYMS and UAP1L1 gene expression levels; The optimized model is used to predict the patient's risk of osimertinib resistance and output the predicted probability value.

2. The method according to claim 1, characterized in that The clinical data of the lung adenocarcinoma patient include gender, age, EGFR mutation site, blood test results, tumor markers, tumor stage, and osimertinib medication time information.

3. The method according to claim 1, characterized in that The osimertinib resistance-related mRNA data in the public database include mRNA data of cell lines before and after osimertinib resistance and mRNA data of lung adenocarcinoma patients.

4. The method according to claim 1, wherein The data preprocessing of the collected clinical data of lung adenocarcinoma patients, screening of indicators significantly associated with drug resistance, and screening of differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm include: The clinical data collected from the lung adenocarcinoma patients were used to screen for indicators significantly associated with drug resistance using the Mann-Whitney U test; The osimertinib resistance-related mRNA data in the public database were sequentially subjected to FPKM normalization, log2 conversion, and pyComBat de-batch processing, and differentially expressed genes were screened using the pyDEG algorithm.

5. The method according to claim 1, wherein The method comprises the following steps: using the results of screening for indicators significantly associated with drug resistance to construct a basic prediction model based on clinical indicators, and using the preset univariate COX regression analysis to screen for statistically significant gene data to construct a prognostic model based on gene expression characteristics, including: Using the results of screening indicators significantly related to drug resistance, the data was processed and modeled using the Xgboost algorithm; The pre-set univariate COX regression analysis was used to screen the statistically significant gene data, and the prognostic model was established using the R software package Mime.

6. The method according to claim 1, characterized in that The gene features screened in the preset prognostic model based on the gene expression features are combined with the basic prediction model based on clinical indicators, and the model performance is optimized through a preset residual correction and stacking integration method, including: Using gene signatures screened in a preset prognostic model based on the gene expression signature; the gene signatures include: TYMS and UAP1 L1; The RNA expression levels of TYMS and UAP1 L1 were examined using pre-specified serum samples; Based on the detection results, the preset TYMS and UAP1 L1 gene expression levels were used as new features and merged with the residual data to train the secondary correction model to obtain the first processing results; Utilizing the first processing result, the outputs of multiple base models are used as input, and the final prediction result is generated through meta-learning integration to optimize its model performance.

7. The method according to claim 6, characterized in that The preset method for measuring the expression levels of the TYMS and UAP1 L1 genes includes: Total RNA was extracted using serum or plasma RNA extraction kit; cDNA was synthesized by reverse transcription, and gene expression levels were detected using SYBR Green fluorescence quantitative PCR; GAPDH was used as the internal reference gene, and the relative expression levels were calculated using the ΔΔCt method.

8. The method according to claim 1, characterized in that The construction of the basic prediction model based on clinical indicators also includes: 5-fold cross validation was used to optimize the Xgboost model parameters, including learning rate, number of trees, and maximum depth; Key clinical indicators, including NLR, lymphocyte percentage, and CEA, were determined by feature importance analysis.

9. The method according to claim 1, characterized in that The prognostic model of gene expression signature further includes: A prognostic model was constructed using various algorithms in the Mime package; Evaluate model performance using C-index and AUC to select the optimal algorithm combination; The various algorithms in the Mime package include: random forest, elastic network, stepwise Cox regression, CoxBoost, partial least squares regression Cox, generalized boosted regression model, survival support vector machine, Ridge regression, and Lasso regression algorithm.

10. A machine learning-based model construction system for predicting osimertinib resistance in lung adenocarcinoma, applied to the machine learning-based model construction method for predicting osimertinib resistance in lung adenocarcinoma according to any one of claims 1 to 9, characterized in that: The method comprises: A data acquisition module is used to collect clinical data of lung adenocarcinoma patients and osimertinib resistance-related mRNA data from public databases; a data preprocessing module for performing data preprocessing on the clinical data collected from patients with lung adenocarcinoma, screening for indicators significantly associated with drug resistance, and screening for differentially expressed genes in the osimertinib resistance-related mRNA data in the public database using a preset algorithm; The model building module is used to use the results of screening indicators significantly associated with drug resistance to build a basic prediction model based on clinical indicators, and to use the preset univariate COX regression analysis to screen statistically significant gene data to build a prognostic model based on gene expression characteristics; a model fusion module for utilizing gene features selected from a preset prognostic model based on the gene expression signature, combining the gene features with the basic prediction model based on clinical indicators, and optimizing the model performance through a preset residual correction and stacking integration method; the gene features selected from the preset prognostic model based on the gene expression signature include: TYMS and UAP1 L1 gene expression levels; The predictive analysis module is used to predict the patient's risk of osimertinib resistance using the optimized model and output the predicted probability value.

Citation Information

Patent Citations

  • Primer group, kit and method for tumor chemotherapeutic drug-related gene TYMS expression level detection

    CN107760786A

  • Kit for predicting lung adenocarcinoma radiotherapy prognosis model and application thereof

    CN115036026A

  • Artificial intelligence system for noninvasive prediction of EGFR / TP53 co-mutation lung cancer patient

    CN115312126A

  • Biomarker for evaluating prognosis, immunization or treatment effect of breast cancer and prediction model based on biomarker

    CN115651983A

  • Overall survival rate prognosis model for lung squamous cell carcinoma patient and application

    CN116153387A

Cited By

  • Primary drug resistance prediction system and method

    CN121366733A

  • Comprehensive prediction model for predicting treatment drug resistance of hepatocellular carcinoma patient as well as construction method and application of comprehensive prediction model

    CN121545777A