A method and system for screening key target points of lung adenocarcinoma based on machine learning

By combining machine learning methods with multi-dimensional validation, high-confidence key targets for lung adenocarcinoma were screened, which solved the problems of insufficient specificity and lack of multi-dimensional validation in existing technologies. This improved the accuracy and reliability of lung adenocarcinoma target screening and allowed for direct assessment of its clinical translation potential.

CN122493936APending Publication Date: 2026-07-31JIAMUSI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIAMUSI UNIVERSITY
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for screening lung adenocarcinoma targets suffer from insufficient specificity and lack multi-dimensional validation, resulting in unclear clinical translational value of the screened targets. Furthermore, these methods are susceptible to the influence of single algorithmic biases, and their stability and reliability need to be improved.

Method used

Using machine learning methods, through differential expression analysis, Cox proportional hazards regression, functional enrichment analysis, screening of multiple machine learning models and multi-dimensional validation, combined with drug database queries, a multi-indicator evaluation system was constructed to screen out key targets for lung adenocarcinoma with high confidence.

Benefits of technology

This improved the accuracy and reliability of screening key targets for lung adenocarcinoma, ensuring that the screening results have clear clinical significance and multi-dimensional validation, thus enhancing the reliability of the results and the targeted nature of drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493936A_ABST
    Figure CN122493936A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for screening key targets in lung adenocarcinoma based on machine learning, relating to the field of intelligent data screening, including the following steps: Step 1: Obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from a public genomics database, perform batch effect correction and standardization processing, and construct a standardized expression matrix; Step 2: Based on the standardized expression matrix, use differential expression analysis to screen differentially expressed genes in lung adenocarcinoma; By constructing a preliminary screening mechanism that associates differential expression with prognosis, a large number of genes that are differentially expressed but not related to patient survival are first excluded, ensuring that the candidate gene set has a clear clinical significance starting point. Different machine learning algorithms are integrated for cross-screening, and a high-confidence core gene set is obtained by taking the intersection, reducing the false positive rate and making the final output target statistically significant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent data screening technology, specifically to a method and system for screening key targets of lung adenocarcinoma based on machine learning. Background Technology

[0002] Lung cancer is one of the leading causes of cancer-related morbidity and mortality worldwide. Lung adenocarcinoma, a major subtype of non-small cell lung cancer, exhibits a complex and highly heterogeneous pathogenesis. With the rapid development of high-throughput sequencing technology, public databases, such as cancer genome atlases and gene expression synthesis databases, have accumulated massive amounts of multi-omics data and corresponding clinical information on lung adenocarcinoma, providing unprecedented resources for systematically exploring the nature of the disease at the molecular level. Simultaneously, the in-depth application of artificial intelligence and machine learning technologies in bioinformatics has made it possible to extract clinically significant patterns from complex biological big data.

[0003] Current technologies for screening targets for lung adenocarcinoma have limitations. Traditional methods are usually based solely on differential expression analysis, resulting in a large number of genes screened, many of which are noise genes unrelated to disease progression, leading to insufficient specificity. Many studies focus only on gene expression differences, lacking direct and systematic validation of their association with patient clinical prognosis, resulting in unclear clinical translational value of the screened targets. Existing methods often employ single algorithm models, whose results are easily affected by specific algorithm preferences or parameter settings, and their stability and reliability need improvement. Most screening processes stop at bioinformatics prediction, lacking systematic validation and integrated evaluation across multiple dimensions such as independent cohorts, protein levels, and the tumor immune microenvironment, making it difficult to comprehensively evaluate the potential application value of targets, thus affecting the efficiency of subsequent experimental validation and the targeting of drug development. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a method and system for screening key targets of lung adenocarcinoma based on machine learning, which can effectively solve the problems of the existing technology.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] This invention discloses a method for screening key targets in lung adenocarcinoma based on machine learning, comprising the following steps:

[0009] Step 1: Obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, perform batch effect correction and standardization, and construct a standardized expression matrix;

[0010] Step 2: Based on the standardized expression matrix, differential expression analysis was used to screen differentially expressed genes in lung adenocarcinoma. Univariate Cox proportional hazards regression analysis was performed on the differentially expressed genes to screen out a subset of prognostic differentially expressed genes that were significantly associated with overall patient survival.

[0011] Step 3: Perform functional enrichment analysis on the subset of differentially expressed genes related to prognosis to clarify the biological processes and signaling pathways they are involved in;

[0012] Step 4: Using the expression levels of the differentially expressed gene subset related to prognosis as features, and patient survival status and time as labels, construct a machine learning model based on LASSO regression, support vector machine recursive feature elimination, and random forest. Perform key gene screening for each model, extract the gene lists from each model, and take their intersection to obtain a core set of key genes with a confidence level higher than the preset threshold.

[0013] Step 5: Perform survival analysis on independent datasets to validate the prognostic value of the core set of key genes; validate their expression specificity in protein expression level databases; analyze the evidence of immune association between their expression and the level of tumor immune microenvironment infiltration.

[0014] Step 6: Integrate the prognostic value, expression specificity, and immune association evidence obtained in Step 5, and combine it with drug database queries to determine its drug potential. Construct a multi-index evaluation system, comprehensively score and rank the genes in the core set of key genes, and output a list of potential key targets for lung adenocarcinoma with the highest priority.

[0015] Furthermore, the public genome database in step 1 includes the TCGA database and the GTEx database, and the preprocessing includes batch correction using the ComBat algorithm and logarithmic transformation and normalization of the expression data.

[0016] Furthermore, in step 2, the limma R package is used for differential expression analysis, and the screening threshold is an adjusted p-value less than 0.05 and an absolute value of a logarithmic 2-fold change greater than 1; the significance threshold for the univariate Cox regression analysis is a p-value less than 0.05.

[0017] Furthermore, in step 3, the clusterProfiler R package is used to perform GO function and KEGG pathway enrichment analysis, with a screening threshold of a false detection rate of less than 0.05.

[0018] Furthermore, in step 4, the LASSO regression uses 10-fold cross-validation to determine the optimal penalty coefficient λ and filters out genes with non-zero coefficients; the support vector machine recursive feature elimination method iteratively removes the features with the smallest weights until a preset number of features is reached; the random forest algorithm evaluates feature importance by calculating the decrease in average precision or the decrease in Gini coefficient, and selects the top N genes in terms of importance.

[0019] Furthermore, the data source for the independent dataset in step 5 is the GEO database; the protein expression level verification is performed using the Human Protein Atlas database; and the immune microenvironment analysis uses the TIMER algorithm or CIBERSORT algorithm to calculate the correlation between key gene expression and immune cell infiltration fraction.

[0020] Furthermore, the drug database in step 6 is the DrugBank database, and the drug potential query includes checking whether the key gene is a known drug target or interacts with a known drug target.

[0021] Furthermore, the operational logic of the multi-index evaluation system in step 6 is as follows:

[0022] For each gene in the core set of key genes, its index values ​​in four dimensions—prognostic risk ratio, expression fold change, immune correlation coefficient, and level of evidence for drug development—are quantified and normalized. Preset weights are assigned to the four normalized indicators, and a comprehensive score for each gene is calculated by weighted summation. All genes are sorted according to the comprehensive scores, and a final priority list of potential key targets for lung adenocarcinoma is output based on preset comprehensive score thresholds and core indicator thresholds.

[0023] A machine learning-based system for screening key targets in lung adenocarcinoma includes:

[0024] 01 is used to obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, and perform batch effect correction and standardization to output a standardized expression matrix;

[0025] 02, used to perform differential expression analysis based on the standardized expression matrix to screen differentially expressed genes, and further perform univariate Cox proportional hazards regression analysis on the differentially expressed genes to output a subset of prognostic-related differentially expressed genes;

[0026] 03, used to perform gene ontology function and Kyoto Encyclopedia of Genes and Genomes enrichment analysis on the subset of differentially expressed genes related to prognosis, and output the biological processes, molecular functions and signaling pathway annotation information of the enriched genes;

[0027] 04, used to use the expression level of the differentially expressed gene subset related to prognosis as a feature and patient survival data as a label, to execute three algorithms in parallel: LASSO regression, support vector machine recursive feature elimination, and random forest, and to take the intersection of the important gene lists output by each algorithm to output a core set of key genes with high confidence.

[0028] 05, used to perform independent survival verification, protein expression level verification, and immune infiltration correlation analysis on the core set of key genes, and output multidimensional verification results including prognostic robustness, expression specificity, and immune correlation;

[0029] 06 is used to integrate the multidimensional validation results and biological function annotation information, and query the drug database to obtain drug-likeness information. Based on the preset multi-index evaluation system, the genes in the core set of key genes are comprehensively scored and ranked, and a list of potential key targets after priority ranking is output.

[0030] Furthermore, 01 is communicatively connected to 02, 02 is communicatively connected to 03 and 04, 04 is communicatively connected to 05, and 06 is communicatively connected to 03 and 05.

[0031] (III) Beneficial Effects

[0032] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects:

[0033] 1. By constructing a preliminary screening mechanism that links differential expression with prognosis, a large number of genes that are differentially expressed but not related to patient survival are excluded, ensuring that the candidate gene set has a clear clinical significance starting point. Different machine learning algorithms are integrated for cross-screening, and a high-confidence core gene set is obtained by taking the intersection. The screening strategy that combines tandem and parallel approaches reduces the false positive rate and makes the final output target statistically significant.

[0034] 2. By designing a multi-dimensional validation mechanism that includes independent cohort survival validation, protein expression validation, immune infiltration association analysis, and pathway phenotype validation, each key gene is subjected to testing from different data levels and technical perspectives, thereby comprehensively characterizing its features as a potential target and greatly enhancing the reliability of the discovery results. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0036] Figure 1 This is a flowchart illustrating the method for screening key targets of lung adenocarcinoma in this invention.

[0037] Figure 2 This is a schematic diagram of the framework of the lung adenocarcinoma key target screening system in this invention.

[0038] The numbers in the diagram represent: 1. Data acquisition module; 2. Preliminary screening module; 3. Enrichment analysis module; 4. Machine learning screening module; 5. Validation module; 6. Target ranking module. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0040] The present invention will be further described below with reference to embodiments.

[0041] This embodiment presents a method for screening key targets in lung adenocarcinoma based on machine learning, such as... Figure 1 As shown, it includes the following steps:

[0042] Step 1: Obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, perform batch effect correction and standardization, and construct a standardized expression matrix; public genome databases include TCGA database and GTEx database, the preprocessing includes batch correction using the ComBat algorithm, and logarithmic transformation and normalization of expression data.

[0043] Step 2: Based on the standardized expression matrix, differential expression analysis was used to screen differentially expressed genes in lung adenocarcinoma. Univariate Cox proportional hazards regression analysis was performed on the differentially expressed genes to screen out a subset of prognostic differentially expressed genes that were significantly associated with overall patient survival. The limma R package was used for differential expression analysis, and the screening threshold was an adjusted p-value of less than 0.05 and an absolute value of a logarithmic 2-fold change greater than 1. The significance threshold for the univariate Cox regression analysis was a p-value of less than 0.05.

[0044] Step 3: Perform functional enrichment analysis on the subset of differentially expressed genes related to prognosis to identify the biological processes and signaling pathways they are involved in; use the clusterProfiler R package to perform GO function and KEGG pathway enrichment analysis, with a screening threshold of false discovery rate less than 0.05.

[0045] Step 4: Using the expression levels of the differentially expressed gene subset related to prognosis as features, and patient survival status and time as labels, construct machine learning models based on LASSO regression, support vector machine recursive feature elimination, and random forest. Perform key gene screening for each model, extract the gene lists selected by each model, and take their intersection to obtain a core set of key genes with a confidence level higher than the preset value. LASSO regression uses 10-fold cross-validation to determine the optimal penalty coefficient λ, screening out genes with a non-zero coefficient. The support vector machine recursive feature elimination method iteratively removes the features with the smallest weights until a preset number of features is reached. The random forest algorithm evaluates feature importance by calculating the decrease in average precision or the decrease in Gini coefficient, and selects the top N genes in terms of importance.

[0046] Step 5: Perform survival analysis on independent datasets to validate the prognostic value of the core set of key genes; validate their expression specificity in a protein expression level database; analyze the evidence of immune association between their expression and tumor immune microenvironment infiltration level; the data source for the independent datasets is the GEO database; the protein expression level validation is performed using the Human Protein Atlas database; the immune microenvironment analysis uses the TIMER algorithm or CIBERSORT algorithm to calculate the correlation between key gene expression and immune cell infiltration score.

[0047] Step 6: Integrate the prognostic value, expression specificity, and immune association evidence obtained in Step 5, and combine it with drug database queries to determine its drug potential. Construct a multi-index evaluation system, comprehensively score and rank the genes in the core set of key genes, and output a list of potential key targets for lung adenocarcinoma with the highest priority. The drug database is the DrugBank database. The drug potential query includes checking whether the key genes are known drug targets or interact with known drug targets.

[0048] The operating logic of the multi-indicator evaluation system is as follows:

[0049] For each gene in the core set of key genes, its index values ​​in four dimensions—prognostic hazard ratio, expression fold change, immune correlation coefficient, and druggability evidence level—were quantified and normalized. Preset weights were assigned to these four normalized indicators, and a weighted summation was used to calculate the comprehensive score for each gene. Prognostic value had the highest weight, followed by expression specificity and immune correlation, and then druggability potential, with the sum of all weights equal to 1. Prognostic value was quantified using the hazard ratio and its 95% confidence interval lower limit from a univariate Cox regression analysis. Expression specificity was quantified using the absolute value of the logarithmic 2-fold change in expression levels between cancerous and normal tissues. Immune correlation was quantified using the absolute value of the Pearson correlation coefficient calculated between gene expression and the infiltration level of at least one major immune cell. Druggability potential was graded and assigned values ​​based on drug database query results. Genes targeted by FDA-approved drugs received the highest value, those in clinical trials received the second highest value, those with clear interactions or belonging to druggable gene families received a base value, and those with no relevant information received zero.

[0050] All genes are sorted according to the comprehensive score, and a final priority list of potential key targets for lung adenocarcinoma is output based on the preset comprehensive score threshold and core indicator threshold. A minimum comprehensive score threshold and a minimum single indicator threshold are set. Only when the comprehensive score of a gene is higher than the minimum comprehensive score threshold, and at least one of its core indicators such as prognostic value or expression specificity is higher than the corresponding minimum single indicator threshold, will it be retained in the final priority target list, so as to ensure that the output targets have both comprehensive advantages and outstanding performance in key dimensions.

[0051] Compared with existing technologies, by integrating multiple machine learning algorithms for cross-validation and feature screening, the overfitting or selection bias that may exist in a single model is effectively overcome, improving the accuracy and reliability of key target screening. A multi-dimensional and systematic validation system is constructed, from gene expression, clinical prognosis to protein level, immune microenvironment and drug-likeness. This not only verifies the biological significance of the targets, but also directly evaluates their clinical translation potential, making the screening results more valuable for clinical application.

[0052] At other levels, this embodiment also provides a machine learning-based screening system for key targets in lung adenocarcinoma, such as... Figure 2 As shown, it includes:

[0053] 01 is used to obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, and perform batch effect correction and standardization to output a standardized expression matrix;

[0054] 02, used to perform differential expression analysis based on the standardized expression matrix to screen differentially expressed genes, and further perform univariate Cox proportional hazards regression analysis on the differentially expressed genes to output a subset of prognostic-related differentially expressed genes;

[0055] 03, used to perform gene ontology function and Kyoto Encyclopedia of Genes and Genomes enrichment analysis on the subset of differentially expressed genes related to prognosis, and output the biological processes, molecular functions and signaling pathway annotation information of the enriched genes;

[0056] 04, used to use the expression level of the differentially expressed gene subset related to prognosis as a feature and patient survival data as a label, to execute three algorithms in parallel: LASSO regression, support vector machine recursive feature elimination, and random forest, and to take the intersection of the important gene lists output by each algorithm to output a core set of key genes with high confidence.

[0057] 05, used to perform independent survival verification, protein expression level verification, and immune infiltration correlation analysis on the core set of key genes, and output multidimensional verification results including prognostic robustness, expression specificity, and immune correlation;

[0058] 06 is used to integrate the multidimensional validation results and biological function annotation information, and query the drug database to obtain drug-likeness information. Based on the preset multi-index evaluation system, the genes in the core set of key genes are comprehensively scored and ranked, and a list of potential key targets after priority ranking is output.

[0059] 01 is connected to 02, 02 is connected to 03 and 04, 04 is connected to 05, and 06 is connected to 03 and 05.

[0060] In summary, this invention performs preliminary screening by linking differential expression with prognosis, ensuring that the target gene set possesses both tumor specificity and clinical relevance. It integrates three complementary machine learning algorithms—LASSO regression, support vector machine recursive feature elimination, and random forest—for ensemble screening, effectively overcoming the bias of single models. By employing an intersection strategy, it identifies the core gene set with high confidence, significantly enhancing the reliability of the results. Furthermore, through a multi-dimensional independent validation mechanism encompassing independent cohort survival analysis, protein expression validation, and immune microenvironment association, it cross-validates the potential value of the target from different biological perspectives.

[0061] By constructing a comprehensive evaluation system that integrates prognostic, expression, immune, and drug-likeness information for priority ranking, the screened targets not only have a solid biological basis but also directly point to the feasibility of drug development. This achieves automated, standardized, and rational screening from massive data to high-value targets, providing efficient and reliable decision support for the precision treatment of lung adenocarcinoma.

[0062] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for screening key targets in lung adenocarcinoma based on machine learning, characterized in that, Includes the following steps: Step 1: Obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, perform batch effect correction and standardization, and construct a standardized expression matrix; Step 2: Based on the standardized expression matrix, differential expression analysis was used to screen differentially expressed genes in lung adenocarcinoma. Univariate Cox proportional hazards regression analysis was performed on the differentially expressed genes to screen out a subset of prognostic differentially expressed genes that are related to the patient's overall survival. Step 3: Perform functional enrichment analysis on a subset of differentially expressed genes related to prognosis; Step 4: Using the expression levels of differentially expressed genes related to prognosis as features and patient survival status and time as labels, construct machine learning models based on LASSO regression, recursive feature elimination of support vector machines, and random forests to screen key genes respectively, extract the gene lists selected by each model, and take their intersection to obtain a core set of key genes with a higher confidence level than the preset level. Step 5: Perform survival analysis on independent datasets to validate the prognostic value of the core set of key genes; validate their expression specificity in protein expression level databases; analyze the evidence of immune association between their expression and the level of tumor immune microenvironment infiltration. Step 6: Integrate the prognostic value, expression specificity, and immune association evidence obtained in Step 5, and combine it with drug database queries to determine its drug potential. Construct a multi-index evaluation system, comprehensively score and rank the genes in the core set of key genes, and output a list of potential key targets for lung adenocarcinoma with the highest priority.

2. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, The public genome databases in step 1 include the TCGA database and the GTEx database. The preprocessing includes batch correction using the ComBat algorithm and logarithmic transformation and normalization of the expression data.

3. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, In step 2, the limma R package is used for differential expression analysis, and the screening threshold is an adjusted p-value less than 0.05 and an absolute value of a logarithmic 2-fold change greater than 1; the significance threshold for the univariate Cox regression analysis is a p-value less than 0.

05.

4. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, In step 3, the clusterProfiler R package is used to perform GO function and KEGG pathway enrichment analysis, with a screening threshold of a false detection rate of less than 0.

05.

5. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, In step 4, the LASSO regression uses 10-fold cross-validation to determine the optimal penalty coefficient λ and filters out genes with non-zero coefficients; the support vector machine recursive feature elimination method iteratively removes the features with the smallest weights until a preset number of features is reached. The random forest algorithm assesses feature importance by calculating the average precision decrease or the Gini coefficient decrease, and selects the top N genes in terms of importance.

6. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, The data source for the independent dataset in step 5 is the GEO database; the protein expression level verification is performed using the Human Protein Atlas Database; and the immune microenvironment analysis uses the TIMER algorithm or CIBERSORT algorithm to calculate the correlation between key gene expression and immune cell infiltration fraction.

7. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, The drug database in step 6 is the DrugBank database, and the drug potential query includes checking whether the key gene is a known drug target or interacts with a known drug target.

8. The method for screening key targets of lung adenocarcinoma based on machine learning according to claim 1, characterized in that, The operating logic of the multi-index evaluation system in step 6 is as follows: For each gene in the core set of key genes, its index values ​​in four dimensions—prognostic risk ratio, expression fold change, immune correlation coefficient, and level of evidence for drug development—are quantified and normalized. Preset weights are assigned to the four normalized indicators, and the comprehensive score of each gene is calculated by weighted summation. All genes are sorted according to the comprehensive score, and a final priority list of potential key targets for lung adenocarcinoma is output based on the preset comprehensive score threshold and core indicator threshold.

9. A machine learning-based system for screening key targets in lung adenocarcinoma, the system being an integrated system based on any one of claims 1-8, characterized in that... include: 01 is used to obtain gene expression profile data and corresponding clinical survival information of lung adenocarcinoma tissue and normal control tissue from public genome databases, and perform batch effect correction and standardization to output a standardized expression matrix; 02, used to perform differential expression analysis based on the standardized expression matrix to screen differentially expressed genes, and further perform univariate Cox proportional hazards regression analysis on the differentially expressed genes to output a subset of prognostic-related differentially expressed genes; 03, used to perform gene ontology function and Kyoto Encyclopedia of Genes and Genomes enrichment analysis on the subset of differentially expressed genes related to prognosis, and output the biological processes, molecular functions and signaling pathway annotation information of the enriched genes; 04, used to use the expression level of the differentially expressed gene subset related to prognosis as a feature and patient survival data as a label, to execute three algorithms in parallel: LASSO regression, support vector machine recursive feature elimination, and random forest, and to take the intersection of the important gene lists output by each algorithm to output a core set of key genes with high confidence. 05, used to perform independent survival verification, protein expression level verification, and immune infiltration correlation analysis on the core set of key genes, and output multidimensional verification results including prognostic robustness, expression specificity, and immune correlation; 06 is used to integrate the multidimensional validation results and biological function annotation information, and query the drug database to obtain drug-likeness information. Based on the preset multi-index evaluation system, the genes in the core set of key genes are comprehensively scored and ranked, and a list of potential key targets after priority ranking is output.

10. A machine learning-based screening system for key targets in lung adenocarcinoma according to claim 9, characterized in that, 01 is connected to 02, 02 is connected to 03 and 04, 04 is connected to 05, and 06 is connected to 03 and 05.