Machine-learning-based method for screening cross-differentially expressed genes between lung cancer and covid-19 infection and correspondingly screening prognostic genes of lung cancer

By using machine learning methods to screen differentially expressed genes between lung cancer and novel coronavirus infection, and identifying prognostic genes for lung cancer, the challenges of early screening and prognostic assessment for lung cancer have been solved, and the cure rate for lung cancer has been improved.

WO2026040166A1PCT designated stage Publication Date: 2026-02-26DALIAN NATIONALITIES UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120523
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2024-09-24
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively explore the relationship between novel coronavirus infection and lung cancer, making early screening and prognostic assessment of lung cancer difficult, and traditional treatments have low cure rates.

Method used

Machine learning methods, including Cox regression and LASSO regression, were used to screen for differentially expressed genes between lung cancer and novel coronavirus infection. Combined with KM survival analysis, GO and KEGG enrichment analysis, prognostic genes for lung cancer were identified and validated by protein interaction network analysis.

Benefits of technology

It improves the efficiency and accuracy of screening prognostic genes for lung cancer, identifies biomarkers with high sensitivity and specificity, and can effectively assess the prognosis of lung cancer, with universal applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120523_26022026_PF_FP_ABST
    Figure CN2024120523_26022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of bioinformatics, and relates to a machine-learning-based method for screening cross-differentially expressed genes between lung cancer and COVID-19 infection and correspondingly screening prognostic genes of lung cancer. The method comprises: acquiring cross-differentially expressed genes between lung cancer and COVID-19; screening the cross-differentially expressed genes between lung cancer and COVID-19 infection; using Cox regression and LASSO regression to correspondingly screen prognostic genes of lung cancer; and using a K-M survival analysis method, GO and KEGG enrichment analysis methods and a protein-protein interaction analysis method to analyze screened-out prognostic genes of lung cancer. The method provided herein can process data, and can also identify intricate patterns in the data, and has relatively high sensitivity and specificity, thereby improving the efficiency and accuracy of screening. The technique can be extended to the research of other diseases, and features universality.
Need to check novelty before this filing date? Find Prior Art

Description

Method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection by machine learning and corresponding screening of lung cancer prognosis genes TECHNICAL FIELD

[0001] The present application belongs to the technical field of bioinformatics, and particularly relates to a method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection by machine learning and corresponding screening of lung cancer prognosis genes. BACKGROUND

[0002] Cancer is one of the main factors threatening human life and health. In 2020, there were about 19.3 million new cancer cases worldwide, and about 9.96 million deaths. Among them, lung cancer incidence was about 2.21 million, and mortality was about 1.8 million, ranking first, far exceeding other cancers. In 2022, there were about 4.8247 million new cancer cases in China, and about 2.5742 million deaths, of which the incidence (about 22%) and mortality (about 28%) of lung cancer ranked first. In the United States in 2023, the top ten major cancer types by gender are predicted, with lung cancer ranking second in incidence in both men and women (about 12% in men and about 13% in women), and ranking first in mortality (about 21% in men and about 21% in women). With the continued spread of the novel coronavirus (COVID-19) infection since 2020, people with lung cancer may experience more severe symptoms, leading to an increase in mortality, so it is necessary to explore the relationship between the two and find cross-differentially expressed genes of lung cancer and novel coronavirus infection.

[0003] According to histopathological classification, lung cancer is mainly divided into non-small cell lung cancer (about 85%) and small cell lung cancer (about 15%). The current treatment for lung cancer is mainly surgical treatment, chemotherapy, radiotherapy and targeted therapy. However, because the early symptoms of lung cancer are not obvious, it is not easy to be found and the recurrence and metastasis rate of late-stage surgery is high, so the traditional treatment method has a low cure rate for lung cancer. Tumor markers can assist in cancer diagnosis, risk stratification and efficacy prognosis evaluation, and exploring markers that are beneficial for early screening and prognosis evaluation of lung cancer is the development direction to improve the cure rate of lung cancer. However, due to the development of sequencing technology, the complexity of lung cancer types, a large amount of genomic data has emerged, bringing great difficulties to screening work, and how to obtain key prognosis gene information of lung cancer has become a problem.

[0004] SUMMARY

[0005] In order to overcome the shortcomings of the prior art, the present application provides a method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection by machine learning and corresponding screening of lung cancer prognosis genes,

[0006] The above object of the application is achieved by the following technical solutions: a method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection by machine learning and corresponding screening of lung cancer prognosis genes, wherein the lung cancer is lung adenocarcinoma (LUAD) and lung squamous carcinoma (LUSC), and the machine learning method is Cox regression and LASSO regression; the method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection by machine learning and corresponding screening of lung cancer prognosis genes is specifically as follows:

[0007] 1. Obtaining cross-differentially expressed genes of lung cancer and novel coronavirus: downloading gene expression data and clinical sample data of LUAD and LUSC, and downloading COVID-19 related gene information;

[0008] 2. Screening cross-differentially expressed genes of lung cancer and novel coronavirus infection: taking the statistical significant difference value range and the differential expression fold range as the screening standard, screening the differentially expressed genes, and drawing a volcano plot, then taking the intersection of the differentially expressed genes of LUAD and LUSC and the COVID-19 related genes, and drawing a Venn diagram;

[0009] 3. Corresponding screening of lung cancer prognosis genes by Cox regression: combining the clinical sample information of LUAD and LUSC patients respectively, and screening the prognosis related genes corresponding thereto by using the single factor Cox regression method;

[0010] 4. Corresponding screening of lung cancer prognosis genes by LASSO regression: the sample number of the training set and the validation set is allocated according to a certain proportion, and then LASSO regression is used for corresponding screening;

[0011] 5. Analyzing the screened lung cancer prognosis genes by using K-M survival analysis method: the survival function is estimated by the following formula

[0012] Wherein S(t) is the survival probability of an individual at time t, n is the number of time points at which the event occurs, d i is the number of individuals who have the event at time t; n i is the number of individuals who still survive at time t, S(0) = 1;

[0013] The 3-7 year survival rates of the patients in the training set and the validation set are analyzed by using K-M survival analysis method respectively, and the performance of the model is evaluated by using the K-M survival curve (p < 0.05) and the AUC value of the ROC curve;

[0014] 6. The lung cancer prognostic genes screened are analyzed using GO and KEGG enrichment analysis methods: the differential expression gene information is saved in the format of vector or data box, GO enrichment analysis is performed using the enrichGO function, and the first 10 results of each function are plotted into a column chart using the ggplot2 function; KEGG enrichment analysis is performed using the enrichKEGG function, and the results of the first 10 pathways are plotted into a column chart using the ggplot2 function;

[0015] 7. The lung cancer prognostic genes screened are analyzed using protein interaction analysis method: STRING database and Cytoscape software are used to explore the protein-protein interaction (PPI) network involved in lung cancer prognostic genes. The screened differential expression genes are imported into the STRING database, and the protein-protein interaction network involved in the key genes is analyzed by combining the Cytoscape software.

[0016] Further, the source of step 1 is: downloading the gene expression data and clinical sample data of LUAD and LUSC from the Cancer Genome Atlas (TCGA) database; downloading COVID-19 related gene information from GeneCards, KEGG, NCBI, and OMIM databases.

[0017] Further, the specific screening criteria of step 2 are: statistical significant difference value p<0.05, and differential expression fold (FC) greater than or equal to 4 times, i.e. |log2FC|≥2, are used as screening criteria, R language DESeq2 package is used to screen differential expression genes, and a volcano plot is drawn; then the differential expression genes of LUAD and LUSC obtained are intersected with COVID-19 related genes, and a Venn diagram is drawn.

[0018] Further, the regression model risk score calculation formula of Cox regression method in step 3 is as follows: h(t,X)=h0(t)exp(β1X1+β2X2+…+β n X n ) (Formula 1)

[0019] Where β is the partial regression coefficient of Cox regression, X is the risk factor, h(t,X) represents the risk of event occurrence at time t under the condition of risk factor X, h0(t) is the base risk when all X is 0. Hazard ratio (HR) refers to the risk ratio of exposure to risk factor group and non-exposure to risk factor group:

[0020] In the Cox regression analysis result, HR>1 is regarded as an increased risk of death, that is, a risk factor; HR<1 is regarded as a reduced risk of death, that is, a protective factor; HR=1 is regarded as no effect on prognosis; for the binary classification problem (survival or death) of single factor Cox regression, X has two values of 1 and 0. At this time, the risk degree is:

[0021] Combined with clinical sample data and according to the single factor Cox regression method, genes related to the prognosis of LUAD and LUSC are screened respectively, and a forest plot is drawn, and a 95% confidence interval is given.

[0022] Further, the LASSO regression target function in step 4 is as follows: minimize||Y-Xw|| 2 +λ|||w||1 (Formula 4)

[0023] Where Y refers to the observed target variable, X is the feature matrix, w is the regression coefficient vector to be estimated, and λ is the parameter controlling the regularization strength. Formula 4 is composed of two terms, where the first term is the residual sum of squares in ordinary least squares, and the second term is the L1 regularization term also known as the penalty term; LASSO regression realizes feature selection and model simplification while fitting data by adjusting λ; when λ is large, smaller coefficients will be compressed to zero, thereby realizing feature selection, and when λ is small, more features will be retained;

[0024] The risk score of each patient is calculated according to the following formula.

[0025] Where β j is the regression coefficient, and x j is the expression amount of the gene.

[0026] The beneficial effects of the present application compared with the prior art are:

[0027] The bioinformatics method and model provided by the present application first explore the relationship between COVID-19 infection and lung cancer, and find more effective biomarkers for prognosis. The present application not only can process data, but also can identify complex patterns in data, has high sensitivity and specificity, and improves the efficiency and accuracy of screening. The technology can be popularized to other disease researches, and has universality. BRIEF DESCRIPTION OF DRAWINGS

[0028] The present application will be further described below in combination with the drawings and specific embodiments

[0029] Fig. 1 is a volcano plot of differentially expressed genes of LUAD (A) and LUSC (B);

[0030] Figure 2 is a Venn diagram of cross-differentially expressed genes of LUAD, LUSC, COVID-19;

[0031] Figure 3 is a forest plot of LUSC prognostic genes;

[0032] Figure 4 is a forest plot of LUAD prognostic genes;

[0033] Figure 5 is the number of key genes screened by LASSO (A) and the relationship between L1 regularization and LASSO regression coefficients (B);

[0034] Figure 6 is the gene expression analysis of ERO1A (A), AHNAK2 (B), NTS (C);

[0035] Figure 7 is the gene expression analysis of LGR4 (A), P3H4 (B), ABCC2 (C), TRIP13 (D), HMMR (E), KRT14 (F), UPK1B (G), IGF2BP1 (H), GRIA1 (I);

[0036] Figure 8 is the gene expression analysis of PTX3 (A), NEFL (B), SALL1 (C), EFNA2 (D), CAMP (E);

[0037] Figure 9 is the risk score distribution (A) and the survival status of patients (B);

[0038] Figure 10 is the K-M survival curve of the training set (A) and the validation set (B);

[0039] Figure 11 is the ROC curve of the training set (A) and the validation set (B);

[0040] Figure 12 is a column chart of GO enrichment analysis of 172 cross genes;

[0041] Figure 13 is a column chart of GO enrichment analysis of 17 prognostic genes;

[0042] Figure 14 is a column chart of KEGG enrichment analysis of 172 cross genes;

[0043] Figure 15 is a column chart of KEGG enrichment analysis of 17 prognostic genes;

[0044] Figure 16 is a protein interaction network relationship diagram involving 172 differentially expressed genes;

[0045] Figure 17 is a key protein interaction network diagram involving EFNA2. DETAILED DESCRIPTION

[0046] The application will be described in detail below with specific examples, but the protection scope of the application is not limited. Unless otherwise specified, the experimental methods used in the application are conventional methods, and the experimental equipment, materials, reagents, etc. used can be obtained from commercial channels.

[0047] Example 1

[0048] The method for screening cross-differentially expressed genes of lung cancer and novel coronavirus infection and corresponding screening lung cancer prognosis genes by machine learning is as follows:

[0049] S1. Obtain cross-differentially expressed genes of lung cancer and novel coronavirus: download gene expression data and clinical sample data of LUAD and LUSC from the Cancer Genome Atlas (TCGA) database; download COVID-19 related gene information from Gene Cards, KEGG, NCBI, and OMIM databases;

[0050] S2. Screen cross-differentially expressed genes of lung cancer and novel coronavirus infection: take statistical significant difference value p < 0.05 and difference expression fold (FC) greater than or equal to 4 times, i.e. |log2FC|≥2, as the screening standard, use the DESeq2 package of R language to screen differentially expressed genes, and draw a volcano plot; then take the intersection of the differentially expressed genes of LUAD and LUSC and the COVID-19 related genes, and draw a Venn diagram; obtain 172 cross-differentially expressed genes of LUAD, LUSC, and COVID-19.

[0051] S3. Corresponding screening lung cancer prognosis genes by Cox regression: respectively combine the clinical sample information of LUAD and LUSC patients, and use the single factor Cox regression method to screen the corresponding prognosis related genes. The risk score calculation formula of the Cox regression model is as follows: h(t,X) = h0(t)exp(β1X1+β2X2+…+β n X n ) (Formula 1)

[0052] Where β is the partial regression coefficient of Cox regression, X is the risk factor, h(t,X) represents the risk of occurrence at time t under the condition of risk factor X, h0(t) is the base risk when all X is 0. The hazard ratio (HR) refers to the risk ratio of exposure to risk factor group and non-exposure to risk factor group:

[0053] HR>1 in the Cox regression analysis result is considered to increase the risk of death, i.e. a risk factor; HR<1 is considered to reduce the risk of death, i.e. a protective factor; HR=1 is considered to have no effect on prognosis

[0076] For binary classification problem (survival or death) of single factor Cox regression, X has two values of 1 and 0. The risk degree at this time is:

[0054] Combine the clinical sample data and according to the single factor Cox regression method, the genes related to the prognosis of LUAD and LUSC are screened respectively and the forest plot is drawn, and the 95% confidence interval is given;

[0055] S4. Use LASSO regression to screen lung cancer prognosis genes: the sample number of training set and validation set (total sample size: 541) is allocated according to 7:3. The LASSO regression objective function is as follows: minimize||Y-Xw|| 2 +λ||w||1 (Formula 4)

[0056] Where Y refers to the observed target variable, X is the feature matrix, w is the regression coefficient vector to be estimated, and λ is the parameter controlling the regularization strength. Formula 4 consists of two parts, where the first part is the residual sum of squares in ordinary least squares, and the second part is the L1 regularization term also known as the penalty term. LASSO regression can achieve feature selection and model simplification while fitting the data by adjusting λ. When λ is large, smaller coefficients will be compressed to zero, thereby achieving feature selection, while when λ is small, more features will be retained. Therefore, LASSO regression can help us identify features that have important influence on the target variable. The risk score of each patient is calculated according to the following formula.

[0057] Where β j is the regression coefficient, and x j is the expression amount of the gene. 17 lung cancer prognosis genes are obtained by using Cox and LASSO regression methods.

[0058] S5. Use K-M survival analysis method to analyze the screened lung cancer prognosis genes: the survival function is estimated by the following formula:

[0059] Where S(t) is the survival probability of an individual at time t, n is the number of time points at which the event occurs, d i is the number of individuals who have events at time t; n i is the number of individuals who are still alive at time t, and S(0) = 1.

[0060] K-M survival analysis was used to analyze the 3-7 year survival rates of patients in the training set and the validation set respectively. The performance of the model was evaluated using K-M survival curve (p<0.05) and AUC value of ROC curve, the larger the AUC value, the higher the prediction accuracy (AUC value greater than 0.7 is considered to be an effective model).

[0061] S6. The lung cancer prognostic genes screened were analyzed using GO and KEGG enrichment analysis methods: the information of 172 differentially expressed genes was saved in vector or data frame format, GO enrichment analysis was performed using the enrichGO function, and the first 10 results of each function were plotted into a column chart using the ggplot2 function; KEGG enrichment analysis was performed using the enrichKEGG function, and the results of the first 10 pathways were plotted into a column chart using the ggplot2 function.

[0062] S7. The lung cancer prognostic genes screened were analyzed using protein interaction analysis method: STRING database and Cytoscape software were used to explore the protein-protein interaction (PPI) network involved in the key genes. The 172 differentially expressed genes screened were imported into the STRING database, and the protein-protein interaction network involved in the key genes was analyzed using Cytoscape software.

[0063] Figure 1 shows the distribution of differentially expressed genes in the form of a volcano plot. The horizontal axis is the logarithmic representation of the differential expression fold (base 2), where the positive and negative values of log2 Fold Change represent the up-regulation (red points in the figure) or down-regulation (green points in the figure) of gene expression in tumor samples, and the larger the difference, the more it is distributed to the ends of the X axis. The gray area in the figure represents genes with no significant changes. The vertical axis is the -log10 transformed representation of p value, and the smaller the p value, the more significant it is, and the more it is distributed upwards. Among the gene expression data (60660) of LUAD, 2597 differentially expressed genes were screened (Figure 1A), and among the gene expression data (60660) of LUSC, 3954 differentially expressed genes were screened (Figure 1B). Taking the intersection with COVID-19 genes (4783), a total of 172 cross-differentially expressed genes of the three lung diseases were obtained (Figure 2).

[0064] The results of single factor Cox regression are shown in the form of a forest plot (Figures 3 and 4). The dotted line with HR=1 is the central axis, the green points to the left of the dotted line are factors that reduce the occurrence of death events, i.e. low-risk genes, the red points to the right of the dotted line are factors that promote the occurrence of death events, i.e. high-risk genes, and the blue line segment where each point is located is the 95% confidence interval.

[0065] After single factor COX regression model calculation, 77 genes related to LUAD prognosis were screened, including 23 tumor suppressor genes and 54 oncogenes (Figure 3); 15 genes related to LUSC prognosis, 4 tumor suppressor genes and 11 oncogenes (Figure 4), among which NTS, CDCA3, CHODL, MEGF10 were highly expressed in lung squamous carcinoma tissues compared with normal tissues, and ABCA3, CD93, MARCO, CXCL2, LPL, ZBTB16, PTX3, C8B, PPBP, F11, CLEC4M were lowly expressed in lung squamous carcinoma tissues.

[0066] Through LASSO regression model calculation, 17 key prognostic genes were screened (Figure 5), which were ERO1A, AHNAK2, NTS, LGR4, P3H4, ABCC2, TRIP13, HMMR, KRT14, UPK1B, IGF2BP1, GRIA1, PTX3, NEFL, SALL1, EFNA2, and CAMP. Further analysis of the expression of 17 key prognostic genes in tumor samples and normal samples was performed by R language (Figures 6, 7, and 8).

[0067] Among them, ERO1A, AHNAK2, LGR4, P3H4, ABCC2, TRIP13, HMMR, KRT14, IGF2BP1, NEFL, SALL1, EFNA2 were highly expressed in lung adenocarcinoma tissues, and NTS, UPK1B, GRIA1, PTX3, CAMP were lowly expressed in lung adenocarcinoma tissues, and a prognosis model was established combining the regression coefficients.

[0068] risk.score <- 0.030753725 x ERO1A + 0.009632997 x AHNAK2 + 0.020734698 x NTS + 0.05924787 x LGR4 + 0.025156105 x P3H4 + 0.016861383 x ABCC2 - 0.067831582 x TRIP13 + 0.135319151 x HMMR + 0.025039238 x KRT14 + 0.006908539 x UPK1B + 0.0245833 x IGF2BP1 - 0.074275033 x GRIA1 + 0.058295983 x PTX3 + 0.035083611 x NEFL + 0.039350275 x SALL1 + 0.039424578 x EFNA2 - 0.076919192 x CAMP

[0069] With the increase of risk score, the number of deaths also increased, that is, the survival rate of patients decreased (Figure 9).

[0070] To further verify the effectiveness of the model, survival analysis was performed on the samples of the training set and the validation set, respectively. The survival status of the high-risk and low-risk groups was displayed by K-M survival curve, and p<0.05 was used to determine whether there was a significant difference. The survival curves of the high-risk and low-risk groups in the training set and the validation set did not cross, and the p values were less than 0.05 (Figure 10), indicating that the prognosis model could significantly distinguish the high-risk and low-risk groups. The 17 key genes screened by the model can be used as the prognosis markers of this type of disease. In the training set, the AUC values of 3-year and 7-year overall survival were 0.732 and 0.739, respectively, and in the validation set, the AUC values of 3-year and 7-year overall survival were 0.704 and 0.709, respectively (Figure 11). The results show that with the increase of time, the prognosis model has achieved good results in the training set and the validation set samples, that is, the model has a certain effectiveness.

[0071] To explore the cell components, biological processes and molecular functions and signal pathways involved in the 172 differentially expressed genes and 17 key prognosis genes, GO and KEGG enrichment analysis was performed on these genes, respectively. GO enrichment results (Figures 12, 13) show that the 172 cross-differentially expressed genes of the three lung diseases are mainly involved in biological processes such as vascular processes in the circulatory system, leukocyte migration, and humoral immune response. The 17 key prognosis genes are mainly involved in biological processes such as cell differentiation of bone remodeling, telencephalon development, synaptic protein complex assembly, and glomerular development. When p<0.05, Ephrin A2 (EFNA2) is involved in more biological processes and cell components. KEGG enrichment results show that the cross-differentially expressed genes are related to coronavirus disease-COVID-19, IL-17 signaling pathway, etc. KEGG enrichment results (Figures 14, 15) show that the cross-differentially expressed genes are related to malaria, coronavirus disease-COVID-19, IL-17 signaling pathway, etc.

[0072] PPI network results (Figures 16, 17) show that EFNA2 protein has a direct interaction relationship with TEK protein. EFNA2 can bind to TEK protein and activate the TEK signaling pathway, thereby affecting the proliferation, migration and vascularization of vascular endothelial cells. This interaction plays an important role in regulating vascular development and homeostasis, and may also play an important role in the process of tumor angiogenesis.

[0073] The above-described embodiments are only preferred embodiments of the present application, and not all the embodiments of the present application that can be implemented. Any obvious modifications made by those skilled in the art without departing from the principles and spirit of the present application should be considered to be included in the protection scope of the claims of the present application.

Claims

1. A method of cross-differentially expressed genes screening lung cancer and novel coronavirus infection and corresponding screening lung cancer prognosis genes by machine learning, characterized in that, The steps are as follows: S1. Cross-differential expression genes of lung cancer and novel coronavirus are obtained: download gene expression data and clinical sample data of LUAD and LUSC, and download COVID-19 related gene information; S2. Cross-differential expression genes of lung cancer and novel coronavirus infection are screened: statistical significant difference value range and differential expression fold range are used as screening criteria to screen differential expression genes, and a volcano plot is drawn, then the intersection of the obtained differential expression genes of LUAD and LUSC and the COVID-19 related genes is taken, and a Venn diagram is drawn; S3. Lung cancer prognostic genes are screened by Cox regression: the clinical sample information of LUAD and LUSC patients is combined respectively, and the single factor Cox regression method is used to screen the prognostic related genes corresponding thereto; S4. Lung cancer prognostic genes are screened by LASSO regression: the sample number of the training set and the validation set is allocated according to a certain proportion, and then LASSO regression is used for corresponding screening; S5. Analysis of the lung cancer prognosis genes screened using K-M survival analysis: the survival function is estimated by the following formula where S(t) refers to the probability of survival of an individual at time t, n refers to the number of time points at which an event occurs, d i refers to the number of individuals that have an event at time t; n i S(t) refers to the number of individuals still alive at time t, S(0) = 1 ; The 3-7 year survival rates of patients in the training set and the validation set are analyzed by K-M survival analysis method, and the performance of the model is evaluated by setting p<0.05 and AUC value of ROC curve of K-M survival curve; S6. The lung cancer prognostic genes screened are analyzed by GO and KEGG enrichment analysis method: the differential expression gene information is saved in the format of vector or data box, GO enrichment analysis is performed by using enrichGO function, and the first 10 results of each function are drawn into a column chart by using ggplot2 function; KEGG enrichment analysis is performed by using enrichKEGG function, and the results of the first 10 pathways are drawn into a column chart by using ggplot2 function; S7. The lung cancer prognostic genes screened are analyzed by protein interaction analysis method: STRING database and Cytoscape software are used to explore the protein interaction network involved in the lung cancer prognostic genes. The screened differential expression genes are imported into STRING database, and the protein interaction network involved in the key genes is analyzed by using Cytoscape software.

2. The method for cross-differentially expressed genes screening lung cancer and novel coronavirus infection and corresponding screening lung cancer prognosis genes according to claim 1, characterized in that, The source of obtaining in step S1 is: downloading the gene expression data and clinical sample data of LUAD and LUSC from the Cancer Genome Atlas database; downloading COVID-19 related gene information from GeneCards, KEGG, NCBI, and OMIM databases. 3.The method of claim 1, wherein the cross-differentially expressed genes between lung cancer and COVID-19 are screened and the lung cancer prognosis genes are screened in response thereto. The specific screening criteria of step S2 are: statistical significant difference value p<0.05, differential expression fold, i.e. FC is greater than or equal to 4 times, i.e. |log2FC|≥2, is used as the screening criterion, DESeq2 package of R language is used to screen differential expression genes, and a volcano plot is drawn; then the intersection of the obtained differential expression genes of LUAD and LUSC and the COVID-19 related genes is taken, and a Venn diagram is drawn. 4.The method of cross-differentially expressed genes screening lung cancer and novel coronavirus infection and corresponding genes screening lung cancer prognosis according to claim 1, wherein, The regression model risk score calculation formula of the Cox regression method in the step S3 is as follows: h(t,X) = h0(t)exp(β1X1+β2X2+…+β n X n ) (Formula 1) where β is the partial regression coefficient of Cox regression, X is the risk factor, h(t, X) represents the risk of occurrence of the event at time t in the presence of the risk factor X, and h0(t) is the base risk when all X is 0. The hazard ratio (HR) refers to the ratio of the risk degree of the exposure group to the risk factor and the non-exposure group to the risk factor: In the Cox regression analysis results, HR>1 is considered as an increased risk of death, i.e. a risk factor; HR<1 is considered as a reduced risk of death, i.e. a protective factor; HR=1 is considered as having no effect on prognosis; for a binary classification problem of single-factor Cox regression, i.e. survival or death, X has two values of 1 and 0. At this time, the risk degree is: The genes related to the prognosis of LUAD and LUSC are screened according to the single factor Cox regression method combined with the clinical sample data, a forest plot is drawn, and a 95% confidence interval is given. 5.The method of cross-differentially expressed genes screening lung cancer and novel coronavirus infection and corresponding genes screening lung cancer prognosis according to claim 1, wherein, The LASSO regression objective function in the step S4 is as follows: minimize||Y-Xw||2+λ||w||1 2 +λ||w||1 (Formula 4) where Y refers to the observed target variable, X is the feature matrix, w is the vector of regression coefficients to be estimated, and λ is a parameter that controls the strength of regularization. Equation 4 consists of two terms, where the first term is the residual sum of squares in ordinary least squares, and the second term is the LI regularization term, also known as the "penalty term"; LASSO regression achieves feature selection and model simplification while fitting the data by adjusting λ; when λ is large, smaller coefficients will be shrunk to zero, thus achieving feature selection, while when λ is small, more features will be retained; ​ The risk score for each patient is calculated according to the following formula. where β j is the regression coefficient, x j is the expression level of the gene.

Citation Information

Patent Citations

  • COVID-19 and lung cancer marker screening and prognosis risk model construction method

    CN115841844A

  • Identifying impact of COVID-19 on HIV patients based on bioinformatics and systematic biology

    CN115954047A

  • Biomarker and model for prognosis risk prediction of colorectal cancer and application of biomarker and model

    CN116030880A

  • Non-small cell lung cancer prognosis model and construction method thereof

    CN117038086A