Construction method of non-small cell lung cancer prognosis model

By collecting and integrating data from small cell lung cancer patients from the cancer genome map database, using Cox regression analysis and random forest method to screen features, constructing and evaluating the prognosis model of non-small cell lung cancer, the problem of existing models being sensitive to data quality is solved, and the accuracy and stability of the model are improved.

CN120199493AInactive Publication Date: 2025-06-24SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510303137.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The accuracy of the existing non-small cell lung cancer prognosis model depends on the quality and integrity of the data set. If there are deviations, omissions or errors in the data, it is easy to affect the model's prediction results.

Method used

By collecting clinical information and molecular data of small cell lung cancer patients from the cancer genome map database, after integration and pretreatment, important prognostic characteristics were screened using Cox regression analysis and random forest method to construct prognostic models, and the stability and generalization ability of the model were evaluated through K-fold cross-validation.

Benefits of technology

The generalization ability of the model and the stable identification ability of important features are improved, the prediction accuracy of the prognostic model is ensured, the stability and reliability of the evaluation results are increased, and the utilization rate of data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199493A_ABST
    Figure CN120199493A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a non-small cell lung cancer prognosis model, and belongs to the technical field of tumor prognosis model construction. The invention relates to a construction method of a non-small cell lung cancer prognosis model. The construction method comprises the following steps: collecting clinical information and molecular data of a small cell lung cancer patient; extracting important prognosis features in the original data set; evaluating the stability and generalization ability of the prognosis model; and verifying the prediction capability of the prognosis model. According to the method, the problem that the prediction result of the model is easily influenced if deviation, omission or wrong records exist in a data set in the prior art is solved, the independent influence of each feature on prognosis can be more accurately evaluated, the flexibility is high, data analysis is more accurate, and the method is suitable for popularization and application. The method improves the generalization ability of the model and the stable recognition ability of important features, guarantees the prediction accuracy of the prognosis model, improves the utilization rate of data, enables the performance evaluation of the prognosis model to be more representative, and improves the stability and reliability of an evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tumor prognosis model construction, and particularly to a method for constructing a non-small cell lung cancer prognosis model. Background Art

[0002] At present, after multidisciplinary comprehensive treatment, the prognosis of non-small cell lung cancer has been greatly improved, and tumor invasion and metastasis are the most important factors affecting prognosis.

[0003] Chinese Patent with publication number CN117038086A discloses a non-small cell lung cancer prognosis model and its construction method. It mainly screens genes with significant differential expression after treatment with G3 helical peptide and Gef combined drugs, further screens genes with significant differential expression that are significantly correlated with patient survival in terms of expression level, verifies the correlation between the screened genes and prognosis risk within the TCGA group and outside the GEO group, and assigns values according to the relationship between the expression level of the significantly correlated genes obtained by screening and the patient survival period; then draws a nomogram through the survival and rms packages in R language to obtain a prognosis score table for non-small cell lung adenocarcinoma patients, which is the non-small cell lung cancer prognosis model. This model is accurate and reliable, can predict the 10-year survival rate of patients with middle and advanced lung adenocarcinoma, and provide reference for targeted drug use by patients.

[0004] In the actual use of the above patent, since the accuracy of the nomogram highly depends on the quality and integrity of the dataset used to construct the model, if there are biases, omissions or incorrect records in the dataset, it is easy to affect the prediction results of the model; therefore, it does not meet the existing needs, and for this reason, we propose a method for constructing a non-small cell lung cancer prognosis model. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for constructing a non-small cell lung cancer prognosis model, which can more accurately evaluate the independent influence of each feature on prognosis, has high flexibility and more accurate data analysis, improves the generalization ability of the model and the stable recognition ability for important features, ensures the accuracy of prognosis model prediction, improves the utilization rate of data, makes the performance evaluation of the prognosis model more representative, and increases the stability and reliability of the evaluation results, and solves the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a non-small cell lung cancer prognosis model, the construction method includes: Collect the clinical information of small cell lung cancer patients from the Cancer Genome Atlas database, and at the same time obtain the genomics, transcriptomics, proteomics and metabolomics data of the patients and perform preprocessing to obtain the molecular data of the patients; Integrate clinical data and molecular data into a unified dataset to obtain the original dataset, and extract important prognostic features from the original dataset; Construct a prognostic model using the important prognostic features, and evaluate the stability and generalization ability of the prognostic model using K-fold cross-validation; After verification, use an independent test set to verify the predictive ability of the prognostic model to obtain a non-small cell lung cancer prognostic model.

[0007] Preferably, collect clinical information of small cell lung cancer patients from the Cancer Genome Atlas database, specifically including: Retrieve clinical information from the Cancer Genome Atlas database according to a preset clinical information retrieval frequency; Real-time monitor the proportion of missing data values corresponding to various data types included in the clinical information; Compare the proportion of missing data values corresponding to each data type in each clinical information retrieval process with a preset proportion threshold; Extract the data type with a proportion of missing data values exceeding the preset value as the target data type in each clinical information retrieval process; Utilize the similarity values of the target data types in every two adjacent clinical information retrieval processes; Utilize the similarity values of the target data types in every two adjacent clinical information retrieval processes to determine whether there is an abnormality in data retrieval.

[0008] Preferably, utilize the similarity values of the target data types in every two adjacent clinical information retrieval processes to determine whether there is an abnormality in data retrieval, specifically including: Extract the similarity values of the target data types in every two adjacent clinical information retrieval processes; Compare the similarity values of the target data types in every two adjacent clinical information retrieval processes with a preset similarity threshold; When the similarity values of the target data types in every two adjacent clinical information retrieval processes are not lower than the preset similarity threshold, then use the first anomaly coefficient acquisition model to obtain the first anomaly coefficient; Among them, the structure of the first anomaly coefficient acquisition model is as follows: Among them, Y 01 represents the first anomaly coefficient obtained by the first anomaly coefficient acquisition model; m represents the number of groups of clinical information retrieved every two adjacent times; S i represents the similarity value corresponding to the i-th group of clinical information; S c represents the preset similarity threshold; S m represents the lowest value of the similarity values corresponding to m groups of clinical information; S dRepresents the standard deviation of the similarity values corresponding to m groups of clinical information; When the similarity value of the target data type in each adjacent two clinical information retrieval processes is lower than the preset similarity threshold, the weighted average value of the same target data type in each adjacent two clinical information retrieval processes is retrieved; Use the weighted average value of the same target data type in each adjacent two clinical information retrieval processes to obtain the second anomaly coefficient by combining with the second anomaly coefficient acquisition model; Among them, the structure of the second anomaly coefficient acquisition model is as follows: Among them, Y 02 Represents the first anomaly coefficient obtained by the second anomaly coefficient acquisition model; m represents the number of groups of clinical information retrieved every two adjacent times; S i Represents the similarity value corresponding to the i-th group of clinical information; S c Represents the preset similarity threshold; S d Represents the standard deviation of the similarity values corresponding to m groups of clinical information; λ i Represents the weighted average value of the same target data type corresponding to the i-th group of clinical information; Compare the first anomaly coefficient and the second anomaly coefficient with their corresponding anomaly coefficient thresholds respectively; When the first anomaly coefficient and the second anomaly coefficient exceed their corresponding anomaly coefficient thresholds, it is determined that there is an abnormality in data retrieval, and an alarm for data retrieval abnormality is given.

[0009] Preferably, obtaining and preprocessing the patient's genomics, transcriptomics, proteomics and metabolomics data specifically includes: Obtain the patient's genomics, transcriptomics, proteomics and metabolomics data from professional omics databases, research institutions or hospitals respectively; Use statistical methods to fill in the missing values in the genomics, transcriptomics, proteomics and metabolomics data and complete the missing values; Screen and sort the genomics, transcriptomics, proteomics and metabolomics data, and identify and process the features related to survival time in the genomics, transcriptomics, proteomics and metabolomics data; Delete the duplicate data values in the genomics, transcriptomics, proteomics and metabolomics data, and keep the first record of the duplicate data when deleting; Perform normalization processing on the genomics, transcriptomics, proteomics and metabolomics data; Preferably, extracting the important prognostic features in the original dataset specifically includes: Screen features related to survival time from the original dataset using statistical test methods; Use Cox regression analysis to screen features related to survival time and select features related to prognosis in a multivariate environment as initial features; Use the random forest method to identify important prognostic features among the initial features.

[0010] Preferably, the step of screening features related to survival time from the original dataset using statistical test methods specifically includes: Set the upper and lower thresholds of features related to survival time, identify features related to survival time in the original dataset, and remove data points beyond the thresholds; Establish a data statistical model, check whether the data points conform to the prediction of the statistical model, and the data points that do not conform are features related to survival time; Divide the data into different data groups, check for isolated points that are significantly different from other data points. For time series data, calculate the statistical characteristics of each data point using a sliding window, and identify features related to survival time based on the characteristics.

[0011] Preferably, the step of using Cox regression analysis to screen features related to survival time and select features related to prognosis in a multivariate environment as initial features specifically includes: Use Cox regression analysis to analyze each feature related to survival time separately and calculate its probability value; Set the threshold of the probability value, select features with a probability value less than the threshold according to the calculation results, and put the selected features into a multivariate Cox regression model; Combining the influence of multiple features on prognosis, use stepwise regression method to optimize the model, remove features with insignificant influence on prognosis in a multivariate environment, and obtain features related to prognosis in a multivariate environment.

[0012] Preferably, the step of using the random forest method to identify important prognostic features among the initial features specifically includes: Select n sample data from the initial features as a training set and generate a decision tree; During the process of generating the decision tree, randomly select features for node splitting, find the best splitting feature, and repeat the above process until a specified number of decision trees are generated; Calculate the purity improvement value brought by each feature related to prognosis in the node splitting of the decision tree, and use the purity improvement value to evaluate the importance of the feature related to prognosis; Average the impurity reduction values of each feature in all decision trees to obtain the importance evaluation value of the feature; According to the importance evaluation value of the feature, sort the features and select the features with higher rankings as important prognostic features.

[0013] Preferably, constructing a prognostic model using important prognostic features and evaluating the stability and generalization ability of the prognostic model using K-fold cross-validation specifically includes: Constructing a prognostic model using the selected important prognostic features and setting initial model parameters; Randomly dividing the original dataset into K non-overlapping subsets with approximately equal sample sizes for each subset; For each K from 1 to K, using the K-th subset as the validation set and the remaining K - 1 subsets as the training set; Training the prognostic model on the training set and evaluating the performance of the model on the validation set, and recording the validation metrics; Repeating the above steps K times to ensure that each subset has the opportunity to be used as the validation set and the remaining subsets as the training set; Collecting the results of all K validations and calculating the average of all K validation results, which is the estimated value of the prognostic model performance, and using the estimated value to evaluate the stability and generalization ability of the prognostic model.

[0014] Preferably, verifying the predictive ability of the prognostic model using an independent test set specifically includes: Obtaining an independent dataset related to the original dataset as the test set, inputting the obtained test set into the evaluated prognostic model, and generating a prediction result; Comparing the generated prediction result with the actual result to obtain the predictive performance of the prognostic model according to the comparison result.

[0015] Compared with the prior art, the beneficial effects of the present invention are: The present invention uses Cox regression analysis to screen features related to prognosis in a multi-variable environment, can simultaneously consider the effects of multiple covariates on survival time, thereby more accurately evaluating the independent impact of each feature on prognosis, making full use of existing information for analysis, with high flexibility and more accurate data analysis. The random forest algorithm can effectively handle high-dimensional data and large-scale datasets, reasonably screen out variables that have a significant impact on the classification result, significantly reduce the risk of overfitting, improve the generalization ability of the model and the stable recognition ability for important features, ensure the accuracy of the prognostic model prediction. Using K-fold cross-validation to evaluate the stability and generalization ability of the prognostic model improves the utilization rate of data, makes the performance evaluation of the prognostic model more representative, increases the stability and reliability of the evaluation result, ensures the objectivity of the verification process, and avoids evaluation biases caused by data overlap. Verifying the predictive ability of the prognostic model using an independent test set can also systematically evaluate the generalization performance of the model. Description of the Drawings

[0016] Figure 1Schematic diagram of the construction method of the non-small cell lung cancer prognosis model of the present invention; Figure 2 Schematic diagram of the extraction of important prognostic features of the construction method of the non-small cell lung cancer prognosis model of the present invention. Detailed implementation manners

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] In order to solve the problem that in the actual utilization process of the prior art, since the accuracy of the nomogram highly depends on the quality and integrity of the dataset used to construct the model, if there are deviations, omissions or incorrect records in the dataset, it is easy to affect the prediction results of the model, please refer to Figure 1 - Figure 2 This embodiment provides the following technical solutions: A construction method of a non-small cell lung cancer prognosis model, the construction method comprising: Collect the clinical information of small cell lung cancer patients from the Cancer Genome Atlas database, and at the same time obtain the genomics, transcriptomics, proteomics and metabolomics data of the patients and perform preprocessing to obtain the molecular data of the patients, so as to facilitate the comparison of data between different sample data; Integrate the clinical data and molecular data into a unified dataset to obtain an original dataset, and extract important prognostic features from the original dataset; Use the important prognostic features to construct a prognosis model, and use K-fold cross-validation to evaluate the stability and generalization ability of the prognosis model; After the verification is completed, use an independent test set to verify the prediction ability of the prognosis model to obtain a non-small cell lung cancer prognosis model. The purpose of constructing the non-small cell lung cancer prognosis model is to predict the survival period or the risk of disease progression of patients by integrating clinical, pathological and molecular features.

[0019] Specifically, collecting the clinical information of small cell lung cancer patients from the Cancer Genome Atlas database specifically includes: Retrieve clinical information from the Cancer Genome Atlas database according to a preset clinical information retrieval frequency; Real-time monitor the proportion of data missing values corresponding to various data types included in the clinical information; Compare the proportion of data missing values corresponding to each data type in each clinical information retrieval process with a preset proportion threshold; During each clinical information retrieval process, extract the data type with a data missing value ratio exceeding a preset value as the target data type; Use the similarity value of the target data type during every two adjacent clinical information retrieval processes; Use the similarity value of the target data type during every two adjacent clinical information retrieval processes to determine whether there is an abnormality in data retrieval.

[0020] The technical effects of the above technical solution are as follows: By monitoring in real time the ratio of data missing values corresponding to various data types in clinical information, it is possible to accurately understand the completeness of each data type in the retrieved clinical information, timely discover which data types have a high risk of data loss, and provide an accurate basis for integrity assessment for subsequent data processing and analysis. Extracting the data type with a data missing value ratio exceeding the preset value as the target data type can focus on the data that may be crucial for the analysis of the clinical information of small cell lung cancer patients, avoid affecting the judgment and research of the patient's condition due to the lack of key data, and help ensure that the collected data can meet the clinical research and diagnosis needs to the greatest extent and improve the usability of the data.

[0021] At the same time, comparing the ratio of data missing values corresponding to each data type during each clinical information retrieval process with the preset ratio threshold and analyzing the similarity value of the target data type during two adjacent clinical information retrieval processes can effectively identify possible abnormal situations in the data retrieval process. If the similarity of the target data type between two adjacent times is abnormally low, it indicates that there are problems in data retrieval, such as data source failures, transmission errors, etc., which helps to correct data errors in a timely manner and ensure data accuracy.

[0022] At the same time, by determining whether there is an abnormality in data retrieval, the quality of the data can be controlled. Once an abnormality is found, measures can be taken in a timely manner to adjust or retrieve the data again, thereby improving the overall quality of the clinical information of small cell lung cancer patients collected from the Cancer Genome Atlas database, providing a reliable data basis for subsequent data analysis and research, and reducing research biases and errors caused by inaccurate data. Using the similarity value of the target data type during every two adjacent clinical information retrieval processes to determine whether there is an abnormality in data retrieval can reflect the stability of the data retrieval process from the side. Stable data retrieval is crucial for continuously and reliably obtaining clinical information and can ensure the smooth progress of clinical research and diagnosis work. According to the stability of data retrieval, parameters such as the preset clinical information retrieval frequency can be optimized and adjusted. If it is found that the data retrieval is unstable or there are many abnormalities, the retrieval frequency can be appropriately adjusted or other optimization measures can be taken to improve the efficiency and stability of data collection and ensure that high-quality clinical information can be stably obtained.

[0023] Specifically, the similarity value of the target data type in each adjacent two clinical information retrieval processes is used to determine whether there is an abnormality in data retrieval, which specifically includes: Extract the similarity value of the target data type in each adjacent two clinical information retrieval processes; Compare the similarity value of the target data type in each adjacent two clinical information retrieval processes with a preset similarity threshold; When the similarity value of the target data type in each adjacent two clinical information retrieval processes is not lower than the preset similarity threshold, the first anomaly coefficient acquisition model is used to obtain the first anomaly coefficient; Among them, the structure of the first anomaly coefficient acquisition model is as follows: Among them, Y 01 represents the first anomaly coefficient obtained by the first anomaly coefficient acquisition model; m represents the number of groups of adjacent two clinical information retrievals; S i represents the similarity value corresponding to the i-th group of clinical information; S c represents the preset similarity threshold; S m represents the minimum value of the similarity values corresponding to m groups of clinical information; S d represents the standard deviation of the similarity values corresponding to m groups of clinical information; When the similarity value of the target data type in each adjacent two clinical information retrieval processes is lower than the preset similarity threshold, the weighted average value of the same target data type in each adjacent two clinical information retrieval processes is retrieved; The weighted average value of the same target data type in each adjacent two clinical information retrieval processes is used in combination with the second anomaly coefficient acquisition model to obtain the second anomaly coefficient; Among them, the structure of the second anomaly coefficient acquisition model is as follows: Among them, Y 02 represents the first anomaly coefficient obtained by the second anomaly coefficient acquisition model; m represents the number of groups of adjacent two clinical information retrievals; S i represents the similarity value corresponding to the i-th group of clinical information; S c represents the preset similarity threshold; S d represents the standard deviation of the similarity values corresponding to m groups of clinical information; λ i represents the weighted average value of the same target data type corresponding to the i-th group of clinical information; Compare the first anomaly coefficient and the second anomaly coefficient with their corresponding anomaly coefficient thresholds respectively; When the first abnormal coefficient and the second abnormal coefficient exceed their corresponding abnormal coefficient thresholds, it is determined that there is an abnormality in data retrieval, and an alarm for abnormal data retrieval is given.

[0024] The technical effects of the above technical solution are as follows: By comparing the similarity value of the target data type in each adjacent two clinical information retrieval processes with the preset similarity threshold, and then using different abnormal coefficient acquisition models (the first abnormal coefficient acquisition model and the second abnormal coefficient acquisition model) to calculate the abnormal coefficient according to the comparison results. This multi-dimensional judgment method comprehensively considers multiple factors such as the similarity value, the lowest similarity value, the standard deviation, and the weighted average value of the same target data type, and can analyze the data retrieval situation more comprehensively and meticulously, so as to more accurately judge whether there is an abnormality in data retrieval and reduce the possibility of misjudgment and missed judgment. At the same time, the two abnormal coefficient acquisition models perform quantitative calculations on relevant factors to obtain specific abnormal coefficients (the first abnormal coefficient and the second abnormal coefficient). These quantified abnormal coefficients provide a clear basis for judging the abnormal degree of data retrieval, enabling the precise determination of the severity of the abnormality in the data retrieval process according to the comparison between the abnormal coefficient and the corresponding threshold, and improving the accuracy and reliability of abnormal detection.

[0025] On the other hand, closely monitor the similarity change of the target data type during each adjacent two clinical information retrieval processes in real time. When the similarity value shows a large fluctuation (below the preset threshold), further analyze by calculating the weighted average value of the same target data type and combining it with the second anomaly coefficient to obtain a model. This timely monitoring and response to the similarity change can quickly detect unstable factors in the data retrieval process, help take timely measures to stabilize the data retrieval, and ensure the continuity and stability of data acquisition. Different anomaly coefficient calculation methods are adopted according to different similarity situations, enabling the evaluation of data retrieval stability to be dynamically adjusted according to the actual situation. When the similarity is high, use the first anomaly coefficient to obtain a model for evaluating anomalies; when the similarity is low, combine the weighted average value and use the second anomaly coefficient to obtain a model for evaluating anomalies. This dynamically adjusted evaluation method can better adapt to various situations in the data retrieval process and ensure the effective guarantee of data retrieval stability. Once it is determined that there is an anomaly in the data retrieval (that is, the first anomaly coefficient and the second anomaly coefficient exceed their corresponding anomaly coefficient thresholds), an alarm for data retrieval anomaly is issued. This can timely remind relevant personnel that there are problems in the data retrieval process, prompt them to take measures to correct the anomalies, and avoid incorrect or incomplete data from flowing into subsequent analysis and application links, thereby ensuring the quality of the data and providing reliable data support for research and diagnosis based on these clinical information. By monitoring and processing data retrieval anomalies, problems existing in the data acquisition process can be continuously discovered, and then the data retrieval process, parameters, etc. can be optimized. For example, adjust the preset similarity threshold, anomaly coefficient threshold, etc. according to the anomaly situation to make the data retrieval more stable and accurate, continuously improve the data quality, and enhance the efficiency and reliability of data acquisition.

[0026] Obtain and preprocess the genomics, transcriptomics, proteomics, and metabolomics data of the patient, specifically including: Respectively obtain the genomics, transcriptomics, proteomics, and metabolomics data of the patient from professional omics databases, research institutions, or hospitals; Use statistical methods to fill in the missing values in the genomics, transcriptomics, proteomics, and metabolomics data and complete the missing values; Screen and sort the genomics, transcriptomics, proteomics, and metabolomics data, and identify and process the features related to survival time in the genomics, transcriptomics, proteomics, and metabolomics data; Delete the duplicate data values in the genomics, transcriptomics, proteomics, and metabolomics data, and retain the first record of the duplicate data when deleting; Perform normalization processing on the genomics, transcriptomics, proteomics, and metabolomics data; Extract important prognostic features from the original dataset, specifically including: Use statistical tests to screen for features related to survival time from the original dataset; Use Cox regression analysis to screen for features related to survival time that are prognostic-related in a multivariate setting as initial features; Use the random forest method to identify important prognostic features among the initial features.

[0027] Use statistical tests to screen for features related to survival time from the original dataset, specifically including: Set upper and lower thresholds for features related to survival time, identify features related to survival time in the original dataset, and remove data points beyond the thresholds; Establish a data statistical model, check whether data points conform to the prediction of the statistical model, and data points that do not conform are features related to survival time; Divide the data into different data groups, check for isolated points that are significantly different from other data points. For time series data, calculate the statistical characteristics of each data point using a sliding window and identify features related to survival time based on the characteristics. The features related to survival time screened by the statistical test method can improve the performance of the model.

[0028] Use Cox regression analysis to screen for features related to survival time that are prognostic-related in a multivariate setting as initial features, specifically including: Use Cox regression analysis to analyze each feature related to survival time individually and calculate its probability value; Set a threshold for the probability value, screen out features with probability values less than the threshold based on the calculation results, and put the screened features into a multivariate Cox regression model; Combine the effects of multiple features on prognosis, use stepwise regression to optimize the model, and remove features that have insignificant effects on prognosis in a multivariate setting to obtain features related to prognosis in a multivariate setting.

[0029] Using Cox regression analysis to screen for features related to prognosis in a multivariate setting can simultaneously consider the effects of multiple covariates on survival time, thus more accurately evaluating the independent impact of each feature on prognosis. Cox regression can also handle censored data, so as to make full use of existing information for analysis, with high flexibility and more accurate data analysis. Use the random forest method to identify important prognostic features among the initial features, specifically including: Select n sample data from the initial features as a training set and generate a decision tree; During the process of generating a decision tree, features are randomly selected for node splitting to find the best splitting feature. The above process is repeated until a specified number of decision trees are generated, and these decision trees together form a random forest; Calculate the purity improvement value brought by each prognosis-related feature in the node splitting of the decision tree, and use the purity improvement value to evaluate the importance of the prognosis-related features; Average the impurity reduction values of each feature in all decision trees to obtain the importance evaluation value of the feature; According to the importance evaluation value of the feature, sort the features, and select the top-ranked features as important prognosis features. The important prognosis features are features that are significantly related to the patient's survival period or treatment response.

[0030] Use the random forest method to identify important prognosis features. The random forest algorithm can effectively handle high-dimensional data and large-scale data sets. By evaluating the importance of features, variables that have a significant impact on the classification result can be reasonably screened. By constructing multiple decision trees and synthesizing their results, the risk of overfitting is significantly reduced, the generalization ability of the model and the stable identification ability of important features are improved, and the accuracy of the prognosis model prediction is guaranteed.

[0031] Construct a prognosis model using important prognosis features, and use K-fold cross-validation to evaluate the stability and generalization ability of the prognosis model, specifically including: Construct a prognosis model using the selected important prognosis features and set the initial model parameters. The initial model parameters include the learning rate, regularization parameter, etc. Randomly divide the original data set into K non-overlapping subsets, and the number of samples in each subset is roughly equal; For each K from 1 to K, use the Kth subset as the validation set and the remaining K - 1 subsets as the training set; Train the prognosis model on the training set and evaluate the performance of the model on the validation set, and record the validation metrics, such as accuracy, recall rate, F1 score, etc.; During the process of model training and validation, continuously optimize and adjust the model parameters. When the evaluation metrics in the training results are within the qualified range, no model tuning is performed. When the evaluation metrics in the training results are not within the qualified range, model tuning is performed; Repeat the above steps K times to ensure that each subset has the opportunity to be used as the validation set and the remaining subsets are used as the training set; Collect the results of all K validations and calculate the average value of all K validation results, which is the estimated value of the prognosis model performance. Use the estimated value to evaluate the stability and generalization ability of the prognosis model.

[0032] The stability and generalization ability of the prognostic model are evaluated using K-fold cross-validation, which allows each sample to have the opportunity to be trained and validated, thereby improving the utilization rate of the data. During the validation process, since each sample is used as the validation set, it helps to reduce overfitting and makes the performance evaluation of the prognostic model more representative. Through multiple trainings and validations, K-fold cross-validation generates K evaluation metrics, which can more accurately evaluate the generalization ability of the model on unseen data and increase the stability and reliability of the evaluation results.

[0033] The predictive ability of the prognostic model is verified using an independent test set, specifically including: An independent dataset related to the original dataset is obtained as the test set, and the obtained test set is input into the evaluated prognostic model to generate prediction results; The generated prediction results are compared with the actual results, and the predictive performance of the prognostic model is obtained based on the comparison results.

[0034] Using an independent test set to verify the predictive ability of the prognostic model, the test set is completely separated from the dataset used for model training, so it is not affected by the training data, ensuring the objectivity of the validation process and avoiding evaluation biases caused by data overlap. Using an independent test set to verify the predictive ability of the prognostic model can also systematically evaluate the generalization performance of the model.

[0035] In summary, for the method for constructing a prognostic model for non-small cell lung cancer of the present invention, Cox regression analysis is used to screen for features related to prognosis in a multi-variable environment, and the effects of multiple covariates on survival time can be considered simultaneously, so as to more accurately evaluate the independent effect of each feature on prognosis. Cox regression can also handle deleted and missing data, so as to make full use of the existing information for analysis, with high flexibility and more accurate data analysis. The random forest method is used to identify important prognostic features. The random forest algorithm can effectively process high-dimensional data and large-scale data sets. By evaluating the importance of features, variables that have a significant impact on the classification result are reasonably screened out. By constructing multiple decision trees and synthesizing their results, the risk of overfitting is significantly reduced, and the generalization ability of the model and the ability to stably identify important features are improved, ensuring the accuracy of the prognostic model prediction. K-fold cross-validation is used to evaluate the stability and generalization ability of the prognostic model, enabling each sample to have the opportunity to be trained and verified, thereby improving the utilization rate of data. During the verification process, since each sample will be used as the validation set, it helps to reduce overfitting and makes the performance evaluation of the prognostic model more representative. K-fold cross-validation generates K evaluation indicators through multiple trainings and validations, and can more accurately evaluate the generalization ability of the model on unseen data, increasing the stability and reliability of the evaluation results. An independent test set is used to verify the prediction ability of the prognostic model. The test set is completely separated from the data set used for model training, so it is not affected by the training data, ensuring the objectivity of the verification process and avoiding evaluation biases caused by data overlap. Using an independent test set to verify the prediction ability of the prognostic model can also systematically evaluate the generalization performance of the model.

[0036] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0037] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

Claims

1. A method for constructing a prognostic model for non-small cell lung cancer, characterized in that: The construction method comprises: Collect clinical information of patients with small cell lung cancer from the Cancer Genome Atlas database, optimize and adjust the preset clinical information retrieval frequency parameters based on the stability of data retrieval, and determine whether there are any abnormalities in data retrieval; Obtain the patient's genomics, transcriptomics, proteomics, and metabolomics data and preprocess them to obtain the patient's molecular data; The clinical data and molecular data were integrated into a unified dataset to obtain the original dataset, and the important prognostic features in the original dataset were extracted. The prognostic model was constructed using the important prognostic features, and the stability and generalization ability of the prognostic model were evaluated using K-fold cross validation. After the validation was completed, an independent test set was used to verify the predictive ability of the prognostic model and obtain a prognostic model for non-small cell lung cancer.

2. The method for constructing a prognostic model for non-small cell lung cancer according to claim 1, characterized in that: Clinical information of patients with small cell lung cancer was collected from the Cancer Genome Atlas database, including: Retrieving clinical information from the cancer genome atlas database according to a preset clinical information retrieval frequency; Real-time monitoring of the proportion of missing data values ​​corresponding to various data types contained in the clinical information; Compare the proportion of missing data values ​​corresponding to each data type in each clinical information retrieval process with the preset proportion threshold; Each time clinical information is retrieved, the data type with a missing value ratio exceeding the preset value is extracted as the target data type; Using the similarity value of the target data type in each of two adjacent clinical information retrieval processes; The similarity value of the target data type in each of two adjacent clinical information retrieval processes is used to determine whether there is an abnormality in the data retrieval.

3. The method for constructing a prognostic model for non-small cell lung cancer according to claim 2, characterized in that: Using the similarity value of the target data type in each of two adjacent clinical information retrieval processes to determine whether there is an abnormality in the data retrieval, specifically includes: Extract the similarity value of the target data type between every two adjacent clinical information retrieval processes; Compare the similarity values ​​of the target data types in each of two adjacent clinical information retrieval processes with a preset similarity threshold; When the similarity value of the target data type in each of two adjacent clinical information retrieval processes is not less than a preset similarity threshold, the first abnormal coefficient acquisition model is used to acquire the first abnormal coefficient; When the similarity value of the target data type in each two adjacent clinical information retrieval processes is lower than a preset similarity threshold, the weighted average value of the same target data type in each two adjacent clinical information retrieval processes is retrieved; The second abnormal coefficient is obtained by using the weighted average value of the same target data type in each of two adjacent clinical information retrieval processes combined with the second abnormal coefficient acquisition model; Compare the first abnormal coefficient and the second abnormal coefficient with their corresponding abnormal coefficient thresholds respectively; When the first abnormal coefficient and the second abnormal coefficient exceed their corresponding abnormal coefficient thresholds, it is determined that there is an abnormality in data retrieval, and a data retrieval abnormality alarm is issued.

4. The method for constructing a prognostic model for non-small cell lung cancer according to claim 1, characterized in that: The step of obtaining the patient's genomics, transcriptomics, proteomics and metabolomics data and preprocessing the data specifically includes: Obtain patients' genomic, transcriptomic, proteomic and metabolomic data from professional data omics databases, research institutions or hospitals; Use statistical methods to fill missing values ​​in genomics, transcriptomics, proteomics, and metabolomics data and complete missing values; Screen and sort genomics, transcriptomics, proteomics, and metabolomics data, and identify and process features related to survival time in genomics, transcriptomics, proteomics, and metabolomics data; Remove duplicate data values ​​in genomics, transcriptomics, proteomics, and metabolomics data, and retain the first record of the duplicate data when removing; Normalize genomics, transcriptomics, proteomics, and metabolomics data.

5. The method for constructing a prognostic model for non-small cell lung cancer according to claim 1, characterized in that: The extracting of important prognostic features from the original data set specifically includes: Statistical testing methods were used to screen features associated with survival time from the original data set; Cox regression analysis was used to screen the features associated with survival time and prognosis in a multivariate environment as initial features; Random forest method was used to identify important prognostic features among the initial features.

6. The method for constructing a prognostic model for non-small cell lung cancer according to claim 5, characterized in that: The statistical test method is used to screen the features related to survival time from the original data set, specifically including: Set upper and lower thresholds of features related to survival time, identify features related to survival time in the original data set, and remove data points that exceed the thresholds; Establish a statistical model for the data and check whether the data points meet the predictions of the statistical model. Data points that do not meet the predictions are features related to survival time. Divide the data into different data groups, check for isolated points that are significantly different from other data points, and for time series data, use a sliding window to calculate the statistical properties of each data point and identify features related to survival time based on the properties.

7. The method for constructing a prognostic model for non-small cell lung cancer according to claim 5, characterized in that: The Cox regression analysis is used to screen the features related to survival time and the features related to prognosis in a multivariate environment as initial features, specifically including: Cox regression analysis was used to analyze each feature associated with survival time separately and calculate its probability value; Set a threshold for the probability value, filter out features with probability values ​​less than the threshold based on the calculation results, and put the filtered features into the multi-factor Cox regression model; Combining the effects of multiple features on prognosis, the model was optimized using the stepwise regression method, and the features that had no significant effect on prognosis in a multivariate environment were eliminated to obtain the features that were related to prognosis in a multivariate environment.

8. The method for constructing a prognostic model for non-small cell lung cancer according to claim 5, characterized in that: The random forest method is used to identify important prognostic features in the initial features, specifically including: Select n sample data from the initial features as a training set and generate a decision tree; In the process of generating decision trees, randomly select features for node splitting, find the best splitting features, and repeat the above process until the specified number of decision trees are generated; Calculate the purity improvement value brought by each prognosis-related feature in the node splitting in the decision tree, and use the purity improvement value to evaluate the importance of the prognosis-related features; The importance assessment value of the feature can be obtained by averaging the impurity reduction values ​​of each feature in all decision trees; According to the importance evaluation value of the features, the features are ranked and the top-ranked features are selected as important prognostic features.

9. The method for constructing a prognostic model for non-small cell lung cancer according to claim 1, characterized in that: The method of constructing a prognostic model using important prognostic features and evaluating the stability and generalization ability of the prognostic model using K-fold cross validation specifically includes: The prognostic model was constructed using the screened important prognostic features, and the initial model parameters were set; The original data set is randomly divided into K non-overlapping subsets, and the number of samples in each subset is roughly equal; For each K=1 to K, the Kth subset is used as the validation set and the remaining K-1 subsets are used as the training set; Train the prognostic model on the training set, evaluate the performance of the model on the validation set, and record the validation indicators; Repeat the above steps K times to ensure that each subset has a chance to be used as a validation set, and the remaining subsets are used as training sets; The results of all K validations were collected, and the average of all K validation results was calculated, which was the estimated value of the prognostic model performance. The estimated value was used to evaluate the stability and generalization ability of the prognostic model.

10. The method for constructing a prognostic model for non-small cell lung cancer according to claim 1, characterized in that: The use of an independent test set to verify the predictive ability of the prognostic model specifically includes: An independent data set related to the original data set is obtained as a test set, the obtained test set is input into the evaluated prognostic model, and a prediction result is generated; The generated prediction results are compared with the actual results, and the prediction performance of the prognostic model is obtained based on the comparison results.

Citation Information

Patent Citations

  • Non-small cell lung cancer prognosis model and construction method thereof

    CN117038086A

  • Non-small cell lung cancer integrated prognosis prediction model and construction method, device and application thereof

    CN113223727A

  • Non-small cell lung cancer patient prognosis survival rate prediction method based on multiple omics

    CN117809838A

  • Postoperative recurrence risk prediction system for I-stage lung adenocarcinoma patient and application of postoperative recurrence risk prediction system

    CN118553418A

  • Intelligent oxygen inhalation device for chronic obstructive pulmonary disease

    CN118846316A