Method for constructing lung cancer screening model based on UVP-TOF-MS

UVP-TOF-MS technology analyzes volatile organic compounds in the exhaled breath and builds an integrated learning model, solving the shortcomings of existing lung cancer screening methods, achieving non-invasive, rapid and accurate lung cancer screening, and having high sensitivity and specific early diagnosis capabilities.

CN118983079BActive Publication Date: 2025-07-22WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411120870.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-07-22
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

The existing lung cancer screening methods have high false positive rates, strong invasiveness, high cost and high radiation risk, making it difficult to achieve early diagnosis, and lack non-invasive, fast and accurate screening methods.

Method used

UVP-TOF-MS technology was used to analyze volatile organic compounds in the exhaled air. By constructing an integrated learning model, potential lung cancer markers were screened out, lung cancer screening model was established, and model performance was evaluated using confusion matrix and ROC curve.

Benefits of technology

It has achieved an effective distinction between lung cancer patients and non-lung cancer patients, has high sensitivity and specificity, and provides a non-invasive, fast and accurate lung cancer screening method, which has potential clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118983079B_ABST
    Figure CN118983079B_ABST
Patent Text Reader

Abstract

Method for constructing lung cancer screening model based on UVP-TOF-MS, comprising the steps of: A. Collecting exhaled breath samples; B. Performing full-spectrum analysis on the collected exhaled breath samples through a UVP-TOF-MS device to form spectral map samples; C. Data preprocessing: including performing various conventional data preprocessing and related calculations on the obtained spectral map samples, and selecting suitable features; D. Constructing a model: constructing an ensemble learning model, ranking the gain importance of each feature by a base classifier to form a feature set of the ensemble learning model; jointly forming a comprehensive lung cancer screening prediction model with a logistic regression model and the ensemble learning model; E. Model performance evaluation: predicting the performance of the lung cancer screening prediction model through a confusion matrix, and then screening out the best-performing lung cancer screening prediction model. Most of the features selected by the present invention have significant differences and can be used as potential lung cancer markers, which is of positive significance for lung cancer screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of model construction for medical use, and specifically relates to a method for constructing a lung cancer screening model based on UVP-TOF-MS. Background Art

[0002] Lung cancer is a major social health problem, causing 16.8% of global deaths. Data released by the World Health Organization (WHO) in 2022 showed that lung cancer (Lc) is the most common cancer in men and the second most common cancer in women, and the mortality rate caused by lung cancer ranks first.

[0003] The mortality rate of lung cancer is high. The five-year survival rate of lung cancer patients in most countries is only 10%-20% after diagnosis, and there has been no significant improvement in the five-year survival rate in the past two decades. The main reason for the high fatality rate is that the early onset of lung cancer is hidden, and it is often in the advanced stage when diagnosed. The prognosis of lung cancer is closely related to its clinical stage. Achieving early diagnosis can improve the prognosis of patients. Therefore, it is very important to select an effective screening method or formulate a suitable screening plan.

[0004] Currently, many methods have been used for lung cancer screening. Commonly used methods in clinical practice include chest X-ray scan, chest CT, sputum culture test, percutaneous biopsy, etc. Currently, chest X-ray scan is not recommended for detecting lung cancer because it is difficult to show lung lesions and cannot reduce the mortality rate. According to the results of the National Lung Screening Trial in China, using low-dose computed tomography (LDCT) to screen for lung cancer is the most effective method to reduce the lung cancer mortality rate. The LDCT guided by the LungRADS algorithm of the American College of Radiology has been verified and performed well in the population of the National Lung Screening Trial in the United States. Data showed that the lung cancer mortality rate decreased by 20%. Therefore, LDCT is applicable to the lung cancer screening of high-risk groups. However, the false positive rate of LDCT is relatively high, which may lead to more misdiagnosis and unnecessary examinations and treatments, and there is also a risk of radiation exposure. Sputum cytology analysis is another lung cancer screening method, but this method usually requires multiple samplings, and its accuracy and sensitivity are relatively low, and it cannot be used as an effective means for routine lung cancer screening. Lung tissue biopsy is the gold standard for cancer diagnosis, and lung tissue can be extracted by means such as fiberoptic bronchoscopy or thoracentesis. However, such methods are invasive, not suitable for multiple examinations, prone to complications and expensive. When using the above traditional diagnostic methods for screening, about 85% of patients are diagnosed in the advanced stage. Therefore, there is an urgent need for a more convenient, accurate, cheap and fast method for early screening of lung cancer in clinical practice.

[0005] Biomarker detection is a new type of lung cancer screening method, including plasma biomarker liquid biopsy and the detection of volatile organic compounds (VOCs) in exhaled breath, etc. VOCs are volatile organic compounds with a melting point below room temperature and a boiling point between 50 - 260 °C. Most of the gases exhaled by the human body are VOCs. When pathological changes occur in the human body, the composition of VOCs may also change, and some are excreted from the lungs through the respiratory tract. The VOCs composition of each disease will change, so specific diseases can be screened out by detecting the VOCs composition. Some studies have also pointed out that VOCs can be used to diagnose lung cancer. Therefore, the diagnostic technology based on VOCs analysis is considered a promising non-invasive early lung cancer screening method, which has the advantages of being fast, non-invasive, highly sensitive, and having good repeatability compared with traditional detection technologies. It is a hot topic in the research of early lung cancer diagnosis in recent years. However, this technology lacks a unified standard, so it has not been widely popularized in clinical practice yet.

[0006] Ultraviolet Photoionization-Mass spectrometry (UVP-MS) is an online mass spectrometry technology commonly used to detect volatile organic compounds in recent years. Ultraviolet Photoionization Time-of-flight Mass spectrometry (UVP-TOF-MS) is a relatively novel detection technology with fast detection speed and high sensitivity, and is less affected by sample humidity compared with other mass spectrometry methods. UVP-TOF-MS can obtain analysis results immediately, skipping laboratory testing and greatly reducing waiting time. Therefore, real-time monitoring of the target can be carried out through online analysis. Therefore, applying UVP-TOF-MS to the exhaled breath detection of patients can theoretically become a potential lung cancer screening method. Summary of the Invention

[0007] The present invention provides a method for constructing a lung cancer screening model based on UVP-TOF-MS, which analyzes the volatile organic compounds in exhaled breath by constructing a corresponding lung cancer screening model to find potential exhaled breath lung cancer biomarkers.

[0008] The method for constructing a lung cancer screening model based on UVP-TOF-MS of the present invention is characterized in that it includes the steps of:

[0009] A. Exhaled breath sample collection: Collect exhaled breath samples of lung cancer patients diagnosed within a set time range, normal exhaled breath samples, and nodule exhaled breath samples. Define normal exhaled breath samples and nodule exhaled breath samples as non-cancer exhaled breath samples;

[0010] B. Perform full-spectrum analysis on each collected exhaled breath sample using a UVP-TOF-MS device to obtain spectral data for each exhaled breath sample and form a spectral sample;

[0011] C. Data preprocessing: Include data cleaning, mean filling of missing values, outlier deletion, data correction of the spectral sample to the same horizontal distribution through calibration gas, environmental background deduction, and peak area calculation for the obtained spectral sample. Divide the spectral sample into a training set, a validation set, and a test set, and select suitable features; Among them,

[0012] When correcting the data of the spectral sample through calibration gas, it includes calculating the peak area of the calibration gas with a mass-to-charge ratio in the range of [77.7, 78.4] per day, and setting the value of the mass-to-charge ratio intensity lower than 300 to 0; Then calculate the calibration gas area of mass-to-charge ratio_78: Set a standard value, and use the quotient of the standard value divided by the area of the calibration gas with mass-to-charge ratio_78 per day as the coefficient value. Finally, multiply the area of each mass-to-charge ratio of each exhaled breath sample by the coefficient value at the corresponding time to obtain the area data of each mass-to-charge ratio after correction of each exhaled breath sample. Select all features in the range of mass-to-charge ratio_15 to mass-to-charge ratio_249, and delete features with a mass-to-charge ratio of 94 and a proportion of values of 0 greater than 90%;

[0013] D. Model construction: Construct an ensemble learning model, and finely adjust the hyperparameters of the base classifier of the ensemble learning model through Bayesian optimization technology to obtain the best combination of model parameters; The base classifier of the ensemble learning model ranks according to the information gain importance of each feature, selects the top N most important features with information gain importance to form the feature set of the ensemble learning model, where N is a preset natural number; Apply the features in the feature set to a logistic regression model, and the logistic regression model and the ensemble learning model together form a comprehensive lung cancer screening prediction model;

[0014] E. Model performance evaluation: The lung cancer screening prediction model is used to predict the test set through a confusion matrix, and the accuracy, sensitivity, and specificity indicators of the lung cancer screening prediction model are calculated to quantify the prediction ability of the lung cancer screening prediction model. Then, by plotting the Receiver Operating Characteristic (ROC) curve of the lung cancer screening prediction model and calculating the Area Under the Curve (AUC), the diagnostic performance of the lung cancer screening prediction model under different parameters is compared. Based on the analysis of the area under the curve, the best-performing lung cancer screening prediction model is selected.

[0015] The present invention is based on Ultraviolet Photoionization Time-of-flight Mass spectrometry (UVP-TOF-MS) to analyze volatile organic compounds (VOCs) in exhaled breath for lung cancer screening. By collecting and analyzing the VOCs of subjects (including non-lung cancer patients and diagnosed lung cancer patients) through UVP-TOF-MS, methods such as confusion matrix and ROC curve are used to evaluate the data set. Data analysis is performed on several different groups of people, such as non-lung cancer and lung cancer, normal people and lung cancer nodules and lung cancer, using an ensemble learning model to find potential exhaled breath lung cancer markers.

[0016] Further, in step B, the step of performing full-spectrum analysis on each collected exhaled breath sample by the UVP-TOF-MS device is as follows:

[0017] After the UVP-TOF-MS device is preheated and the state is stable, the collected exhaled breath samples are connected to the injection pipeline one by one. In the full-spectrum analysis of the exhaled breath samples, two spectral data points are collected for each exhaled breath sample, and the second spectral data of each exhaled breath sample is taken as the spectral sample of the exhaled breath sample.

[0018] Further, in step C, the peak areas under the curves of all qualitative VOCs in the spectral samples are normalized to the interval (0, 1), and a data matrix is generated through data processing. The spectral samples are divided into a training set, a validation set, and a test set at a set ratio, and then the data is corrected.

[0019] Further, when correcting the data, calculate the peak area and the standard gas area of the calibration gas data of each day within the specified mass-to-charge ratio range, and then obtain the area data of each mass-to-charge ratio of each exhaled breath sample after correction. Then, select all the features within the set mass-to-charge ratio area range.

[0020] The beneficial effects of the present invention include:

[0021] 1. Most of the different numbers of features selected by the model have significant differences and can be used as potential lung cancer markers, which is of positive significance for lung cancer screening.

[0022] 2. It can well distinguish between two groups of people: lung cancer patients and non-lung cancer patients. Description of the Drawings

[0023] Figure 1 This is a flowchart of the method for constructing a lung cancer screening model based on UVP-TOF-MS of the present invention.

[0024] Figure 2 This is a graph of the percentage contribution of the top 30 features in Example 1.

[0025] Figure 3 a is a prediction graph of the confusion matrix for the validation set samples in Example 1.

[0026] Figure 3 b is an ROC curve graph of the validation set samples in Example 1.

[0027] Figure 4 a is a prediction graph of the confusion matrix for the test set samples in Example 1.

[0028] Figure 4 b is an ROC curve graph of the test set samples in Example 1.

[0029] Figure 5 This is a graph of 28 features with significant differences in Example 1.

[0030] Figure 6 This is a graph of the percentage contribution of the top 30 features in Example 2.

[0031] Figure 7 a is a prediction graph of the confusion matrix for the validation set samples in Example 2.

[0032] Figure 7 b is an ROC curve graph of the validation set samples in Example 2.

[0033] Figure 8 a is a prediction graph of the confusion matrix for the test set samples in Example 2.

[0034] Figure 8 b is an ROC curve graph of the test set samples in Example 2.

[0035] Figure 9 This is a graph of 22 features with significant differences in Example 2.

[0036] Figure 10 This is a graph of the percentage contribution of the top 30 features in Example 3.

[0037] Figure 11 a is the prediction graph of the confusion matrix for the validation set samples in Example 3.

[0038] Figure 11 b is the ROC curve graph of the validation set samples in Example 3.

[0039] Figure 12 a is the prediction graph of the confusion matrix for the test set samples in Example 3.

[0040] Figure 12 b is the ROC curve graph of the test set samples in Example 3.

[0041] Figure 13 are 28 feature graphs with significant differences in Example 3. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0043] As Figure 1 shown, the method for constructing a lung cancer screening model based on UVP-TOF-MS of the present invention is characterized by including the steps of:

[0044] A. Collection of exhaled breath samples: Collect exhaled breath samples of patients diagnosed with lung cancer, normal exhaled breath samples, and nodular exhaled breath samples within a set time range, and define the normal exhaled breath samples and nodular exhaled breath samples as non-cancer exhaled breath samples.

[0045] In this embodiment, the exhaled breath samples were collected from patients who visited West China Hospital of Sichuan University from October 2022 to April 2023 and had space-occupying lesions in the lungs shown by chest CT as the research objects, and healthy subjects were recruited as the control group at the same time.

[0046] Inclusion criteria: (1) Aged 18 - 70 years old, regardless of gender; (2) Complete clinical information, including inpatient / outpatient number, age, gender, clinical diagnosis, etc.; (3) Meeting one of the following conditions: a) Subjects suspected of having lung cancer; b) Subjects with other benign respiratory diseases; c) Patients with other cancers, such as gastric cancer, esophageal cancer, colorectal cancer; d) Healthy subjects without respiratory diseases and signs. When a subject meets all of the above criteria, they can be included.

[0047] Exclusion criteria: (1) Patients previously diagnosed with lung cancer and receiving lung cancer treatment; (2) Patients who ate after 22:00 the previous night; (3) Patients with ventilation or gas exchange disorders; (4) Patients with severe mental illness or cognitive impairment; (5) Females in pregnancy or lactation; (6) Patients with a history of drug use or who are currently using drugs; (7) Patients whom the researcher deems unsuitable to participate in this study. When a subject meets any of the above criteria, they cannot be included.

[0048] This study has established the withdrawal criteria and procedures for subjects to protect the rights and interests of the subjects. The subject has the right to withdraw from the trial at any time without any reason. If the subject decides to withdraw from the trial, the researcher should try to find out the reason for withdrawal as much as possible and record this reason as the original recorded text data. To protect the rights and interests of the subjects, ensure the quality of the trial, and avoid unnecessary economic losses, if a subject has the following situations and the researcher determines that the subject cannot continue with this clinical trial, the researcher has the right to suspend the subject's participation in the trial and require the subject to withdraw from the trial in advance: (1) The subject meets the exclusion criteria (newly emerged or previously undetected) and cannot continue to participate in the study; (2) The subject has other situations that the researcher deems unable to continue with this clinical trial; (3) The subject is unwilling to continue to participate in this clinical trial for any reason.

[0049] In this example, a total of 4325 cases of lung cancer samples and normal samples were collected from West China Hospital of Sichuan University, including 360 cases of lung cancer (abbreviated as Lc) samples, 995 cases of normal samples, and 2970 cases of node (Node) samples. The normal samples and node samples were defined as non-lung cancer (abbreviated as Non-Lc) samples, with a total of 3965 cases.

[0050] In this study, the clinical data of patients were mainly collected from the HIS medical record system (Hospital Information System, HIS), and the clinical data of the health examination population were collected from the physical examination system. It mainly includes the basic information of patients and examination result data: gender, age, exposure history of risk factors such as smoking history, imaging examination, pathological report (histological type, pathological stage), and tumor location and size, laboratory examination (including: blood routine, biochemistry, tumor markers, etc.).

[0051] To ensure that the quality of exhaled breath collection meets the requirements, the collected exhaled breath is collected through an exhaled breath sample collection bag, and the following conditions should be met: 1) Collect exhaled breath before treatments such as surgery, chemotherapy, and radiotherapy; 2) Fast on an empty stomach for more than 8 hours in the morning; 3) Do not smoke or drink alcohol within 2 hours; 4) Use an offline sampling device to collect alveolar gas samples from the subjects. A strict standard process for collecting and detecting and analyzing exhaled breath samples of subjects has been formulated for the processes of collecting exhaled breath samples, collecting environmental air, and transporting exhaled breath sample collection bags to ensure the accuracy and scientific nature of the experiment.

[0052] B. Perform a full-spectrum analysis on each collected exhaled breath sample through a UVP-TOF-MS device to obtain the spectral data of each exhaled breath sample and form a spectral sample:

[0053] After sending the exhaled breath sample collection bag containing the exhaled breath sample to the laboratory according to the standard process, the exhaled breath sample can be detected and analyzed. Mass spectrometry conditions and parameter settings of the UVP-TOF-MS device: Device model: UVP-TOF-MSPLUS, inlet tube and ion source temperature 70 °C, ion source gas pressure 500 Pa, ion source current 1 mA, integration time 50 s; Test analysis experimental process of exhaled breath samples: (1) After the UVP-TOF-MS device is preheated and in a stable state, connect the exhaled breath sample collection bag to the inlet pipeline, open the air bag valve, and start the full-spectrum analysis of the exhaled breath sample. Two spectral data points are collected for each exhaled breath sample, and the second spectral data of each exhaled breath sample is taken as the result; (2) After the spectral collection is completed, close the air bag valve, open the nitrogen cleaning pipeline controlled by the Mass Flow Controller (MFC), and execute the pipeline cleaning process for the above two data points to discharge the gas residues in the pipeline through clean high-purity nitrogen; (3) After the pipeline cleaning process is completed, the above steps can be repeated to start the injection analysis process of the next exhaled breath sample.

[0054] C. Data preprocessing: including data cleaning, filling missing values with the mean value, deleting outliers, calibrating the data of the spectral sample to the same horizontal distribution through calibration gas, subtracting the environmental background, and calculating the peak area:

[0055] Normalize the peak areas (VOCs content) of all qualitative VOCs (volatile organic compounds) in the spectral sample to the interval (0, 1). After data processing, a data matrix is generated. The spectral sample is divided into a training set, a validation set, and a test set in a ratio of 3:1:1, and then the data is calibrated.

[0056] When calibrating the data, calculate the peak area and the standard gas area of the calibration gas data for each day where the M / Z (mass-to-charge ratio) is in the range of [77.7, 78.4]. Set the M / Z values with intensity values lower than 300 to 0, and calculate the standard gas area of M / Z_78: Take 3,000,000 as the standard value, retain the quotient of dividing the standard value by the area of the daily standard gas M / Z_78 as the coefficient value. Finally, multiply the area of each M / Z of each exhaled breath sample by the coefficient value at the corresponding time to obtain the area data of each M / Z after correcting the exhaled breath sample. Select all features in the range of M / Z_15 to M / Z_249, and a total of 235 features are selected. Finally, delete the contaminated M / Z_94 existing in the sampling bag of the exhaled breath sample itself and the features with a proportion of 0 greater than 90%.

[0057] D. Build a model: Build an ensemble learning model, which can include models such as Random Forest (RF) model, CatBoost model, or XGBoost model. Perform precise modeling through the training set and the validation set to ensure the reliability and accuracy of the model. During the training process of the model, use the training set to train the base classifiers of the model and use the validation set to evaluate their performance. Then, apply Bayesian optimization techniques to finely tune the hyperparameters (such as the learning rate and the maximum depth of the tree) of these base classifiers to ensure the selection of the best combination of model parameters.

[0058] To further improve the efficiency and interpretability of the model, the information gain importance scores of each feature will be based on the base classifiers. By training the base classifiers and collecting these scores, a feature importance curve is drawn. This curve arranges the features in descending order of their importance and shows their cumulative contributions. Then analyze this curve to identify those features whose contribution to the improvement of the model performance decreases, which is usually reflected in a significant inflection point of the curve. Based on this inflection point, select the top N most important features to form the feature set of the model, where N is a preset natural number.

[0059] Apply the features in the feature set to the logistic regression model. The logistic regression model is a statistical learning method used to solve classification problems. The training objective of the logistic regression model is to maximize the likelihood function, and usually optimization algorithms such as gradient descent are used to solve the parameters. Once the model training is completed, it can be used to classify new samples and output the corresponding class probabilities. The logistic regression model and the ensemble learning model together form a comprehensive lung cancer screening prediction model. By analyzing the weights of the independent variables in the logistic regression model, key factors closely related to the occurrence of lung cancer can be identified. These factors and their weights will become the core indicators for predicting the individual lung cancer risk, thus enabling a more accurate prediction of the likelihood of lung cancer occurrence.

[0060] E. Model Performance Evaluation: The lung cancer screening prediction model is used to predict the test set through a confusion matrix, and the accuracy, sensitivity, and specificity indicators of the lung cancer screening prediction model are calculated to quantify the prediction ability of the lung cancer screening prediction model. Then, by plotting the Receiver Operating Characteristic (ROC) curve of the lung cancer screening prediction model and calculating the Area Under the Curve (AUC), the diagnostic performance of the lung cancer screening prediction model under different parameters is compared. Based on the analysis of the area under the curve, the best-performing lung cancer screening prediction model is selected. The closer the AUC value is to 1, the higher the diagnostic accuracy of the model. Based on the analysis of the AUC value, the best-performing diagnostic model can be selected.

[0061] The confusion matrix, also known as the error matrix or likelihood matrix, is a standard format for representing accuracy evaluation in the field of machine learning. The confusion matrix is a visualization tool, especially for supervised learning. In structured accuracy evaluation, it is mainly used to compare the classification results with the actual measured values, and the accuracy of the classification results can be displayed in a confusion matrix.

[0062] The meaning expressed by the confusion matrix: Each column of the confusion matrix represents the predicted category, and the total number of each column represents the number of data predicted as this category; each row represents the true belonging category of the data, and the total number of data in each row represents the number of data instances of this category; the values in each column represent the number of true data predicted as this category. The structure of the confusion matrix is generally shown in Table 1.

[0063] Table 1:

[0064]

[0065] In Table 1, TP represents True Positive, that is, the number of samples that are truly positive and are correctly predicted as positive; FN represents False Negative, that is, the number of samples that are truly positive but are wrongly predicted as negative; FP represents False Positive, that is, the number of samples that are truly negative but are wrongly predicted as positive; TN represents True Negative, that is, the number of samples that are truly negative and are correctly predicted as negative.

[0066] Sensitivity = True Positive / (True Positive + False Negative) × 100% = TP / (TP + FN) × 100%;

[0067] Specificity = True Negative / (False Positive + True Negative) × 100% = TN / (FP + TN) × 100%;

[0068] Accuracy rate = (True positive + True negative) / (True positive + False negative + False positive + True negative) × 100% = (TP + TN) / (TP + FP + FN + TN) × 100%.

[0069] These metrics together constitute a comprehensive evaluation of the performance of the classification model. The accuracy rate gives the overall proportion of correct predictions by the model, while sensitivity and specificity measure the model's ability to identify true positives and true negatives respectively.

[0070] AUC (Area Under the Curve) is the area under the ROC (Receiver Operating Characteristic) curve, usually ranging between 0.5 and 1. This metric is used to intuitively evaluate the performance of a classifier, where the value closer to 1 indicates better performance of the classifier. AUC is an important performance evaluation criterion because, compared to the ROC curve itself, it provides a more explicit comparison of classifier performance in the form of a single value.

[0071] The calculation of AUC usually adopts the trapezoidal rule, that is, calculating the total area under the ROC curve. The value of AUC is used to judge the quality of the classifier, and the specific explanations are as follows:

[0072] AUC = 1: Indicates that the classifier is perfect and can make perfect predictions by setting thresholds.

[0073] 0.5 < AUC < 1: Indicates that the performance of the classifier is better than random guessing and has practical predictive value.

[0074] AUC = 0.5: Indicates that the prediction effect of the classifier is the same as random guessing, that is, it has no practical predictive value.

[0075] AUC < 0.5: Indicates that the performance of the classifier is worse than random guessing, but if its prediction results can be interpreted in reverse, it may be better than random guessing.

[0076] This method makes AUC a concise and effective tool for measuring and comparing the performance of different classifiers when dealing with imbalanced datasets.

[0077] The present invention is based on Ultraviolet Photoionization Time-of-flight Mass spectrometry (UVP-TOF-MS) to analyze volatile organic compounds (VOCs) in exhaled breath for lung cancer screening. By collecting and analyzing the VOCs of subjects (including non-lung cancer patients and confirmed lung cancer patients) through UVP-TOF-MS, methods such as confusion matrix and ROC curve are used to evaluate the dataset. For several different paired populations such as non-lung cancer and lung cancer, normal people and lung cancer nodules and lung cancer, an integrated learning model is used for data analysis to find potential exhaled breath lung cancer markers.

[0078] The following further illustrates the present invention through multiple clinical data:

[0079] Example 1:

[0080] The prediction model based on UVP-TOF-MS can effectively distinguish lung cancer patients from non-lung cancer patients (healthy people and nodule patients), as follows:

[0081] The collected lung cancer samples and non-lung cancer samples (a total of 877 cases) are divided into a training set, a validation set, and a test set according to a ratio of 6:2:2: 526 cases in the training set; 176 cases in the validation set, including 90 lung cancer samples and 86 non-lung cancer samples; 175 cases in the test set, including 81 lung cancer samples and 94 non-lung cancer samples. An XGBoost (eXtreme Gradient Boosting) algorithm is used to establish a binary classification model.

[0082] The top 30 sorted results of the feature importance of the model are as Figure 2 shown in the Feature contribution percentage chart. It can be seen from Figure 2 that the feature M / Z_105_area has the highest percentage contribution to the model, the feature M / Z_167_area has the second highest percentage contribution to the model, the feature M / Z_96_area has the third highest percentage contribution to the model, and so on for other features, and the feature M / Z_92_area has the lowest percentage contribution to the model.

[0083] The feature contribution percentage graph is a visualization method that shows the impact of each feature in the model on the prediction result. Features are labeled as M / Z plus a series of numbers. The importance of each feature is represented as a percentage and presented in the graph in terms of length. The longer the bar, the greater the impact of the feature on the prediction result of the model. The longest bar represents the most important feature, while the shortest represents the feature with less impact.

[0084] Use the trained model to predict the samples in the validation set, and through the confusion matrix ( Figure 3 a) and the ROC curve graph ( Figure 3 b) for evaluating the diagnostic performance, the prediction of the samples in the validation set is as Figure 3 shown in a and Figure 3 b.

[0085] Among the Figure 3 90 lung cancer samples (Lc in the vertical True value) in a, 4 are predicted as non-lung cancer samples (Non_lc in the horizontal Predicted value), and among the 86 non-lung cancer samples (Non_lc in the vertical True value), 1 is predicted as a lung cancer sample (Lc in the horizontal Predicted value). The overall accuracy is 97.16%, the sensitivity is 95.56%, and the specificity is 98.84%; Figure 3 In b, the AUC value of the ROC area is 0.9938, which is very close to 1, indicating that the modeling effect is very excellent.

[0086] Use the trained model to predict the samples in the test set to evaluate the model performance, and through the confusion matrix ( Figure 4 a) and the ROC curve graph ( Figure 4 b) for the prediction of the samples in the test set, as Figure 4 shown in a and Figure 4 b.

[0087] Among the Figure 4 81 lung cancer samples (Lc in the vertical True value) in a, 1 is predicted as a non-lung cancer sample (Non_lc in the horizontal Predicted value), and among the 94 non-lung cancer samples (Non_lc in the vertical True value), 5 are predicted as lung cancer samples (Lc in the horizontal Predicted value). The overall accuracy is 96.57%, the sensitivity is 98.77%, and the specificity is 94.68%; Figure 4 In b, the AUC value of the ROC area is 0.9971, which is very close to 1, indicating that the modeling effect is very excellent.

[0088] For non-lung cancer patients and cancer patients, the non-parametric Mann-Whitney U test was performed on the top 30 features before selecting the XGBoost model. There were a total of 25 features with a significant P-value less than 0.05. Box plots of these 25 features for non-lung cancer and lung cancer patients and the P-values of the non-parametric Mann-Whitney U test were shown on the figure titles, as Figure 5 shown.

[0089] It can be Figure 5 seen that the distribution differences of 25 specific mass-to-charge ratio (M / Z) features between the "non-lung cancer" (Nor_lc) and "lung cancer" (Lc) groups. The P-values of each feature indicated its statistical significance and could well distinguish these two groups of people. The median of most features was higher in the lung cancer group, indicating that the expression levels of these features were higher in lung cancer. Moreover, the degree of dispersion of the features showed certain differences between the two groups, especially the relatively consistent expression in the lung cancer group might point to specific biological processes. For example, the features of M / Z_121_area and M / Z_135_area showed obvious significant differences between non-lung cancer and lung cancer patients, which meant that they might be closely related to specific biological processes of lung cancer. In addition, the outliers of some features suggested the existence of possible subgroups, which might play an important role in future lung cancer subtype classification or individualized treatment.

[0090] Overall results: The accuracy rate of the training set was 100%, the accuracy rate of the validation set was 97.16%, and the accuracy rate of the test set was 96.57%; the prediction effects of the three data sets were all very good, and the correct rates were all greater than 95%, indicating that the generalization ability of the prediction model was very good. It was proved that the prediction model based on UVP-TOF-MS could effectively distinguish lung cancer patients from non-lung cancer patients (healthy people and nodule patients).

[0091] Example 2:

[0092] The prediction model based on UVP-TOF-MS could effectively distinguish lung cancer patients from healthy people, as follows:

[0093] The collected lung cancer samples and normal samples were divided into a training set, a validation set, and a test set at a ratio of 6:2:2: 525 samples in the training set; 176 samples in the validation set, including 78 lung cancer samples and 98 normal samples; 175 samples in the test set, including 91 lung cancer samples and 84 normal samples. A binary classification model was established using the XGBoost algorithm.

[0094] The top 30 ranking results of the feature importance of the model were as Figure 6As shown in the Feature contribution percentage chart, the Feature M / Z_105_area has the highest percentage value of contribution to the model, the Feature M / Z_227_area has the second highest percentage value of contribution to the model, the Feature M / Z_240_area has the third highest percentage value of contribution to the model, and so on for other features.

[0095] The trained model is used to predict the validation set samples, and the prediction of the validation set samples is shown in the confusion matrix ( Figure 7 a) and the ROC curve graph ( Figure 7 b) for evaluating the diagnostic performance, as shown in Figure 7 a and Figure 7 b.

[0096] As can be seen from Figure 7 a, among the 78 lung cancer samples (Lc in the vertical True value) in the validation set, 4 were predicted as normal samples (Normal in the horizontal Predicted value), and among the 98 normal samples (Normal in the vertical True value), 3 were predicted as lung cancer samples (Lc in the horizontal Predicted value). The overall accuracy rate is 96.02%, the Sensitivity is 94.87%, and the Specificity is 96.94%. Figure 7 The area under the ROC curve in b is 0.9898, and the AUC is greater than 0.90, indicating that the classifier has a good effect.

[0097] The trained model is used to predict the test set samples to evaluate the model performance, and the prediction of the test set samples is shown in the confusion matrix ( Figure 8 a) and the ROC curve graph ( Figure 8 b) for evaluating the diagnostic performance, as shown in Figure 8 a and Figure 8 b.

[0098] As can be seen from Figure 8 a, among the 91 lung cancer samples (Lc in the vertical True value) in the test set, 4 were predicted as normal samples (Normal in the horizontal Predicted value), and among the 84 normal samples (Normal in the vertical True value), 4 were predicted as lung cancer samples (Lc in the horizontal Predicted value). The overall accuracy rate is 95.43%, the Sensitivity is 95.60%, and the Specificity is 95.24%. Figure 8The area under the ROC curve in b is 0.9958, and the AUC is greater than 0.90, indicating that the classifier has a very good effect.

[0099] For normal people and cancer patients, the top 30 features of the XGBoost model were selected for the non-parametric Mann-Whitney U test. There were a total of 20 features with a significant P-value less than 0.05. Box plots of these 20 features for normal and lung cancer patients and the P-values of the non-parametric Mann-Whitney U test were shown on the graph titles, as Figure 9 shown.

[0100] In Figure 9 each box plot presents the distribution of a specific m / z feature between the normal population and lung cancer patients. By the position of the median, we can observe a significant increase or decrease of some features in the lung cancer group, which may indicate that they play an important biological role in the development of lung cancer. The width of the interquartile range shows the dispersion of the data within each group. A narrower interquartile range indicates that the feature values are more concentrated within the group. The P-value of each feature is also marked on the graph. A P-value below 0.05 usually indicates that this difference is statistically significant, and an extremely low P-value close to 0 more strongly indicates a significant difference between the two groups. For example, the features of m / z 105_area and m / z 167_area show extreme significant differences between the normal population and lung cancer patients, which means they may be closely related to the biomarkers of lung cancer. In addition, the outliers on the box plot reflect specific individual differences, which may be related to the heterogeneity of the disease or different lung cancer subtypes. In summary, this set of box plots provides important information about the distribution differences of m / z features between the normal population and lung cancer patients, which is of great significance for understanding the molecular mechanism of lung cancer, developing new diagnostic methods, and personalized treatment strategies. These features are expected to become key biomarkers in future lung cancer diagnosis and treatment.

[0101] Overall results: The accuracy of the training set is 100%, the accuracy of the validation set is 96.02%, and the accuracy of the test set is 95.43%; the prediction effects of the three datasets are all very good, and the correct rates are all greater than 95%, indicating that the model has good generalization ability. It is proved that the prediction model based on UVP-TOF-MS can effectively distinguish lung cancer patients from healthy people.

[0102] Example 3:

[0103] The prediction model based on UVP-TOF-MS can effectively distinguish lung cancer patients from patients with benign lung nodules, as follows:

[0104] The collected lung cancer samples and nodule samples (876 cases in total) were divided into a training set, a validation set, and a test set at a ratio of 6:2:2: 525 cases in the training set; 176 cases in the validation set, including 78 lung cancer samples and 98 nodule samples; 175 cases in the test set, including 91 lung cancer samples and 84 nodule samples. An XGBoost algorithm was used to establish a binary classification model.

[0105] The top 30 sorted results of the feature importance of the model are as Figure 10 shown in the Feature contribution percentage chart, and it can be seen from Figure 10 this that the feature M / Z_105_area has the highest contribution to the model, the feature M / Z_206_area has the second highest contribution to the model, the feature M / Z_179_area has the third highest contribution to the model, and so on for other features, and the feature M / Z_118_area has the lowest contribution to the model.

[0106] The trained model was used to predict the validation set samples, and the predictions of the validation set samples are shown in Figure 11 a) and the ROC curve graph ( Figure 11 b) used to evaluate the diagnostic performance as shown in Figure 11 a and Figure 11 b.

[0107] Figure 11 In Figure 11 a, among the 78 lung cancer samples (Lc in the vertical True value), 3 were predicted as nodule samples (Node in the horizontal Predicted value), and among the 98 nodule samples (Node in the vertical True value), 3 were predicted as lung cancer samples (Lc in the horizontal Predicted value). The overall accuracy was 96.59%, the sensitivity was 96.15%, and the specificity was 96.94%;

[0108] The trained model was used to predict the test set samples, and the predictions of the test set samples are shown in Figure 12 a) and the ROC curve graph ( Figure 12 b) used to evaluate the diagnostic performance as shown in Figure 12 a and Figure 12 b.

[0109] In Figure 12Among the 91 lung cancer samples in a (Lc vertically for True value), 2 were predicted as nodule samples (Node horizontally for Predicted value), and among the 84 nodule samples (Node vertically for True value), 1 was predicted as a lung cancer sample (Lc horizontally for Predicted value). The overall accuracy was 98.29%, the sensitivity was 97.80%, and the specificity was 98.81%; in Figure 12 In b, the AUC value of the ROC curve was 0.9966, very close to 1, indicating excellent modeling performance.

[0110] For nodule patients and cancer patients, the first 30 features of the XGBoost model were selected for the non-parametric Mann-Whitney U test, and the significant P value was less than 0.05. There were a total of 26 features. Box plots of nodule patients and lung cancer patients for these 26 features and the P values of the non-parametric Mann-Whitney U test were shown on the figure titles, as Figure 13 shown.

[0111] In Figure 13 each box plot presented reveals the distribution of a specific m / z feature in nodule and lung cancer samples. The medians of these features were significantly different between the two types of samples, which may indicate that certain specific mass spectrometry features are closely related to the development of lung cancer. For example, the expression levels of some features in the lung cancer group were significantly higher than those in the nodule group, which may reflect the expression patterns of lung cancer biomarkers. The presence of outliers indicates individual differences in feature expression in both populations, which may be related to different disease subtypes or biological behaviors. Especially in lung cancer samples, higher or lower outliers may be related to disease heterogeneity or patient biomarker specificity. Statistically, the P value of each feature was marked on the corresponding box plot. Features with P values below 0.05 indicate that the differences between nodule and lung cancer samples are statistically significant, meaning these features can be used to distinguish two different biological states. In summary, the m / z features shown in these box plots have important biological and clinical significance for distinguishing nodule and lung cancer samples. They may become valuable biomarkers for early diagnosis of lung cancer, monitoring disease progression, and guiding treatment decisions. Therefore, further research on these features will help us better understand the molecular mechanisms of lung cancer and provide support for precision medicine of lung cancer.

[0112] Overall results: the accuracy rate of the training set was 100%, the accuracy rate of the validation set was 96.59%, and the accuracy rate of the test set was 98.29%; the prediction effects of the three datasets were all good, the correct rates were all greater than 90%, and the model generalization ability was good. It was proved that the prediction model based on UVP-TOF-MS could effectively distinguish lung cancer patients from nodule patients.

[0113] In the above embodiments, the obtained exhaled breath samples were respectively divided into lung cancer samples, nodule samples and normal samples, the VOCs in the exhaled breath samples were analyzed using UVP-TOF-MS, a model was built, and the model was evaluated using a confusion matrix and an ROC curve.

[0114] From the evaluation results, it was known that in the three groups of non-lung cancer and lung cancer, normal people and lung cancer research, and nodules and lung cancer, the lung cancer screening model of the present invention had relatively excellent prediction ability, had the potential to be promoted to the clinic and used as an early lung cancer screening method, and the characteristic results provided ideas for possible potential lung cancer biomarkers. The established lung cancer screening model provided a new path for a convenient, accurate, cheap and fast early lung cancer screening method required clinically. At the same time, further research and improvement were needed in the optimization of the algorithm and the screening of characteristic markers in order to achieve better screening effects.

[0115] The above embodiments only represent the specific implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. For those of ordinary skill in the art, several deformations and improvements made without departing from the concept of the present application all belong to the protection scope of the present application.

Claims

1. Method for constructing lung cancer screening model based on UVP-TOF-MS It is characterized by including the steps of: A. Exhaled breath sample collection: Collect exhaled breath samples of confirmed lung cancer, normal exhaled breath samples, and nodule exhaled breath samples within a set time range, and define the normal exhaled breath samples and nodule exhaled breath samples as non-cancer exhaled breath samples; B. Perform full-spectrum analysis on each collected exhaled breath sample through a UVP-TOF-MS device to obtain the spectral data of each exhaled breath sample and form a spectral sample; C. Data preprocessing: Include data cleaning, mean filling of missing values, outlier deletion, calibration of the spectral sample data to the same horizontal distribution through calibration gas, environmental background deduction, and peak area calculation for the obtained spectral sample, divide the spectral sample into a training set, a validation set, and a test set, and select suitable features; among them, When calibrating the spectral sample data through the calibration gas, it includes calculating the peak area of the calibration gas with a mass-to-charge ratio in the range of [77.7, 78.4] every day, setting the value with a mass-to-charge ratio intensity lower than 300 to 0; then calculating the peak area of the calibration gas with a mass-to-charge ratio of 78, setting a standard value, and using the quotient of the standard value divided by the area of the calibration gas with a mass-to-charge ratio of 78 every day as the coefficient value. Finally, multiply the area of each mass-to-charge ratio of each exhaled breath sample by the coefficient value corresponding to the time to obtain the area data of each mass-to-charge ratio of each exhaled breath sample after calibration. Select all features in the range of mass-to-charge ratio from 15 to 249, and delete features with a mass-to-charge ratio of 94 and a proportion of peak area values of 0 greater than 90%; D. Model construction: Construct an ensemble learning model, and finely adjust the hyperparameters of the base classifier of the ensemble learning model through Bayesian optimization technology to obtain the best model parameter combination; the base classifier of the ensemble learning model selects the top N most important features according to the information gain importance ranking of each feature to form the feature set of the ensemble learning model, where N is a preset natural number; apply the features in the feature set to a logistic regression model, and the logistic regression model and the ensemble learning model together form a comprehensive lung cancer screening prediction model; E. Model performance evaluation: Through the confusion matrix, use the lung cancer screening prediction model to predict the test set, calculate the accuracy, sensitivity, and specificity indicators of the lung cancer screening prediction model to quantify the prediction ability of the lung cancer screening prediction model; then draw the receiver operating characteristic curve of the lung cancer screening prediction model and calculate the area under the curve to compare the diagnostic performance of the lung cancer screening prediction model under different parameters. Based on the analysis of the area under the curve, select the best-performing lung cancer screening prediction model.

2. The method for constructing a lung cancer screening model based on UVP-TOF-MS according to claim 1, wherein: In step B, the steps of performing full-spectrum analysis on each collected exhaled breath sample through a UVP-TOF-MS device are: After the UVP-TOF-MS device is preheated and the state is stable, connect the collected exhaled breath samples to the injection pipeline one by one. In the full-spectrum analysis of the exhaled breath samples, two spectral data points are collected for each exhaled breath sample, and the second spectral data of each exhaled breath sample is taken as the spectral sample of the exhaled breath sample.

3. The method for constructing a lung cancer screening model based on UVP-TOF-MS according to claim 1, characterized in that: In step C, the areas under the peaks of all qualitative VOCs in the spectral samples are normalized to the interval (0, 1). After data processing, a data matrix is generated. The spectral samples are divided into a training set, a validation set, and a test set at a set ratio, and then the data is corrected.

4. The method for constructing a lung cancer screening model based on UVP-TOF-MS according to claim 3, wherein: When correcting the data, calculate the peak area and the standard gas area of the calibration gas data for each day within the specified mass-to-charge ratio range, so as to obtain the area data of each mass-to-charge ratio after correction for each exhaled breath sample. Then, select all features within the set mass-to-charge ratio area range.

Citation Information

Patent Citations

  • Pulmonary tuberculosis risk assessment method and system based on exhaled gas mass spectrum detection

    CN114324549A

  • Self-adaptive lung cancer screening method based on expiration measurement data element learning

    CN117711633A