Training method of early lung cancer saliva metabolism fingerprint detection model

By training saliva metabolic fingerprint data with an integrated model, the problems of high false positive rate in low-dose CT screening and large noise in saliva diagnosis were solved, achieving efficient and accurate early screening and diagnosis of lung cancer and improving sensitivity and specificity.

CN120748760APending Publication Date: 2025-10-03HANGZHOU FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510876921.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing technology has a high false positive rate in low-dose CT screening for lung cancer, and invasive diagnosis is frequent. Saliva metabolic fingerprint diagnosis faces problems of high noise and baseline drift, making it difficult to effectively identify the characteristics of early lung cancer.

Method used

An integrated model training method was adopted, combined with salivary metabolic fingerprint data, and feature selection was performed using the SFFS method. A weighted voting mechanism for multiple models such as random forest, logistic regression, and CatBoost was constructed. The performance was evaluated using the SHAP method, and a comprehensive evaluation was performed in combination with CA125 and CEA test values.

Benefits of technology

It has achieved efficient and accurate non-invasive early screening for lung cancer, with significantly improved sensitivity and specificity, ROC 0.849-0.850, sensitivity 81.69%-83.33% and specificity 74.23%-74.39%, which is superior to traditional tumor biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748760A_ABST
    Figure CN120748760A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of an early lung cancer saliva metabolism fingerprint detection model, and the method at least comprises the steps: obtaining saliva metabolism fingerprint data, and dividing the saliva metabolism fingerprint data into a discovery set, a verification set and a test set; inputting the discovery set into at least two preparatory models, performing feature selection in a K-fold cross validation framework by adopting an SFFS method, and performing parameter optimization on each preparatory model based on sensitivity, specificity, accuracy, F1 score and AUC; constructing an integrated model, wherein the integrated model comprises at least two preparation models and a weighted voting mechanism; and evaluating the performance of the integrated model on the verification set and the test set by using an SHAP method. The integrated model constructed by the method provided by the invention can efficiently, accurately and quickly perform non-intrusive screening and early detection on the lung cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical diagnosis, artificial intelligence and machine learning, and in particular to a training method for a saliva metabolic fingerprint detection model for early lung cancer. Background Art

[0002] Low-dose CT (LDCT) is the recommended lung cancer screening method. Clinical trials have shown that LDCT can reduce lung cancer (LC) mortality by 20%. However, the false-positive rate of LDCT screening is as high as 96.4%, resulting in unnecessary invasive diagnoses, severe psychological stress, and financial burdens for patients. Currently, most suspected lung cancer cases detected through LDCT screening still require further invasive confirmation. Therefore, there is an urgent need for effective differential methods for large-scale LC screening and early detection.

[0003] Lung cancer diagnosis based on serum tumor markers such as carcinoembryonic antigen (CEA) and carbohydrate antigen 125 (CA125) has high sensitivity and specificity, but this method requires blood collection from the patient. Most lung cancer patients are not diagnosed until the late stages of the disease, partly because early symptoms are not obvious and partly because blood collection requires an invasive procedure, which reduces the willingness to undergo early screening. Humans secrete saliva at high daily volumes, and this sample does not require invasive collection, making it crucial for the widespread use of early lung cancer diagnosis.

[0004] Despite this, early diagnosis of lung cancer using saliva metabolic fingerprints faces significant challenges. Firstly, saliva samples are significantly influenced by a patient's lifestyle and dietary habits, such as smoking, alcohol consumption, and differences in oral microbiome. Secondly, metabolic fingerprint data often suffers from high noise and baseline drift, making it difficult to effectively identify specific features using saliva metabolic fingerprints. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a method for training a model for early lung cancer detection based on saliva metabolic fingerprint data to achieve the desired sensitivity and specificity.

[0006] Technical solution: To achieve the above-mentioned purpose of the invention, the present invention provides a training method for a saliva metabolic fingerprint detection model for early lung cancer, comprising the following steps:

[0007] S100: obtaining saliva metabolic fingerprint data from lung cancer patients, benign lung lesion patients, and healthy subjects, and dividing the saliva metabolic fingerprint data into a discovery set, a validation set, and a test set;

[0008] S200: inputting the discovery set into at least two preliminary models, performing feature selection using the SFFS method within a K-fold cross-validation framework, and optimizing parameters of each preliminary model based on sensitivity, specificity, accuracy, F1 score, and AUC;

[0009] S300: Constructing an integrated model, wherein the integrated model includes at least two preliminary models and a weighted voting mechanism;

[0010] S400: Evaluate the performance of the integrated model on the validation set and the test set using the SHAP method.

[0011] The step 200 uses 10-fold cross validation to perform feature selection, evaluates the average AUC of each fold feature subset, and selects the feature subset with the highest AUC value as the feature subset.

[0012] In the evaluation of each fold feature subset, the feature with the lowest AUC value is deleted after each feature is added through a floating mechanism.

[0013] The weighted voting mechanism includes using an iterative algorithm to determine the weight distribution of the maximum AUC, thereby determining the optimal combination of the integrated model.

[0014] The output of the integrated model also includes a comprehensive evaluation of the probability value predicted by the integrated model and the CA125 test value and the CEA test value through the DCA decision voting mechanism.

[0015] The at least two preparatory models are selected from at least two of the following: Random Forest, Logistic Regression, CatBoost, Gradient Boosting, AdaBoost, LightGBM, k-NN, Support Vector Machine, and XGBoost models. Specifically, Random Forest generates multiple decision trees by randomly extracting data and features, and combines the results by voting or averaging to enhance the model's robustness and generalization capabilities. Logistic Regression, based on linear regression, utilizes the Sigmoid function to map linear outputs to probabilities for classification prediction. CatBoost utilizes symmetric decision trees and a gradient boosting framework to reduce bias and overfitting by sorting target variables and unordering categorical features. Furthermore, Gradient Boosting iteratively optimizes negative gradients, gradually constructing weak learners and performing weighted combinations to approximate complex functions. AdaBoost gradually adjusts sample weights, focusing on erroneous samples and gradually building a strong classifier. LightGBM, based on the gradient boosting framework, utilizes histogram optimization and GOSS sampling to improve computational efficiency. k-NN (k-NearestNeighbors) calculates the distance between samples and selects the nearest k neighbors to vote to determine the classification or regression results; Support Vector Machine (SVM) maximizes the classification boundary by finding the optimal hyperplane in high-dimensional space, making it suitable for handling nonlinear problems; XGBoost improves model performance and efficiency by optimizing the gradient boosting framework and combining regularization, sparse perception and parallel computing.

[0016] Furthermore, the step S100 further includes:

[0017] S110: Obtain saliva sample;

[0018] S120: performing a separation process on the saliva sample to separate at least mucin, bacteria, cell debris, and oral insoluble matter / impurities;

[0019] S130: Obtain analytical samples by protein precipitation;

[0020] S140: Obtaining structural information of the analysis sample based on a mass spectrometry analysis method.

[0021] Furthermore, the saliva metabolic fingerprint data includes the structural information and feature annotation information, wherein the feature annotation information is obtained by matching the m / z value in the HMDB database (http: / / www.hmdb.ca / ).

[0022] Furthermore, the method further includes a step of preprocessing the saliva metabolic fingerprint data obtained in step S100, wherein the preprocessing includes at least one or more of smoothing, baseline correction, intensity normalization, alignment, peak detection, and peak binning.

[0023] Furthermore, the method further includes a step of performing missing value processing on the saliva metabolic fingerprint data obtained in step S100.

[0024] Beneficial Effects: This study uses saliva fingerprints as input and, through machine learning, provides an integrated model for the early diagnosis of lung cancer. It also identifies 35 metabolic signatures associated with early lung cancer. In both validation and test sets, the model achieved a receiver operating characteristic (ROC) of 0.849-0.850, a sensitivity of 81.69%-83.33%, and a specificity of 74.23%-74.39%, all outperforming traditional tumor biomarkers. This demonstrates that the model, constructed based on machine learning extraction, can efficiently, accurately, and rapidly perform non-invasive screening and early detection of lung cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of a training method for a saliva metabolic fingerprint detection model for early lung cancer according to the present invention;

[0026] Figure 2 It is a schematic diagram of the principle of model training of the present invention;

[0027] Figure 3 The ROC curve comparison of the nine SFFS-based models in the validation set is shown in the following example;

[0028] Figure 4 The figure shows the comparison of AUC values ​​in the validation set of the five models using SFFS and p-value filtering methods for feature selection in the embodiment;

[0029] Figure 5 The ROC curve of all features of the model used in the embodiment on the validation set;

[0030] Figure 6 A line graph showing the average AUC values ​​of the top three SFFS models selected for the embodiment with different numbers of features in the ten-fold cross validation;

[0031] Figure 7 The distribution of up-regulated features in the LC group and the non-LC group in the examples is shown. The data are presented as the mean ± standard error (SE). The differences in metabolites between the LC group and the non-LC group were evaluated by the Mann-Whitney U test, and FDR correction was performed to determine the p.value adjustment value (* < 0.05, ** < 0.01, *** < 0.001);

[0032] Figure 8 Figure 2 shows the distribution of down-regulated features (B) in the LC group and the non-LC group in the examples. The data are presented as mean ± standard error (SE). The differences in metabolites between the LC group and the non-LC group were evaluated by Mann-Whitney U test, and FDR correction was performed to determine the adjusted p.value (* < 0.05, ** < 0.01, *** < 0.001).

[0033] Figure 9 To evaluate the analytical accuracy of the integrated model for specific metabolites in the examples, the bar graph shows the inter-batch CV values ​​of 35 selected characteristic metabolites in 46 batches of QC samples (n=3 per batch);

[0034] Figure 10 This is a radar chart comparing various indicators of the integrated model and the single model on the validation set in the embodiment;

[0035] Figure 11 This is a radar chart comparing various indicators of the integrated model and the single model on the test set in the embodiment;

[0036] Figure 12 3 is a comparison of the ROC curves of the integrated model and the tumor markers CA125 and CEA regarding sensitivity in the embodiment, where (A) represents the validation set and (B) represents the test set. DETAILED DESCRIPTION

[0037] In order to make the technical solution of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] Example 1

[0039] This example provides a training method for a saliva metabolic fingerprint detection model for early lung cancer. Figure 1 and Figure 2 , the method comprises at least the following steps:

[0040] S100: Obtain saliva metabolic fingerprint data from lung cancer patients, benign lung lesion patients, and healthy subjects, and divide the saliva metabolic fingerprint data into a discovery set, a validation set, and a test set;

[0041] S200: Input the discovery set into at least two preliminary models, perform feature selection using the SFFS method within the K-fold cross-validation framework, and optimize the parameters of each preliminary model based on sensitivity, specificity, accuracy, F1 score, and AUC;

[0042] S300: Build an integrated model, which includes at least two preliminary models and a weighted voting mechanism;

[0043] S400: Use the SHAP method to evaluate the performance of the integrated model on the validation set and test set.

[0044] The above method is described in detail below.

[0045] S110: Obtaining a saliva sample

[0046] The samples in this example were collected from saliva of patients with lung cancer (LC), benign lung lesions (BD), and healthy subjects (HC). Saliva samples were collected in the morning after the participants finished brushing their teeth. For at least one hour before sampling, typically one to two hours, participants were required to refrain from eating, drinking, brushing their teeth, and any form of exercise. On the day of sampling, participants were asked not to smoke, use oral sprays, or wear lipstick. To ensure sample purity, volunteers rinsed their mouths thoroughly three times with purified water, swallowed any remaining water, and then waited five minutes to eliminate any dilution effects.

[0047] S120: Separation and processing of saliva samples

[0048] The above separation process can be performed during sample collection using a specific collection device, or after collection, in a laboratory for separation / purification. A preferred embodiment utilizes a collection device capable of separating mucin, bacteria, cell debris, and oral insoluble matter / impurities. For example, the collection device incorporates a hydrophobic porous polymer filter plate (1-3 mm thick, 10-100 μm pore size) to retain mucin, cell debris, and insoluble matter in saliva. Bacteria and other viscous substances are further retained using a filter membrane with a pore size of 0.2-0.5 μm. Those skilled in the art can adjust the parameters of the filter plate or membrane as needed based on the size and adsorption properties of the desired retained substance.

[0049] In addition, the device can also be pre-loaded with a biological sample preservation solution. As a preferred embodiment, the preservation solution includes at least one C1-C3 monohydric alcohol and acetonitrile. After the biological sample is collected and separated, it is stored at -80°C.

[0050] S130: Obtaining analytical samples by protein precipitation

[0051] This step was performed using the ASSIST PLUS automated platform (INTEGRA), with 29 clinical samples and 3 quality control samples added per batch. The specific procedure was as follows: After thawing, the saliva samples were arranged in rows of 12 samples each in the cold zone of the platform (4°C). A 96-well polymerase chain reaction (PCR) plate was prepared, and 20 μL of acetonitrile / methanol (1:1, v / v) was added to each well. Subsequently, 10 μL of saliva sample was transferred to the plate using a robotic arm and vortexed at 1500 rpm for 5 minutes. The supernatant was then centrifuged at 2200 g for 5 minutes to separate the supernatant. 12 μL of the supernatant was then transferred to a new PCR plate, and 4 μL of ultrapure water was added. After thorough mixing, each sample solution was automatically deposited onto a Met-Si array chip (Hangzhou Huijian Technology Co., Ltd.), which contains a vertical SiNW array. The samples were then dried at a controlled humidity of 40-50% for 30-40 minutes. To verify the performance of the Met-Si array chip with the organic matrix, succinic acid, pyroglutamic acid, L-histidine, cysteic acid, and N-acetylhistidine were mixed with quality control samples to obtain a final concentration of 0.033 μg / μL. The stability of the Met-Si array chip was evaluated by laser desorption ionization (LDI).

[0052] S140: Obtaining structural information of the analyzed sample based on mass spectrometry

[0053] A Met-Si array chip was loaded onto an AutoFlex Max MALDI-TOF / TOF mass spectrometer (Bruker Daltonics Inc.) to acquire salivary metabolic fingerprints. Metabolic profiling of the cohort samples was performed in automated batch processing mode. Metabolite signatures were annotated by comparing the mass-to-charge ratios of the obtained precursor and fragment ions with those of standard metabolites and searching the HMDB (http: / / www.hmdb.ca / ). The relative error of the mass-to-charge ratio was set to 35 ppm.

[0054] S150: Saliva metabolic fingerprint data preprocessing

[0055] Saliva metabolic fingerprint data were preprocessed using MALDIquant in the R language, including smoothing, baseline correction, intensity normalization, alignment, peak detection, and binning. A total of 646 metabolite peaks were detected in the spectra, and these peaks were expressed in more than 80% of the samples (signal-to-noise ratio >5). The robustness of the entire workflow was assessed by calculating the coefficient of variation (CV) for QC samples. Principal component analysis (PCA), t-SNE, and Unified Mapping (UMAP) were used to assess the distribution of samples. These analyses were also performed in R using the prcomp function for PCA, the Rtsne package for t-SNE, and the umap package for UMAP. Differential metabolites were evaluated using the Mann-Whitney U test, followed by FDR correction to obtain adjusted p-values. Correlation analysis was performed using Spearman correlation in the psych package. Pathway analysis was performed using the MetaboAnalyst R package, and OR analysis was performed using the logistic regression (LR) algorithm to assess the impact of features.

[0056] In addition, transcriptome data of tumor and normal tissues were downloaded from the public TCGA database (https: / / www.cancer.gov / tcga) to investigate LC-related metabolic disorders. To determine the appropriate sample size and predictive power, machine learning-based power analysis was applied in the discovery set.

[0057] S200: Input the discovery set into at least two preliminary models, perform feature selection using the SFFS method within the K-fold cross-validation framework, and optimize the parameters of each preliminary model based on sensitivity, specificity, accuracy, F1 score, and AUC. Specific methods include the following:

[0058] The entire sample cohort was divided into a discovery set, a validation set, and a test set, with a ratio of approximately 7:3:3. Features with a coefficient of variation (CV) exceeding 25% in the quality control samples were subsequently excluded. Within the discovery set, feature selection was performed using nine algorithms (random forest, logistic regression, CatBoost, gradient boosting, AdaBoost, LightGBM, k-NN, support vector machine, and XGBoost) within a K-fold cross-validation framework combined with SFFS. The specific steps involved were as follows:

[0059] S210: Starting from the original feature set, the initial feature set size is n = 1;

[0060] S220: Generate all feature subsets of size n + 1 = 2, evaluate the AUC of each feature subset through 10-fold cross validation, and use the average AUC of all folds to measure its performance; iterate;

[0061] S230: Floating search, adding features while removing the feature with the smallest AUC;

[0062] S240: Repeat steps 2 and 3 until the maximum number of features is reached;

[0063] S250: Based on all the iteration results, the subset with the highest AUC value is selected as the final feature subset.

[0064] The present invention identified metabolic features associated with LC using a machine learning strategy. Features with a coefficient of variation (CV) exceeding 25% in QC samples were excluded. Next, the SFFS method was applied to further screen the discovery set to optimize the retained features. By iteratively adding or removing features based on the model's AUC value, the SFFS method effectively sorted the feature subsets. In comparison, features were extracted using a traditional p-value feature selection method. Using different algorithms, nine SFFS-based models were established with AUC values ​​ranging from 0.703 to 0.804, indicating that these models performed well in identifying LC and non-LC in the validation set ( Figure 3 The first five algorithms selected from the SFFS model were also used to construct a diagnostic model based on p-value selection features. Figure 4 As shown in Figure 2, among the five algorithms, the AUC values ​​of the SFFS-based model (0.761-0.804) in the validation set were consistently better than those of the p-value-based model (0.723-0.782). Figure 5 As shown, the top five SFFS models outperformed all feature-based models (AUC = 0.720-0.802). This indicates that the feature subsets filtered by the SFFS method performed better. In distinguishing LC and non-LC based on saliva metabolic fingerprints, features were identified by common p-value screening methods or complete feature sets. Based on the AUC ranking, three algorithms, RF, LR, and CatBoost, were determined. Among the top three SFFS models, subsets of 12, 22, and 33 features were selected, respectively. These models were named RF, LR, and CatBoost, because after the AUC values ​​initially increased to an optimal point, as each algorithm continued to add features, the AUC values ​​began to fluctuate slightly or decrease ( Figure 6 ). Finally, a total of 35 features were obtained.

[0065] These 35 features were annotated by MS / MS fragmentation and comparison of the m / z values ​​of precursor and fragment ions with those of standard metabolites, as well as searching the HMDB database. Compared with the non-LC group, 10 metabolites were significantly upregulated and 10 were significantly downregulated (e.g. Figure 7 、 Figure 8 ), while the other 15 showed no significant differences in LC patients.

[0066] S300: Build Random Forest (RF), Logistic Regression (LR) and CatBoost ensemble models.

[0067] After S200 training is complete, each base model makes predictions on the validation set and outputs its own predicted probability value. To combine these independent predictions into the final predicted probability, a weighted average soft voting strategy is employed. In this strategy, the weights of each model in the ensemble are optimized based on its AUC value on the validation set. The optimization process involves exploring multiple weight combinations to identify the optimal weight allocation that maximizes the AUC.

[0068] Specifically, three weights w1, w2, and w3 are given. Their values ​​are constrained to range between 0 and 1, and must satisfy the condition w1 + w2 + w3 = 1. To facilitate optimization, these weights are discretized into multiple candidate combinations within the specified range. Each weight combination is applied to the ensemble model, and the corresponding AUC value is calculated to assess its effectiveness. By traversing all possible weight combinations, the combination with the highest AUC value is selected as the optimal weight allocation.

[0069] This strategy significantly improves the predictive performance of the ensemble model by optimizing the weight distribution. Through an iterative optimization process, a weighting scheme that maximizes the AUC of the ensemble model is ultimately determined, further enhancing the performance of the overall model.

[0070] S400: Use the SHAP method to evaluate the performance of the integrated model on the validation set and test set.

[0071] In this step, we further used SHAP analysis (Shapley Additive ExPlanation) in Python and TreeExplainer to calculate and explain the contribution of the 35 metabolites to the model prediction. A positive SHAP value indicates a positive contribution to the prediction result, while a negative SHAP value indicates a negative contribution to the prediction result.

[0072] like Figure 9 As stated, Figure 9 The analytical precision evaluation of the above ensemble model is shown, with the bar graph showing the inter-batch coefficient of variation (CV) values ​​of 35 selected characteristic metabolites in 46 batches of QC samples (n = 3 per batch).

[0073] like Figure 10As shown in the figure, the sensitivity, specificity, accuracy, and F1 score of the ensemble model provided by the present invention are superior to those of the basic model in the validation set, indicating the advantage of adopting an ensemble strategy in LC identification by voting. In the test set, the ensemble model maintained high performance with an AUC of 0.849, a sensitivity of 81.69%, a specificity of 74.23%, and an accuracy of 76.50%, respectively. Figure 11 As shown in Table 1, the ensemble model also outperforms the base model in the test set, which indicates that combining multiple algorithms can improve diagnostic accuracy. In addition, Table 1 shows that the ensemble model outperforms the other algorithms - the SFFS model - in terms of sensitivity, specificity, accuracy, and F1 score in the test set.

[0074] Table 1. Performance comparison of nine algorithm models on the validation set and test set after feature selection using the SFFS method

[0075]

[0076] like Figure 12 As shown, the ensemble model provided by the present invention achieved AUC values ​​of 0.850 and 0.849 in the validation and test sets, respectively, significantly higher than those of CA125 (0.724, 0.667) and CEA (0.565, 0.581). Although the specificity of CA125 and CEA was higher than that of the ensemble model (74.23%-74.39%), the ensemble model outperformed CA125 (1.61%-4.23%) and CEA (9.68%-12.50%) in terms of sensitivity (81.69%-83.33%). These data suggest that the ensemble model provided by the present invention may have greater clinical advantages than traditional serological biomarkers in the diagnosis of lung cancer.

[0077] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A training method for a saliva metabolic fingerprint detection model for early lung cancer, characterized by: S100: obtaining saliva metabolic fingerprint data from lung cancer patients, benign lung lesion patients, and healthy subjects, and dividing the saliva metabolic fingerprint data into a discovery set, a validation set, and a test set; S200: inputting the discovery set into at least two preliminary models, performing feature selection using the SFFS method within a K-fold cross-validation framework, and optimizing parameters of each preliminary model based on sensitivity, specificity, accuracy, F1 score, and AUC; S300: Constructing an integrated model, wherein the integrated model includes at least two preliminary models and a weighted voting mechanism; S400: Evaluate the performance of the integrated model on the validation set and the test set using the SHAP method.

2. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 1, characterized in that: The step 200 uses 10-fold cross validation to perform feature selection, evaluates the average AUC of each fold feature subset, and selects the feature subset with the highest AUC value as the feature subset.

3. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 2, characterized in that: In the evaluation of each fold feature subset, the feature with the lowest AUC value is deleted after each feature is added through a floating mechanism.

4. The method for training a saliva metabolic fingerprint detection model for early lung cancer according to any one of claims 1 to 3, characterized in that: The weighted voting mechanism includes using an iterative algorithm to determine the weight distribution of the maximum AUC, thereby determining the optimal combination of the integrated model.

5. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 1, characterized in that: The output of the integrated model also includes a comprehensive evaluation of the probability value predicted by the integrated model and the CA125 test value and the CEA test value through the DCA decision voting mechanism.

6. A method for training a saliva metabolic fingerprint detection model for early lung cancer according to any one of claims 1 to 3 and 5, characterized in that: The at least two preliminary models are selected from at least two of random forest, logistic regression, CatBoost, gradient boosting, AdaBoost, LightGBM, k-NN, support vector machine and XGBoost models.

7. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 6, characterized in that: The step S100 further includes: S110: Obtain saliva sample; S120: performing a separation process on the saliva sample to separate at least mucin, bacteria, cell debris, and oral insoluble matter / impurities; S130: Obtain analytical samples by protein precipitation; S140: Obtaining structural information of the analysis sample based on a mass spectrometry analysis method.

8. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 7, characterized in that: The saliva metabolic fingerprint data includes the structural information and feature annotation information, wherein the feature annotation information is obtained by matching the m / z value in the HMDB database.

9. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 1, characterized in that: The method further includes a step of preprocessing the saliva metabolic fingerprint data obtained in step S100, wherein the preprocessing includes at least one or more of smoothing, baseline correction, intensity normalization, alignment, peak detection, and peak binning.

10. The training method for a saliva metabolic fingerprint detection model for early lung cancer according to claim 1, characterized in that: The method further includes a step of performing missing value processing on the saliva metabolic fingerprint data obtained in step S100.