Lung cancer biomarker screening method based on neural network and SHAP interpretation

By using multi-layer perceptron neural network and SHAP method to screen lung cancer biomarkers in early diagnosis of lung cancer, the radiation risk, high cost or invasive problems present in the diagnosis process in the prior art is solved, and high-precision, interpretability and low-cost early diagnosis of lung cancer is achieved.

CN120148646APending Publication Date: 2025-06-13DALIAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510361216.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has radiation risks, high costs or invasive problems in early diagnosis of lung cancer. Traditional statistical models and machine learning methods cannot effectively capture nonlinear relationships and lack interpretability, resulting in incomplete marker screening and insufficient accuracy.

Method used

Multi-layer perceptron (MLP) neural network combined with Shapley Additive exPlanations (SHAP) method is used to screen the top 10% of metabolites with contributions as candidate markers through data acquisition and preprocessing, model training and verification, thereby improving the accuracy and interpretability of lung cancer biomarker screening.

Benefits of technology

High-precision early diagnosis of lung cancer was achieved, the average AUC of the classification model reached 0.991 and the variance was as low as 0.007, which significantly improved the accuracy and stability of marker screening and reduced the cost of clinical testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148646A_ABST
    Figure CN120148646A_ABST
Patent Text Reader

Abstract

The invention discloses a lung cancer biomarker screening method based on a neural network and SHAP interpretation, and belongs to the technical field of marker screening. According to the method, a multi-layer perceptron (MLP) and a Shapley interpretability algorithm are mainly used, expired gas metabolite data are obtained through a TD-GC-MS technology, a neural network is used for fitting the data, and a high-contribution-degree marker is screened through an SHAP method. Compared with a traditional method, the prediction accuracy of the classification model is remarkably improved. The invention provides a high-precision and high-stability solution for early diagnosis of non-invasive lung cancer, and has remarkable clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of biomedical engineering and artificial intelligence, and specifically relates to a method and system for screening lung cancer biomarkers by combining a non - linear machine learning model (neural network) with interpretability analysis (Shapley Additive exPlanations, SHAP), which is used to improve the accuracy of early lung cancer diagnosis and the interpretability of the model. Background Art

[0002] Lung cancer is one of the malignant tumors with the highest mortality rate globally, and early diagnosis is crucial for improving the cure rate. In the existing technology, low - dose CT (LDCT) and lung biopsy have problems such as radiation risk, high cost, or invasiveness, making it difficult to promote. Non - invasive detection methods based on metabolites (such as exhaled breath) have become a research hotspot, but traditional statistical models (such as volcano plots) and traditional machine learning methods (such as principal component analysis, PCA) have the following defects when screening biomarkers: Linear hypothesis limitation: Traditional statistical models cannot capture complex non - linear relationships, resulting in incomplete biomarker screening; Black - box problem: Machine learning models (such as neural networks) lack interpretability, making it difficult to verify the biological significance of biomarkers; Performance bottleneck: The biomarker combinations screened by existing methods have insufficient accuracy (AUC < 0.99) and stability (variance > 0.01) in classification models. Summary of the Invention

[0003] To solve the problems existing in the prior art, the present application provides a method for screening lung cancer biomarkers based on neural networks and Shapley explanations to meet the current requirements for lung cancer biomarker screening methods. This method has high precision, high interpretability, and is suitable for screening lung cancer biomarkers from complex biological data, with strong application value and practicality.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides a method for screening lung cancer biomarkers, including the following steps: Data collection and pre - processing: Obtain exhaled breath metabolite data of lung cancer patients and healthy volunteers through thermal desorption gas chromatography - mass spectrometry (TD - GC - MS) technology, and normalize the data;

[0005] Construct a multi - layer perceptron (MLP) neural network, input all metabolite data, and train the model through a cross - entropy loss function and an Adam optimizer; Use the SHAP method to analyze the output of the neural network, calculate the contribution degree of each metabolite to the classification result, and screen the metabolites with the top 10% contribution degrees as candidate biomarkers; Model Validation and Optimization: Based on the selected biomarker combinations, neural network models with the same architecture are trained, and the performance of different biomarker combinations is evaluated by comparing the area under the ROC curve (AUC), sensitivity, and specificity.

[0006] The beneficial effects of the present invention are as follows: High Precision: By fitting the non-linear relationship through a neural network and combining the SHAP interpretation to select biomarker combinations, the average AUC of the classification model reaches 0.991, and the variance is as low as 0.007. Interpretability: The SHAP method quantifies the impact of metabolites on the diagnostic results and clarifies the importance degree of each biomarker. High Efficiency: The number of selected biomarkers (36) is significantly less than that of traditional methods (69), reducing the clinical detection cost. Description of the Drawings

[0007] Figure 1 Screening method for lung cancer biomarkers.

[0008] Figure 2 Comparison between the improved method and the selected traditional method.

[0009] Figure 3 Lung cancer biomarker combination screened by the volcano plot method.

[0010] Figure 4 Lung cancer biomarker combination screened by the principal component analysis method.

[0011] Figure 5 Lung cancer biomarker combination screened by the improved feature screening method. Detailed Implementation Modes

[0012] Hardware: Configure a computer system with GPU acceleration to support neural network training. Software: Integrate the Python platform, TensorFlow deep learning framework, and SHAP library. In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail with reference to the accompanying drawings.

[0013] 1. Data Collection and Preprocessing Sample Source: The experimental data is from the Liaoning Cancer Hospital, with a total of 125 samples, including 73 lung cancer patients confirmed by pathology and 52 healthy volunteers. All subjects need to sign an informed consent form, and exhaled breath samples are collected after fasting for 12 hours.

[0014] Sampling method: Use a Teflon gas sampling bag to collect the end-expiratory gas of the subject in a single breath (excluding the gas in the first 4 seconds of exhalation), and enrich volatile organic compounds (VOCs) through a Tenax TA stainless steel enrichment tube to avoid metabolite decomposition.

[0015] Instrument parameters: A thermal desorption gas chromatography-mass spectrometry (TD-GC-MS, Agilent) was used. The chromatographic column was Agilent db624 (30m × 0.250mm × 1.4μm). The thermal desorption conditions were as follows: The carrier gas was helium (purity 99.8%), the desorption temperature was 250°C, the cold trap temperature was -30°C, and the desorption time was 10 minutes. During mass spectrometry analysis, the initial temperature was maintained at 40°C for 5 minutes, then increased to 250°C at a rate of 10°C / min and maintained for 4 minutes, and then increased to 260°C at a rate of 10°C / min and maintained for 10 minutes, with a total duration of 41.5 minutes.

[0016] Data normalization: To prevent the model from misjudging the importance of metabolites due to concentration differences, the concentrations of 1755 metabolites finally detected by GC-MS were first normalized. The formula is:

[0017] 2. Biomarker screening Neural network construction: Input layer: 1755 metabolite features.

[0018] Hidden layer structure: The first fully connected layer: 2048 neurons, with the activation function ReLU and the L2 regularization coefficient 0.06; Dropout layer: The random inactivation rate is 0.3; The second fully connected layer: 4096 neurons, with the activation function ReLU and the L2 regularization coefficient 0.06; Dropout layer: The random inactivation rate is 0.3; The third fully connected layer: 2048 neurons, with the activation function ReLU and the L2 regularization coefficient 0.06; Dropout layer: The random inactivation rate is 0.3; The fourth fully connected layer: 1024 neurons, with the activation function ReLU and the L2 regularization coefficient 0.06 Output layer: 2 neurons, and the Softmax function outputs the classification probabilities of lung cancer patients and healthy people.

[0019] Regularization strategy: All fully connected layers adopt L2 weight regularization (coefficient 0.06), and the Dropout layer suppresses overfitting.

[0020] Model compilation: Optimizer: Adam algorithm; Loss function: categorical cross-entropy; Evaluation metric: classification accuracy.

[0021] SHAP contribution calculation: Using the trained model, parse the neural network decision through the KernelExplainer of the SHAP library; Calculate the Shapley value of each metabolite feature for the classification result, and screen the top 10% (36 kinds) of metabolites with the highest contribution as candidate biomarkers.

[0022] Figure 1 The design of a lung cancer biomarker screening method based on neural network and Shapley interpretation includes a data preprocessing module, a model fitting module, a model interpretation module and a validation module. The present invention proposes an optimized lung cancer biomarker screening method, mainly replacing the commonly used statistical methods in the model fitting module with machine learning methods, and using the SHAP interpretability algorithm in the model interpretation module. The improved machine learning algorithm can not only extract the linear relationship between input and output, but also extract the non-linear relationship between input and input, input and output. Compared with traditional statistical methods, it can extract more key information. The SHAP interpretability algorithm is specifically used to explain the black-box relationship in the neural network and screen out lung cancer biomarkers.

[0023] As Figure 2 shown, to demonstrate the advantages of this method, the traditional statistical method - volcano plot method and the traditional machine learning method - principal component analysis (PCA) are used to screen lung cancer markers. Among them, the specific screening rule of the volcano plot method is: use one-way ANOVA to screen metabolites with a P value less than 0.01. The principal component analysis method ranks the metabolites in the sample according to their contribution and screens the top 10% of metabolites with the highest contribution. In the present invention, the multi-layer perceptron + SHAP feature extraction method will calculate the contribution of each feature and screen the top 10% of metabolites with the highest contribution.

[0024] As Figure 3 shown, using the volcano plot method, 64 eligible metabolites can be screened.

[0025] As Figure 4 shown, using the principal component analysis method, 69 eligible metabolites can be screened.

[0026] As Figure 5 shown, using the multi-layer perceptron + SHAP model interpretability algorithm, 36 eligible metabolites can be screened.

[0027] Table 1 Screening method Volcano plot Principal component analysis Neural network + SHAP Metabolites screened by all three methods C statistic 0.946±0.022 0.990±0.017 0.991±0.007 0.991±0.012 Accuracy rate (%) 0.838±0.008 0.937±0.012 0.965±0.012 0.964±0.017 Sensitivity (%) 0.905±0.122 0.942±0.038 0.970±0.020 0.967±0.027 Specificity (%) 0.744±0.168 0.931±0.065 0.961±0.031 0.957±0.030 As shown in Table 1, models with the same architecture were used to test the discrimination ability of these metabolite combinations. Finally, the biomarker combination screened by the multi-layer perceptron + SHAP interpretability algorithm had the best discrimination ability, and the AUC value was 0.991 and the accuracy rate was 0.965 with the least number of features used. It can be shown that this method can screen out key biomarker combinations.

Claims

1. A method for screening lung cancer biomarkers, characterized in that: The following steps are involved: S1. Obtain exhaled breath metabolite data through TD-GC-MS technology and perform normalization processing; S2. Use neural network combined with SHAP interpretation method to screen candidate markers.

2. The method for screening lung cancer biomarkers according to claim 1, characterized in that: The step S2 is specifically as follows: Construct a multi-layer perceptron neural network, input all metabolite data, and train the model using the cross entropy loss function and Adam optimizer; The SHAP method was used to analyze the neural network output, calculate the contribution of each metabolite to the classification results, and screen the top 10% metabolites as candidate markers.

3. The method for screening lung cancer biomarkers according to claim 2, characterized in that: The multi-layer perceptron neural network has a hidden layer activation function of ReLU, an output layer activation function of Softmax, a loss function of cross entropy, and an optimizer of Adam.

4. The method for screening lung cancer biomarkers according to claim 3, characterized in that: The multi-layer perceptron neural network: Hidden layer structure: The first fully connected layer: 2048 neurons, the activation function is ReLU, the L2 regularization coefficient is 0.06; Dropout layer: random inactivation rate is 0.3; The second fully connected layer: 4096 neurons, the activation function is ReLU, the L2 regularization coefficient is 0.06; Dropout layer: random inactivation rate is 0.3; The third fully connected layer: 2048 neurons, the activation function is ReLU, the L2 regularization coefficient is 0.06; Dropout layer: random inactivation rate is 0.3; The fourth fully connected layer has 1024 neurons, the activation function is ReLU, and the L2 regularization coefficient is 0.

06. Output layer structure: 2 neurons, Softmax function.

Citation Information

Patent Citations

  • Multi-model fusion-based interpretable breast cancer recurrence prediction method and system

    CN117438089A

  • Neural network-based gastric cancer disease risk assessment method and device

    CN118039147A

  • Pleural effusion auxiliary diagnosis method based on biomarker

    CN119314660A