A serum metabolic fingerprint-based multi-omics differential diagnosis model for lung benign and malignant nodules and a construction method thereof
Patent Information
- Application Number
- CN202211138968.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-09-19
AI Technical Summary
国内外研究者基于机器学习、深度学习等计算机辅助技术提出了不同的影像学肺结节风险预测模型,但多中心异质影像导致的模型难以泛化,影响了真实临床场景的使用
[0029]1、本发明提供的诊断模型通过获取样本的血清代谢指纹、CEA含量信息和患者胸部CT影像,实现代谢组学、肿瘤蛋白标志物和CT影像特征的三模态数据分析,最终快速、便捷、准确的区分患者肺部结节的良恶性。
Smart Images

Figure CN115458173B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of analysis, and particularly relates to a lung benign and malignant nodule multi-omics differential diagnosis model based on serum metabolic fingerprints and a construction method thereof. BACKGROUND
[0002] Lung cancer is one of the malignant tumors with the highest incidence and mortality in China. In the past three decades, the mortality rate of lung cancer has increased by about 5 times, and a major reason is that the early characteristics of lung cancer are not obvious, and 75% of cancer patients are diagnosed in the middle and late stages. The early detection rate of lung cancer is less than 25%, but the 5-year survival rate of early lung cancer can reach more than 90%.
[0003] Currently, low-dose helical computed tomography (LDCT) screening for early lung cancer has reached a consensus. With the popularization of LDCT screening, 50% of the population receiving LDCT screening will detect pulmonary nodules (PN). Pulmonary nodules refer to local, round, and density-increased lung shadows with a diameter of less than or equal to 30 mm. According to the imaging density, they are divided into solid nodules (SN), part-solid nodules (PSN), and pure ground glass nodules (pGGN). 95% of pulmonary nodules are caused by benign lesions, including granulomas, lymph nodes, chronic inflammation, and hamartoma, while malignant lesions mainly include adenocarcinoma and squamous cell carcinoma. Through in-depth clinical research, many countries have established corresponding diagnosis and treatment standards and guidance paths. Currently, major guidelines / consensus are based on clinical information (age, smoking history, and family history of tumors) and imaging features (size, density, and boundary characteristics of nodules) for evaluation. For nodules with different malignant risks, different clinical management measures such as no intervention, imaging follow-up, non-surgical biopsy, and surgical resection are developed. International and domestic clinical centers have also constructed lung nodule malignant risk prediction models based on clinical cohort data. The American College of Chest Physicians (ACCP) guidelines recommend using the Mayo model for nodule risk assessment, and the British Thoracic Society (BTS) guidelines recommend using the Brock model to calculate the malignant risk of nodules larger than 8 mm. However, the evaluation of imaging indicators (diameter, burr, composition classification, etc.) and clinical data (smoking history) in the model has certain subjectivity, with significant inter-reader differences and variability. In actual clinical practice, some information may be missing, leading to certain difficulties and challenges in the actual clinical application of the model. Currently, international clinical centers based on the above path and model have an over-diagnosis rate of 18-25% for pulmonary nodules. Especially for patients with solid imaging and sub-centimeter diameter, about 30% of them are benign after surgery, which not only causes serious psychological burden to patients but also leads to huge waste of health resources.
[0004] With the rapid development of omics and bioinformatics analysis technologies, detection and analysis methods based on biology and imaging have shown great promise in the field of in vitro diagnostics. Researchers both domestically and internationally have proposed various imaging-based lung nodule risk prediction models using computer-aided techniques such as machine learning and deep learning. However, the heterogeneous imaging from multiple centers makes these models difficult to generalize, affecting their use in real-world clinical scenarios. In the field of bioomics, molecular markers based on tumor cell-free DNA, autoantibodies, and proteins have also shown potential application value, but due to limitations in the sensitivity of detection technologies and the differences in methodologies and data analysis, they are currently still in the clinical evaluation stage. Summary of the Invention
[0005] To address the problems in the prior art, this invention discloses a multi-omics differential diagnostic model for benign and malignant pulmonary nodules based on serum metabolic fingerprinting and its construction method. This model can realize trimodal data detection, improve the sensitivity and accuracy of benign and malignant pulmonary nodule differentiation, and the model construction method is simple, convenient and fast, with high detection throughput and low detection cost, making it easy to promote and apply in clinical practice.
[0006] To address the aforementioned problems, this invention first provides a method for constructing a multi-omics differential diagnostic model for benign and malignant pulmonary nodules based on serum metabolic fingerprinting, comprising the following steps:
[0007] S1, obtain the raw metabolic fingerprint profiles and CEA protein content of serum samples from lung adenocarcinoma and control serum samples;
[0008] S2, based on the original metabolic fingerprint profiles of the two samples, a single-modal diagnostic model is constructed using machine learning methods; further, the score of the single-modal diagnostic model and the CEA content are combined as inputs to construct a dual-modal diagnostic model using machine learning methods.
[0009] S3: Acquire CT images of patients with pulmonary nodules and input them into the CT image-assisted diagnostic model;
[0010] S4. Using the scores of the dual-modal diagnostic model obtained in step S2 and the CT image-assisted diagnostic model obtained in step S3 as input, a three-modal diagnostic model is constructed using machine learning methods, namely the multi-omics differential diagnostic model for benign and malignant lung nodules based on serum metabolic fingerprint.
[0011] Preferably, the CT image-assisted diagnostic model is a risk prediction model for benign or malignant pulmonary nodules constructed solely based on CT image information.
[0012] Preferably, the machine learning method includes any one or more of support vector machines, neural networks, Gaussian Naive Bayes, or AdaBoost algorithms.
[0013] Preferably, in step S1, MALDI-MS technology is used to perform metabolic detection on the lung adenocarcinoma serum sample and the control serum sample to obtain the original metabolic fingerprint profile.
[0014] Furthermore, the specific steps for obtaining the raw metabolic fingerprint using MALDI-MS technology include:
[0015] (1) Collect serum samples from patients with lung adenocarcinoma and non-lung adenocarcinoma controls, and prepare nanomatrix materials;
[0016] (2) The serum sample and the nanomatrix material were diluted with deionized water to prepare the serum sample and nanomatrix suspension to be tested;
[0017] (3) Spot the serum sample to be tested on the LDI-MS mass spectrometry target plate, dry it at room temperature, then spot the matrix suspension, and dry it at room temperature;
[0018] (4) The serum sample to be tested was detected in LDI-TOF-MS to obtain the original metabolic fingerprint of the serum sample.
[0019] Preferably, in step (2), the serum sample is diluted 10 times and the concentration of the nanomatrix suspension is 1 mg / mL.
[0020] In some embodiments, the nanomatrix material includes metallic nanomaterials such as iron nanoparticles, silver nanoparticles, and gold nanoparticles, as well as composite nanomaterials combining metals and inorganic materials, to ensure high specific heat, low thermal conductivity, plasmon resonance effect, good ultraviolet light absorption, and a rough porous structure. The nanomatrix material can be commercially available or prepared in a laboratory.
[0021] In another aspect, the present invention provides a multi-omics differential diagnostic model for benign and malignant pulmonary nodules based on serum metabolic fingerprints, constructed according to the construction method described in any of the preceding claims.
[0022] In another aspect, the present invention provides a method for using the aforementioned multi-omics differential diagnostic model for benign and malignant pulmonary nodules based on serum metabolic fingerprinting, characterized by comprising the following steps:
[0023] (1) Take a serum sample to be tested and analyze it using MALDI-MS technology to obtain the original metabolic fingerprint of the serum sample to be tested; at the same time, detect the CEA content in the serum sample to be tested and obtain the patient's chest CT image.
[0024] (2) Perform chromatographic preprocessing on the original metabolic fingerprint to obtain the serum metabolic fingerprint of the sample;
[0025] (3) Input CT images into the CT image-assisted diagnostic model;
[0026] (4) Input the CEA content, serum metabolic fingerprint and the score of the CT image-assisted diagnostic model in step (3) into the multi-omics differential diagnosis model for benign and malignant lung nodules. The model gives a score of 0-1 based on the probability of malignancy.
[0027] Furthermore, when the model score is greater than the cutoff value, it indicates that the patient's lung nodule has a high probability of malignancy and requires further treatment or follow-up; when the model score is not higher than the cutoff value, it indicates that the patient's lung nodule has a low risk of malignancy.
[0028] Compared with the prior art, the beneficial effects of the present invention are:
[0029] 1. The diagnostic model provided by this invention obtains serum metabolic fingerprints, CEA content information and chest CT images of patients to achieve trimodal data analysis of metabolomics, tumor protein markers and CT image features, and finally quickly, conveniently and accurately distinguishes between benign and malignant lung nodules in patients.
[0030] 2. Compared with traditional single-modal CT imaging diagnostic methods, the diagnostic model of this invention improves the sensitivity, accuracy and throughput of identifying benign and malignant pulmonary nodules. Moreover, the model construction method is simple, convenient and quick, which greatly reduces the workload and cost of screening, making it suitable for large-scale screening and easy to promote and apply in clinical practice. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating the construction process of the multi-omics differential diagnosis model for benign and malignant lung nodules based on serum metabolic fingerprinting, as described in this invention.
[0032] Figure 2 The results show the characterization of iron nanoparticles, among which,
[0033] a is the SEM image;
[0034] b is the TEM image;
[0035] c represents dynamic light scattering measurement;
[0036] d represents the optical absorption spectrum of the material.
[0037] Figure 3 This is a schematic diagram of the established local serum metabolic fingerprint database; where,
[0038] a shows typical serum metabolic mass spectrometry of lung adenocarcinoma and benign lung diseases (granulomas) (inset shows H&E staining images of tissues confirmed by pathology);
[0039] b is the serum metabolic fingerprint extracted from the original metabolic fingerprint after preprocessing.
[0040] Figure 4 Comparison of ROC curves for different diagnostic models of benign and malignant pulmonary nodules, among which,
[0041] The green dashed line represents the ROC curve of the three-modal diagnostic model constructed in this invention;
[0042] The blue solid line represents the ROC curve of the CT image-assisted diagnostic model Image-AI.
[0043] The orange solid line represents the ROC curve of the clinical diagnostic model VA;
[0044] The green solid line represents the ROC curve of the Mayo clinical diagnostic model;
[0045] The purple dashed line represents the ROC curve of the tumor marker CEA. Detailed Implementation
[0046] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0047] As mentioned above, in view of the shortcomings of the prior art, the applicant of this invention, after long-term research, proposes the technical solution of this invention, the preparation process of which is as follows: Figure 1 As shown: First, MALDI-MS technology was used to capture metabolites from complex biological samples, thereby sensitively and selectively collecting metabolites (100–1000 Da) and establishing a local serum metabolic fingerprint database. Then, machine learning methods were used to learn the metabolic fingerprint, constructing a single-modal diagnostic model and outputting its score. Next, the sample's clinical indicator, CEA content, was combined with machine learning to construct a dual-modal diagnostic model. Chest CT images of patients with pulmonary nodules were acquired, and after analysis using the CT image-assisted diagnostic model, the scores from the dual-modal model were further combined with machine learning to construct a tri-modal diagnostic model, resulting in a multi-omics differential diagnostic model for benign and malignant pulmonary nodules.
[0048] the term
[0049] The “MALDI-MS technology” mentioned in this invention refers to LDI-MS technology based on nanomaterial matrix.
[0050] The "CT image-assisted diagnostic model" described in this invention is a risk prediction model for benign or malignant lung nodules based solely on CT image information. In some embodiments, commercially available diagnostic models can be used, such as the Image-AI model developed by Sanmed Biotech, the United Imaging Intelligence CT image-assisted lung nodule detection software developed by United Imaging Intelligence, and the lung nodule CT image-assisted detection software developed by Infervision. In some embodiments, a risk prediction model for benign or malignant lung nodules based solely on CT image information, constructed by the research group using machine learning methods, can also be used.
[0051] Example 1
[0052] The experimental process and results of this invention will be described in detail below.
[0053] The construction method and efficacy verification of the multi-omics differential diagnostic model for benign and malignant pulmonary nodules based on serum metabolic fingerprinting of this invention are as follows:
[0054] 1. Research Subjects
[0055] This study included 2276 participants who visited or underwent physical examinations at Shanghai Chest Hospital between November 2016 and May 2018. Among them, 320 were patients with benign lung diseases, 958 were patients with lung adenocarcinoma, and 998 were healthy controls. Benign diseases included pneumonia, chronic obstructive pulmonary disease, and tuberculosis. Lung adenocarcinoma patients were confirmed by histopathology and / or cytopathology, and staging was based on the 8th edition of the TNM staging system. Healthy controls were outpatients undergoing physical examinations. Patients lacking histopathological diagnosis, acute illness history, or other malignant tumors were excluded. All participants signed informed consent forms. This study has completed clinical trial registration (ChiCTR2000036938).
[0056] In constructing the unimodal and bimodal diagnostic models described in this invention, all subjects in this study were randomly assigned to the training set (2 / 3 of the total sample) and the test set (1 / 3 of the total sample). The lung adenocarcinoma patients enrolled in this study were primarily early-stage lung adenocarcinoma patients (stage I and II), accounting for 71.2% (458 / 643) in the training set and 75.2% (237 / 315) in the test set.
[0057] In constructing the aforementioned trimodal diagnostic model, 480 patients with pulmonary nodules on CT scans were selected and randomly assigned to the pulmonary nodule training set (4 / 5 of the total sample) and the pulmonary nodule test set (1 / 5 of the total sample). Relevant information for these 480 patients with pulmonary nodules is shown in Table 1.
[0058] Table 1. Relevant information of 480 patients with pulmonary nodules
[0059]
[0060]
[0061] 2. Establishing a serum metabolic fingerprint database using MALDI-MS
[0062] Whole blood samples were collected from subjects after a one-night fast to eliminate dietary interference. Serum was obtained by centrifugation at 3500 rpm for 10 min at 4°C and stored at -80°C. Raw mass spectrometry data were acquired using an Autoflex speed time-of-flight mass spectrometry (Bruker, Germany) instrument.
[0063] 2.1 Instruments and Equipment
[0064] The experimental instruments included: an ultrapure water system (Milli-Q, Millipore, USA), a mass spectrometer (Autoflex SpeedTOF / TOF, Bruker, Germany), a transmission electron microscope (2100F, JEM, Japan), a scanning electron microscope (S-4800, Hitachi, Japan), and a nanoparticle size potentiometer (Mastersizer 3000, Malvern, UK).
[0065] Experimental consumables include: pipettes, 10μL pipette tips, 100μL pipette tips, 1.5mL centrifuge tubes, markers, gloves, and masks.
[0066] 2.2 Preparation and Characterization of Nanoscale Matrix Materials
[0067] (1) Weigh 0.60g of ferric chloride hexahydrate, 0.15g of trisodium citrate and 0.96g of sodium acetate and dissolve them in ethylene glycol solution in sequence. The solution is then sonicated to make it homogeneous.
[0068] (2) The above mixed solution was transferred to a reactor with a capacity of 50 mL and heated to 200 °C for 10 h to obtain trivalent iron nanoparticles.
[0069] (3) The obtained ferric nanoparticles were washed several times with ethanol and deionized water until the supernatant was colorless. The final product was then dried at 60°C for 12 hours and stored in a vacuum for later use.
[0070] To characterize the trivalent iron nanoparticles prepared above, SEM images were obtained using an S-4800 scanning electron microscope; transmission electron microscopy (TEM) images were recorded using a JEM-2100F instrument; dynamic light scattering measurements were performed on a Nano ZS instrument (Malvern, Worcestershire, UK); and optical absorption spectra of the materials were collected on a UV1900 UV-Vis spectrometer (Aucybest, China). The characterization results are as follows: Figure 2 As shown.
[0071] 2.3 Serum MALDI-MS detection
[0072] Metabolic fingerprints of serum samples were obtained using enhanced laser desorption / ionization time-of-flight mass spectrometry (MALDI-MS) based on the trivalent iron nanoparticles prepared above, specifically including the following steps:
[0073] (1) Preparation of nanomatrix suspension: The trivalent iron nanoparticles obtained in step 2.2 were diluted with deionized water to 1 mg / mL;
[0074] (2) Dilute the subject's serum sample 10 times with deionized water;
[0075] (3) Sample preparation on the mass spectrometry target plate: Spot 1 μL of each diluted serum sample or standard and dry at room temperature;
[0076] (4) Matrix preparation on the mass spectrometry target plate: 1 μL of each matrix suspension was spotted and dried at room temperature;
[0077] (5) Serum metabolic fingerprints were collected using LDI-TOF-MS.
[0078] The laser source was an Nd:YAG laser with a wavelength of 355 nm and a maximum frequency of 2 kHz. The mass spectrometry data acquisition mode was set to positive ion reflectance mode, and the molecular weight range of the oligonucleotides to be detected was set to 100 to 1000 Da. During the experiment, the standard parameters were set to a laser frequency of 1000 Hz and a laser intensity of 70%. The data obtained for each experiment were superimposed spectra obtained from 2000 laser shots. Standard molecules were used for mass calibration to ensure accurate mass measurement and avoid intra-plate bias. All serum samples were randomly dropped onto multiple 384-well target plates to reduce systematic errors and inter-plate differences caused by uneven sample type distribution. In addition, five independent experiments were performed to eliminate intra-individual bias in order to enhance the reproducibility and stability of the diagnostic results. During metabolite identification, only signals with a mass spectrometry signal-to-noise ratio (S / N) greater than 3 were used for molecular identification and recognition based on accurate mass alignment. In addition to the precise mass comparison method (±0.05 Da), for selected specific small molecules, the molecular peaks of the secondary mass spectra of their mass spectra (from biological samples and standards) are compared with each other to finally confirm the metabolite to be tested.
[0079] 2.4 Determination of serum tumor markers
[0080] The CEA content of serum samples from 2276 subjects was detected using a Roche Cobas e601 instrument (electrochemiluminescence assay kit). The cutoff values were obtained from the kit instructions.
[0081] 2.5 Construction and Score Calculation of a Multi-omics Differential Diagnosis Model for Benign and Malignant Lung Nodules
[0082] 2.5.1 Model training for serum metabolic fingerprinting
[0083] (1) In step 2.3, the serum metabolic fingerprint profiles of 2276 subjects were imaged, and the imaging results are as follows: Figure 3 As shown in Figure a, the upper part of the figure is the fingerprint spectrum of disease samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry; the lower part of the figure is the fingerprint spectrum of samples with benign lung diseases detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry.
[0084] (2) Preprocessing of 2276 serum metabolic fingerprint profiles: First, Gaussian filtering with σ=1 was used for noise reduction and curve smoothing; then, Top-Hat operation was used for baseline correction; finally, local maximum processing was used to extract the final metabolic molecular features. A total of 2316 feature signals in the range of 100-1000 Da were obtained, constituting the serum metabolic fingerprint database described in this invention. The feature signal profiles are shown below.Figure 3 As shown in Figure b, the upper part of the figure is the characteristic signal spectrum of serum samples from 958 patients with lung adenocarcinoma as diseased samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry, while the lower part of the figure is the characteristic signal spectrum of serum samples from 1318 healthy samples and benign lung disease samples as control samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry.
[0085] (3) The 2276 preprocessed samples were randomly divided into a training set (2 / 3 of the total samples) and a test set (1 / 3 of the total samples). The training set included 643 patients with lung adenocarcinoma, 214 patients with benign lung diseases, and 669 healthy controls. The test set included 315 patients with lung adenocarcinoma, 106 patients with benign lung diseases, and 329 healthy controls. It should also be noted that the lung adenocarcinoma patients enrolled in this study were mainly early-stage lung adenocarcinoma patients (stage I and II), accounting for 71.2% (458 / 643) in the training set and 75.2% (237 / 315) in the test set.
[0086] (4) Perform principal component analysis (PCA) on the training set. Multiple principal components are fitted by PCA principal component analysis. The first 75 principal components (PC1-PC75) are initially selected for further analysis. After Pearson correlation analysis, PC38 is removed, and the remaining 74 principal components are input into support vector machine (SVM) for model training to obtain a single-modal initial diagnostic model.
[0087] Specifically: 10-fold cross-validation was used to train the SVM model on the training set. The SVM parameters were: C = 2.794450000000003, tol = 0.0001, coef0 = 0, kernel = rbf, class_weight = balanced, degree = 3, gamma = auto, probability = True.
[0088] (5) Test the test set in the single-modal initial diagnostic model to obtain the single-modal diagnostic model, which is used to obtain the predicted score of the metabolic molecule scoring module.
[0089] 2.5.2 Constructing a dual-modal diagnostic model by combining the single-modal diagnostic model score and CEA content
[0090] (1) In the training set, the predicted score of the metabolic molecule scoring module obtained in step 2.5.1 (i.e. the single-modal diagnostic model score) and the serum CEA content of the sample obtained in step 2.4 are used as joint inputs. The Gaussian Naive Bayes (GaussianNB) algorithm is used to train the model and obtain the dual-modal initial diagnostic model.
[0091] (2) The test set was tested in the bimodal initial diagnostic model to obtain a bimodal diagnostic model. This bimodal diagnostic model can be used for screening early lung adenocarcinoma and has excellent performance. Given the excellent performance of the bimodal diagnostic model in diagnosing early lung adenocarcinoma, the applicant of this invention applied the bimodal model to a training set and a test set of 480 patients with lung nodules to test its performance in differentiating between benign and malignant lung nodules.
[0092] In the lung nodule group test set, the diagnostic performance of the dual-modal diagnostic model in differentiating between benign and malignant lung nodules is shown in Table 2: Compared with the traditional method of detecting CEA content to diagnose benign and malignant lung nodules, the AUC value of the dual-modal diagnostic model provided in this embodiment is significantly improved, reaching 0.705.
[0093] Table 2. Comparison of diagnostic performance of CEA and Example 1 bimodal diagnostic model on the pulmonary nodule test set.
[0094]
[0095] However, as shown in Table 2, both the traditional methods for detecting CEA content and the bimodal diagnostic model provided by this invention have very low sensitivity in diagnosing benign and malignant pulmonary nodules. Therefore, the applicant of this invention, based on the bimodal diagnostic model, continued to research, construct, and validate the diagnostic model by combining it with CT imaging features. Specifically, as follows:
[0096] 2.5.3 Constructing a multi-omics differential diagnostic model for benign and malignant pulmonary nodules by combining bimodal diagnostic model scores and CT imaging features.
[0097] (1) Obtain chest CT images of the subjects. 480 patients with pulmonary nodules on CT were selected and randomly assigned to the pulmonary nodule group training set (accounting for 4 / 5 of the total sample) and the pulmonary nodule group test set (accounting for 1 / 5 of the total sample).
[0098] (2) Raw CT image data were input into Image-AI, a CT image-assisted diagnostic model. This model is an artificial intelligence software based on a deep convolutional neural network model for lung nodule detection and classification, developed by Sanmed Biotech and previously clinically validated. 3D-Unet was used for nodule detection and segmentation, while downstream tasks, including nodule type classification and cancer risk score prediction networks, were implemented using 3D ResNet. A total of 480 raw chest CT data were transferred to the workstation, and the software system automatically performed batch lung nodule identification and annotation.
[0099] (3) In the lung nodule group training set, the bimodal diagnostic model score and the Image-AI diagnostic model score obtained in step 2.5.3 were used as joint inputs, and the AdaBoost algorithm was used to train the model to obtain the trimodal initial diagnostic model. The AdaBoost algorithm parameters are as follows: n_estimators: 50, learning_rate: 0.1, algorithm: SAMME.
[0100] (2) The lung nodule test set was tested in a three-modal initial diagnostic model to obtain the three-modal diagnostic model, namely the multi-omics differential diagnostic model for benign and malignant lung nodules described in this invention. The performance of the three-modal diagnostic model in the lung nodule test set is shown in Table 3. Figure 4 As shown in the results, the trimodal diagnostic model constructed in this embodiment significantly outperforms single-modal diagnostic models such as Image-AI and CEA, as well as clinical diagnostic models Mayo and VA. The area under the ROC curve (AUC value) reaches 0.866, and the sensitivity and specificity are 88.73% and 60%, respectively.
[0101] Table 3. Comparison of diagnostic performance of the trimodal diagnostic model with single-modal models such as Image-AI and CEA on the lung nodule test set.
[0102]
[0103] 2.6 Conclusion
[0104] When using the model of this invention to identify benign and malignant lung nodules, the steps are as follows:
[0105] (1) Collect the serum to be tested and perform MALDI-MS detection according to step 2.3 above to obtain the original metabolic fingerprint; perform serum CEA content detection according to step 2.4 above; and obtain chest CT images of the patient to be tested according to step 2.5.3.
[0106] (2) The original metabolic fingerprint of the sample serum was preprocessed according to the above steps to obtain the metabolic fingerprint database of the serum to be tested.
[0107] (3) The patient's CT images are input into the CT image-assisted diagnostic model Image-AI. The obtained score, combined with serum metabolic fingerprint data and serum CEA content, is input into the multi-omics differential diagnostic model for benign and malignant lung nodules based on serum metabolic fingerprint described in this invention to obtain a trimodal diagnostic model score. In practical applications, if the trimodal diagnostic model score of the subject's serum is greater than 0.515, it indicates that the patient's lung nodule has a high probability of malignancy and requires further treatment or follow-up; if the trimodal diagnostic model score of the subject's serum is less than or equal to 0.515, it indicates that the patient's lung nodule has a low risk of malignancy.
[0108] Example 2
[0109] Based on the training and test sets of all subjects in Example 1, this example also uses a deep learning-based neural network to construct a single-modal diagnostic model and a dual-modal diagnostic model, and further uses a random forest to construct a trimodal diagnostic model, specifically:
[0110] (1) The training set model is trained using a neural network to obtain a single-modal initial diagnostic model: the serum metabolic fingerprint is extracted through six feature extraction blocks. Each feature extraction block contains a fully connected layer with 1024 hidden units. After Dropout operation and LeakyReLU activation, the diagnostic score is calculated through a fully connected layer.
[0111] (2) Test the test set in the single-modal initial diagnostic model to obtain the single-modal diagnostic model, which is used to obtain the predicted score of the metabolic molecule scoring module.
[0112] (3) In the training set, the predicted score of the metabolic molecule scoring module obtained in step (2) of this embodiment and the serum CEA content of the sample are used as joint inputs. A fully connected layer is used to calculate the final probability and obtain the dual-modal initial diagnosis model.
[0113] (4) The test set was tested on the bimodal initial diagnostic model to obtain the bimodal diagnostic model. During the training of the bimodal model, binary cross-entropy was used as the loss function to guide gradient descent. The initial learning rates of the Adam optimizer were 0.0001, 0.9β1, and 0.999β2. The entire model training process was carried out for 1000 epochs on an Nvidia GeForce RTX 2070 GPU (Nvidia Corporation, California, USA).
[0114] (5) In the training set of the pulmonary nodule group, ten-fold cross-validation was used to train the bimodal diagnostic model score and the CT image-assisted diagnostic model Image-AI score data, and a trimodal initial diagnostic model was constructed using random forest.
[0115] (6) The lung nodule test set was tested in a three-modal initial diagnostic model to obtain a three-modal diagnostic model, namely the multi-omics differential diagnostic model for benign and malignant lung nodules described in this invention. In the lung nodule test set, the diagnostic model constructed in this embodiment achieved an AUC value of 0.912 in differentiating between benign and malignant lung nodules. At the optimal cutoff value (0.697), the sensitivity and specificity were 81.69% and 92%, respectively. The Image-AI diagnostic model had an AUC value of 0.719, with a sensitivity and specificity of 52.11% and 84.00%, respectively.
[0116] In summary, this invention discloses a multi-omics diagnostic model for the differential diagnosis of benign and malignant pulmonary nodules based on serum metabolic fingerprinting and its construction method. The diagnostic model constructed by this method acquires serum metabolic fingerprinting, CEA content information, and CT image information of the sample, realizing trimodal analysis of metabolomics, tumor protein markers, and CT image features. Compared with traditional single-modal analysis methods, the diagnostic model of this invention significantly improves the sensitivity, accuracy, and detection throughput of the identification of benign and malignant pulmonary nodules. Moreover, the model construction method is simple, convenient, and quick, and is easy to promote and apply in clinical practice.
[0117] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A method for constructing a multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting, characterized in that, Includes the following steps: S1, obtain the raw metabolic fingerprint profiles and CEA protein content of serum samples from lung adenocarcinoma and control serum samples; S2, based on the original metabolic fingerprint profiles of two samples, a single-modal diagnostic system is constructed using machine learning methods; further, by combining the score and CEA content of the single-modal diagnostic system as input, a dual-modal diagnostic system is constructed using machine learning methods; specifically including: The original metabolic fingerprint map is preprocessed to obtain the serum metabolic fingerprint of the sample. The serum metabolic fingerprints of the diseased serum sample and the control serum sample are divided into corresponding training set and test set. Principal component analysis and Pearson correlation analysis were used sequentially to select features from serum metabolic fingerprints on the training set; and support vector machine was further used to build a model on the training set to obtain a single-modality initial diagnostic system. The single-modal initial diagnostic system was tested on the test set to obtain the single-modal diagnostic system. Using the score and CEA content of the single-modal diagnostic system as input, a bimodal initial diagnostic system is constructed on the training set using the Gaussian Naive Bayes algorithm. The dual-modal initial diagnostic system was tested on the test set to obtain the dual-modal diagnostic system. S3: Acquire CT images of patients with pulmonary nodules and input them into the CT image-assisted diagnosis system; S4. Using the scores of the dual-modal diagnostic system obtained in step S2 and the CT image-assisted diagnostic system obtained in step S3 as input, a trimodal diagnostic system is constructed using the AdaBoost algorithm, namely the multi-omics differential diagnostic system for benign and malignant lung nodules based on serum metabolic fingerprints.
2. The method for constructing a multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting as described in claim 1, characterized in that, The CT image-assisted diagnostic system is a risk prediction system for benign and malignant pulmonary nodules based solely on CT image information.
3. The method for constructing a multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting as described in claim 1, characterized in that, In step S1, LDI-MS technology based on nanomatrix materials is used to perform metabolic detection on the lung adenocarcinoma serum sample and the control serum sample to obtain the original metabolic fingerprint profile.
4. The method for constructing a multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting as described in claim 3, characterized in that, The specific steps for obtaining raw metabolic fingerprints using LDI-MS technology based on nanomaterials include: (1) Collect serum samples from patients with lung adenocarcinoma and non-lung adenocarcinoma controls, and prepare nanomatrix materials; (2) The serum sample and the nanomatrix material were diluted with deionized water to prepare the serum sample and nanomatrix suspension to be tested; (3) Spot the serum sample to be tested on the LDI-MS mass spectrometry target plate, dry it at room temperature, then spot the matrix suspension and dry it at room temperature; (4) The serum sample to be tested was detected in LDI-TOF-MS to obtain the original metabolic fingerprint of the serum sample.
5. The method for constructing a multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting as described in claim 4, characterized in that, In step (2), the serum sample was diluted 10 times and the concentration of the nanomatrix suspension was 1 mg / mL.
6. A multi-omics diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting, constructed according to any one of claims 1-5.
7. A method of using the multi-omics differential diagnostic system for benign and malignant pulmonary nodules based on serum metabolic fingerprinting as described in claim 6, characterized in that, Includes the following steps: (1) Take a serum sample to be tested and analyze it using LDI-MS technology based on nanomatrix materials to obtain the original metabolic fingerprint of the serum sample to be tested; at the same time, detect the CEA content in the serum sample to be tested and obtain the patient's chest CT image; (2) Perform chromatographic preprocessing on the original metabolic fingerprint to obtain the serum metabolic fingerprint of the sample; (3) Input the patient's CT images into the CT image-assisted diagnostic system; (4) Input the CEA content, serum metabolic fingerprint and the score of the CT image-assisted diagnostic system in step (3) into the multi-omics differential diagnosis system for benign and malignant lung nodules. The system gives a score of 0-1 based on the probability of malignancy.
8. The method of use as described in claim 7, characterized in that, When the system score is greater than the cutoff value, it indicates that the patient's lung nodule has a high probability of malignancy and requires further treatment or follow-up; when the system score is not higher than the cutoff value, it indicates that the patient's lung nodule has a low risk of malignancy.