A lung adenocarcinoma multi-omics diagnosis model based on serum metabolic fingerprint and a construction method thereof

CN115472293BActive Publication Date: 2026-09-29SHANGHAI FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211139619.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2026-09-29
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

[0006]为解决现有技术中肺腺癌早期筛查灵敏度和准确性差、检测通量低以及多模态检测模型构建困难等问题,本发明公开了一种基于血清代谢指纹的肺腺癌多组学诊断模型及其构建方法,该模型能实现双模态数据检测,提高了早期肺腺癌筛查的灵敏度和准确度,还解决了传统检测方法检测通量低的问题

Benefits of technology

[0036]1、本发明提供的诊断模型通过获取样本的血清代谢指纹和CEA含量信息,实现代谢组学和蛋白质肿瘤标志物CEA的双模态分析,最终快速、便捷的筛选出肺腺癌高发人群,缩小了肺腺癌筛查的范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472293B_ABST
    Figure CN115472293B_ABST
Patent Text Reader

Abstract

The application discloses a lung adenocarcinoma multi-omics diagnosis model based on serum metabolic fingerprints and a construction method thereof, and the construction method comprises the following steps: detecting diseased serum samples and control serum samples by using a MALDI-MS technology to obtain original metabolic fingerprint spectra of the two kinds of samples; performing spectrum pretreatment on the original fingerprint metabolic spectra to obtain serum metabolic fingerprints; performing machine learning on the serum metabolic fingerprints to construct a single-mode diagnosis model; obtaining CEA protein contents of the two kinds of serum samples; combining single-mode diagnosis model scores and CEA contents, and adopting a machine learning method to construct the lung adenocarcinoma multi-omics diagnosis model. The diagnosis model provided by the application realizes double-mode analysis of metabolomics and a protein tumor marker CEA, greatly improves the sensitivity and accuracy of lung adenocarcinoma screening, and has the advantages of simple model construction method, convenience, low detection cost, large-scale screening, easy clinical popularization and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of analytical technology, specifically relating to a multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting and its construction method. Background Technology

[0002] The prognosis of lung adenocarcinoma is closely related to its stage, and early diagnosis is a prerequisite for improving survival rates. Low-dose spiral CT screening in high-risk groups has reduced the 5-year mortality rate of lung cancer by 24%. However, with the widespread use of CT scans, the problem of overdiagnosis and treatment of accompanying pulmonary nodules has become increasingly prominent, with a false positive rate of 96%. This not only places a heavy psychological burden on patients but also results in a huge waste of national medical and health resources. Therefore, early diagnosis of lung adenocarcinoma and differential diagnosis of benign and malignant pulmonary nodules are pain points and research hotspots in clinical diagnosis and treatment, as well as major needs for the construction of a healthy China and economic development.

[0003] Currently, there are several methods for diagnosing lung cancer in clinical practice: Histopathology is the gold standard for tumor diagnosis, but early-stage lung cancer (in situ and stage IA lung cancer) is often less than 1 cm in diameter, and the lesions often move due to respiration, making localization difficult. Repeated punctures can cause serious trauma and complications, so histopathological testing is not suitable for early diagnosis of lung cancer; Low-dose spiral CT has a high false positive rate and carries the risk of radiation exposure and overdiagnosis, and is only suitable for lung cancer screening in high-risk groups, not for early diagnosis; Clinically routinely used serum tumor markers such as CEA, SCC, Cfra21-1, and NSE have some value in the auxiliary diagnosis and differential diagnosis of tumors, but their sensitivity and specificity for early lung cancer diagnosis are very low when used alone, and the detection rate of combined detection in stage I lung cancer is less than 20%, which is far from meeting clinical needs; Sputum cytology, although convenient, economical, non-invasive, and highly accepted by patients, has very low sensitivity and can only play a suggestive role in the diagnosis of lung cancer.

[0004] Recently, metabolomics has been considered an extension of genomics and proteomics, and the ultimate direction of "omics" research. It bypasses the complex and inefficient regulatory processes within living organisms, providing final, holistic results through the analysis of metabolites, which is its advantage in terms of its enormous application potential in health assessment, disease diagnosis, and efficacy evaluation. However, the development and progression of lung cancer involve complex biological mechanisms, and single-modal data analysis of pathogenic factors still has many limitations. Systematic models that combine serum metabolic fingerprints (SMF) with clinically accessible data (e.g., traditional tumor markers or CT features) will be superior to single biomarkers or a selected set of biomarkers. However, due to the inconsistency in the dimensionality of multiple modalities and the inherent heterogeneity of biological systems, traditional approaches have failed to couple SMF with other data for clinical application.

[0005] Therefore, there is an urgent need in this field to build a comprehensive multimodal platform based on serum metabolic fingerprinting, which is of great significance for precision diagnosis. Summary of the Invention

[0006] To address the problems of poor sensitivity and accuracy, low throughput, and difficulty in constructing multimodal detection models in existing technologies for early screening of lung adenocarcinoma, this invention discloses a multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting and its construction method. This model can achieve dual-modal data detection, improving the sensitivity and accuracy of early lung adenocarcinoma screening and solving the problem of low throughput in traditional detection methods. Furthermore, the model construction method is simple, convenient, and fast, with low detection costs, making it easy to promote and apply clinically.

[0007] To address the above problems, this invention first provides a method for constructing a multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting, comprising the following steps:

[0008] S1, Metabolic detection was performed on diseased serum samples and control serum samples using MALDI-MS technology to obtain the original metabolic fingerprint profiles of the two samples;

[0009] S2, perform pattern preprocessing on the original fingerprint metabolic pattern to obtain the serum metabolic fingerprint of the sample;

[0010] S3, The serum metabolic fingerprint is trained using machine learning methods to construct a single-modal diagnostic model;

[0011] S4, Obtain the CEA protein content of the diseased serum sample and the control serum sample;

[0012] S5. Using the score of the single-modal diagnostic model in step S3 and the CEA content obtained in step S4 as input, a dual-modal diagnostic model is constructed using machine learning methods, namely the multi-omics diagnostic model for lung adenocarcinoma.

[0013] Preferably, the machine learning method includes any one or more of support vector machines, neural networks, or Gaussian Naive Bayes algorithms.

[0014] Preferably, in step S2, the spectral preprocessing includes: noise reduction, curve smoothing, baseline correction, and peak extraction.

[0015] In some embodiments, step S3 specifically includes:

[0016] S3.1, the serum metabolic fingerprints of the diseased serum samples and control serum samples obtained in step S2 are divided into corresponding training sets and test sets;

[0017] S3.2, Principal component analysis and Pearson correlation analysis are used sequentially to select features from serum metabolic fingerprints on the training set to obtain metabolic features; and support vector machine is further used to build a model on the training set to obtain a single-modal initial diagnostic model;

[0018] S3.3, Test the single-modal initial diagnostic model on the test set to obtain the single-modal diagnostic model.

[0019] In some embodiments, step S5 specifically includes:

[0020] S5.1, Using the score of the single-modal diagnostic model in step S3 and the CEA content obtained in step S4 as input, the Gaussian Naive Bayes algorithm is used to construct a dual-modal initial diagnostic model on the training set.

[0021] S5.2, Test the bimodal initial diagnostic model on the test set to obtain the bimodal diagnostic model.

[0022] Preferably, step S1 specifically includes:

[0023] S1.1, Collect serum samples from patients with lung adenocarcinoma and non-lung adenocarcinoma controls, and prepare nanomatrix materials;

[0024] S1.2, The serum sample and the nanomatrix material were diluted with deionized water to prepare the serum sample and nanomatrix suspension to be tested;

[0025] S1.3, Spot the serum sample to be tested on the LDI-MS mass spectrometry target plate, dry it at room temperature, then spot the matrix suspension and dry it at room temperature;

[0026] S1.4, Detect the serum sample to be tested in LDI-TOF-MS to obtain the original metabolic fingerprint of the serum sample.

[0027] Preferably, in step S1.2, the serum sample is diluted 10 times, and the concentration of the nanomatrix suspension is 1 mg / mL.

[0028] In some embodiments, the nanomatrix material includes metallic nanomaterials such as iron nanoparticles, silver nanoparticles, and gold nanoparticles, as well as composite nanomaterials combining metals and inorganic materials, to ensure high specific heat, low thermal conductivity, plasmon resonance effect, good ultraviolet light absorption, and a rough porous structure. The nanomatrix material can be commercially available or prepared in a laboratory.

[0029] In another aspect, the present invention provides a method for constructing a multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprints, which is constructed according to the construction method described in any of the preceding claims.

[0030] In another aspect, the present invention provides a method for using the aforementioned multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting, comprising the following steps:

[0031] (1) Take the serum sample to be tested and analyze it using MALDI-MS technology to obtain the original metabolic fingerprint of the serum sample to be tested; at the same time, detect the CEA content in the serum sample to be tested.

[0032] (2) Perform chromatographic preprocessing on the original metabolic fingerprint to obtain the serum metabolic fingerprint of the sample;

[0033] (3) Input the CEA content and serum metabolic fingerprint of the sample into the lung adenocarcinoma multi-omics diagnostic model. The model gives a score of 0-1 based on the probability of malignancy.

[0034] Furthermore, if the model score is greater than the cutoff value, it indicates that the subject has a risk of developing lung adenocarcinoma and needs further CT examination or surgical biopsy; if the model score is not higher than the cutoff value, it indicates that the subject does not have a risk of developing lung adenocarcinoma and does not need further examination.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] 1. The diagnostic model provided by this invention obtains the serum metabolic fingerprint and CEA content information of the sample, realizes dual-modal analysis of metabolomics and protein tumor marker CEA, and finally quickly and conveniently screens out high-incidence populations of lung adenocarcinoma, thus narrowing the scope of lung adenocarcinoma screening.

[0037] 2. Compared with traditional single-modal analysis methods, the diagnostic model of this invention improves the sensitivity, accuracy and throughput of lung adenocarcinoma screening. Moreover, the model construction method is simple, convenient and quick, which greatly reduces the workload and cost of screening. It is suitable for large-scale screening and easy to promote and apply in clinical practice.

[0038] 3. The serum metabolomics analysis in the diagnostic model of this invention is mostly based on serum samples from patients with early-stage lung adenocarcinoma (over 70%). Therefore, the serum metabolomics data of this invention has a high degree of matching with the serum metabolomics data of the high-risk population to be screened, making it more suitable for early screening of lung adenocarcinoma. Attached Figure Description

[0039] Figure 1 This is a schematic diagram illustrating the construction process of the multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting, as described in this invention.

[0040] Figure 2 The results show the characterization of iron nanoparticles, among which,

[0041] a is the SEM image;

[0042] b is the TEM image;

[0043] c represents dynamic light scattering measurement;

[0044] d represents the optical absorption spectrum of the material.

[0045] Figure 3 This is a schematic diagram of the established local serum metabolic fingerprint database; where,

[0046] a shows typical serum metabolic mass spectrometry of lung adenocarcinoma and benign lung diseases (granulomas) (inset shows H&E staining images of tissues confirmed by pathology);

[0047] b is the serum metabolic fingerprint extracted from the original metabolic fingerprint after preprocessing.

[0048] Figure 4 Visualization of the dimensionality reduction results of the random neighborhood embedding (t-SNE) of the serum metabolic fingerprint t-distribution, where,

[0049] 'a' represents the results in the training set;

[0050] b represents the results in the test set. Detailed Implementation

[0051] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0052] As mentioned above, in view of the shortcomings of the prior art, the applicant of this invention, after long-term research, proposes the technical solution of this invention, the preparation process of which is as follows: Figure 1 As shown: First, MALDI-MS technology is used to capture metabolites from complex biological samples, thereby sensitively and selectively collecting metabolic fingerprints of metabolites (100-1000 Da) to establish a local serum metabolic fingerprint database; then, machine learning methods are used to learn the metabolic fingerprints, construct a single-modal diagnostic model, and output the score of the single-modal diagnostic model; further, by combining the sample clinical indicator CEA content, a dual-modal diagnostic model is constructed through machine learning to achieve early diagnosis of LUAD (lung adenocarcinoma).

[0053] the term

[0054] The “MALDI-MS technology” mentioned in this invention refers to LDI-MS technology based on nanomaterial matrix.

[0055] Example 1

[0056] The experimental process and results of this invention will be described in detail below.

[0057] The construction method and efficacy verification of the multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting of this invention are as follows:

[0058] 1. Research Subjects

[0059] This study included 2276 participants who visited or underwent physical examinations at Shanghai Chest Hospital between November 2016 and May 2018. Among them, 320 were patients with benign lung diseases, 958 were patients with lung adenocarcinoma, and 998 were healthy controls. Benign diseases included pneumonia, chronic obstructive pulmonary disease, and tuberculosis. Lung adenocarcinoma patients were confirmed by histopathology and / or cytopathology, and staging was based on the 8th edition of the TNM staging system. Healthy controls were outpatients undergoing physical examinations. Patients lacking histopathological diagnosis, acute illness history, or other malignant tumors were excluded. All participants signed informed consent forms. This study has completed clinical trial registration (ChiCTR2000036938).

[0060] All participants in this study were randomly assigned to the training set (2 / 3 of the total sample) and the test set (1 / 3 of the total sample). The lung adenocarcinoma patients enrolled in this study were mainly early-stage lung adenocarcinoma patients (stage I and II), accounting for 71.2% (458 / 643) in the training set and 75.2% (237 / 315) in the test set.

[0061] 2. Establishing a serum metabolic fingerprint database using MALDI-MS

[0062] Whole blood samples were collected from subjects after a one-night fast to eliminate dietary interference. Serum was obtained by centrifugation at 3500 rpm for 10 min at 4°C and stored at -80°C. Raw mass spectrometry data were acquired using an Autoflex speed time-of-flight mass spectrometry (Bruker, Germany) instrument.

[0063] 2.1 Instruments and Equipment

[0064] The experimental instruments included: an ultrapure water system (Milli-Q, Millipore, USA), a mass spectrometer (Autoflex SpeedTOF / TOF, Bruker, Germany), a transmission electron microscope (2100F, JEM, Japan), a scanning electron microscope (S-4800, Hitachi, Japan), and a nanoparticle size potentiometer (Mastersizer 3000, Malvern, UK).

[0065] Experimental consumables include: pipettes, 10μL pipette tips, 100μL pipette tips, 1.5mL centrifuge tubes, markers, gloves, and masks.

[0066] 2.2 Preparation and Characterization of Nanoscale Matrix Materials

[0067] (1) Weigh 0.60g of ferric chloride hexahydrate, 0.15g of trisodium citrate and 0.96g of sodium acetate and dissolve them in ethylene glycol solution in sequence. The solution is then sonicated to make it homogeneous.

[0068] (2) The above mixed solution was transferred to a reactor with a capacity of 50 mL and heated to 200 °C for 10 h to obtain trivalent iron nanoparticles.

[0069] (3) The obtained ferric nanoparticles were washed several times with ethanol and deionized water until the supernatant was colorless. The final product was then dried at 60°C for 12 hours and stored in a vacuum for later use.

[0070] To characterize the trivalent iron nanoparticles prepared above, SEM images were obtained using an S-4800 scanning electron microscope; transmission electron microscopy (TEM) images were recorded using a JEM-2100F instrument; dynamic light scattering measurements were performed on a Nano ZS instrument (Malvern, Worcestershire, UK); and optical absorption spectra of the materials were collected on a UV1900 UV-Vis spectrometer (Aucybest, China). The characterization results are as follows: Figure 2 As shown.

[0071] 2.3 Serum MALDI-MS detection

[0072] Metabolic fingerprints of serum samples were obtained using enhanced laser desorption / ionization time-of-flight mass spectrometry (MALDI-MS) based on the trivalent iron nanoparticles prepared above, specifically including the following steps:

[0073] (1) Preparation of nanomatrix suspension: The trivalent iron nanoparticles obtained in step 2.2 were diluted with deionized water to 1 mg / mL;

[0074] (2) Dilute the subject's serum sample 10 times with deionized water;

[0075] (3) Sample preparation on the mass spectrometry target plate: Spot 1 μL of each diluted serum sample or standard and dry at room temperature;

[0076] (4) Matrix preparation on the mass spectrometry target plate: 1 μL of each matrix suspension was spotted and dried at room temperature;

[0077] (5) Serum metabolic fingerprints were collected using LDI-TOF-MS.

[0078] The laser source was an Nd:YAG laser with a wavelength of 355 nm and a maximum frequency of 2 kHz. The mass spectrometry data acquisition mode was set to positive ion reflectance mode, and the molecular weight range of the oligonucleotides to be detected was set to 100 to 1000 Da. During the experiment, the standard parameters were set to a laser frequency of 1000 Hz and a laser intensity of 70%. The data obtained for each experiment were superimposed spectra obtained from 2000 laser shots. Standard molecules were used for mass calibration to ensure accurate mass measurement and avoid intra-plate bias. All serum samples were randomly dropped onto multiple 384-well target plates to reduce systematic errors and inter-plate differences caused by uneven sample type distribution. In addition, five independent experiments were performed to eliminate intra-individual bias in order to enhance the reproducibility and stability of the diagnostic results. During metabolite identification, only signals with a mass spectrometry signal-to-noise ratio (S / N) greater than 3 were used for molecular identification and recognition based on accurate mass alignment. In addition to the precise mass comparison method (±0.05 Da), for selected specific small molecules, the molecular peaks of the secondary mass spectra of their mass spectra (from biological samples and standards) are compared with each other to finally confirm the metabolite to be tested.

[0079] 2.4 Determination of serum tumor markers

[0080] The CEA content of serum samples from 2276 subjects was detected using a Roche Cobas e601 instrument (electrochemiluminescence assay kit). The cutoff values ​​were obtained from the kit instructions.

[0081] 2.5 Construction and Score Calculation of a Multi-omics Diagnostic Model for Lung Adenocarcinoma

[0082] 2.5.1 Model training based on serum metabolic fingerprinting was used to achieve preliminary prediction of lung adenocarcinoma.

[0083] (1) In step 2.3, the serum metabolic fingerprint profiles of 2276 subjects were imaged, and the imaging results are as follows: Figure 3 As shown in Figure a, the upper part of the figure is the fingerprint spectrum of disease samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry; the lower part of the figure is the fingerprint spectrum of samples with benign lung diseases detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry.

[0084] (2) Preprocessing of 2276 serum metabolic fingerprint profiles: First, Gaussian filtering with σ=1 was used for noise reduction and curve smoothing; then, Top-Hat operation was used for baseline correction; finally, local maximum processing was used to extract the final metabolic molecular features. A total of 2316 feature signals in the range of 100-1000 Da were obtained, constituting the serum metabolic fingerprint database described in this invention. The feature signal profiles are shown below. Figure 3 As shown in Figure b, the upper part of the figure is the characteristic signal spectrum of serum samples from 958 patients with lung adenocarcinoma as diseased samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry, while the lower part of the figure is the characteristic signal spectrum of serum samples from 1318 healthy samples and benign lung disease samples as control samples detected by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry.

[0085] (3) The 2276 preprocessed samples were randomly divided into a training set (2 / 3 of the total samples) and a test set (1 / 3 of the total samples). The training set included 643 patients with lung adenocarcinoma, 214 patients with benign lung diseases, and 669 healthy controls. The test set included 315 patients with lung adenocarcinoma, 106 patients with benign lung diseases, and 329 healthy controls. It should also be noted that the lung adenocarcinoma patients enrolled in this study were mainly early-stage lung adenocarcinoma patients (stage I and II), accounting for 71.2% (458 / 643) in the training set and 75.2% (237 / 315) in the test set.

[0086] (4) Perform principal component analysis (PCA) on the training set. Multiple principal components are fitted by PCA principal component analysis. The first 75 principal components (PC1-PC75) are initially selected for further analysis. After Pearson correlation analysis, PC38 is removed, and the remaining 74 principal components are input into support vector machine (SVM) for model training to obtain a single-modal initial diagnostic model.

[0087] Specifically: 10-fold cross-validation was used to train the SVM model on the training set. The SVM parameters were: C = 2.794450000000003, tol = 0.0001, coef0 = 0, kernel = rbf, class_weight = balanced, degree = 3, gamma = auto, probability = True.

[0088] (5) Test the test set in the single-modal initial diagnostic model to obtain the single-modal diagnostic model, which is used to obtain the predicted score of the metabolic molecule scoring module.

[0089] 2.5.2 Constructing a multi-omics diagnostic model for lung adenocarcinoma by combining metabolic molecule scoring module and CEA content.

[0090] (1) In the training set, the predicted score of the metabolic molecule scoring module obtained in step 2.5.1 (i.e. the single-modal diagnostic model score) and the serum CEA content of the sample obtained in step 2.4 are used as joint inputs. The Gaussian Naive Bayes (GaussianNB) algorithm is used to train the model and obtain the dual-modal initial diagnostic model.

[0091] (2) The test set was tested in the bimodal initial diagnostic model to obtain the bimodal diagnostic model, namely the lung adenocarcinoma multi-omics diagnostic model of the present invention. The diagnostic performance of the model in the test set is shown in Table 1: Compared with the traditional method of detecting CEA content, the AUC value of the diagnostic model provided in this embodiment is significantly improved, reaching 0.753.

[0092] Table 1. Comparison of diagnostic performance of CEA and Example 1 diagnostic model on the test set.

[0093]

[0094] 2.6 Comparative Experiment

[0095] For the serum metabolic fingerprints of lung adenocarcinoma patients and control patients (healthy individuals and patients with benign lung diseases) extracted from the training and test sets in this embodiment, the unsupervised method t-distributed random neighborhood embedding (t-SNE) was applied for dimensionality reduction and visualization. The results are as follows: Figure 4 As shown, no clear distinction was found between adenocarcinoma and non-adenocarcinoma in the training and test sets.

[0096] This comparative experiment shows that traditional linear dimensionality reduction methods cannot effectively distinguish between adenocarcinoma and non-adenocarcinoma.

[0097] In addition, this invention also constructs a comparative model that is not based on a single-modal diagnostic model: based on the same training set and test set, 74 principal components selected after principal component analysis and Pearson correlation analysis of the training set are directly combined with the CEA results, and the comparative model is obtained after training with a support vector machine. The diagnostic performance of the comparative model is shown in Table 2. Its area under the ROC curve (AUC value) is lower than that of the diagnostic model constructed by the method of this invention.

[0098] Table 2. Comparison of diagnostic performance of CEA, Example 1 diagnostic model, and comparative model on the test set.

[0099]

[0100] 2.7 Conclusion

[0101] When using the model of this invention for the early diagnosis of lung adenocarcinoma, the steps are as follows:

[0102] (1) Collect serum samples to be tested and perform MALDI-MS detection according to step 2.3 above to obtain the raw metabolic fingerprint; perform serum CEA content detection according to step 2.4 above.

[0103] (2) The original metabolic fingerprint of the sample serum was preprocessed according to the above steps to obtain the metabolic fingerprint database of the serum to be tested.

[0104] (3) Serum metabolic fingerprint data and serum CEA levels are input into the lung adenocarcinoma diagnostic model to obtain a model score, thereby assisting in the diagnosis of lung adenocarcinoma through human metabolic status and serum tumor marker CEA levels. In practical applications, if the subject's serum lung adenocarcinoma diagnostic model score is greater than 0.0368, it indicates that the subject has a risk of developing lung adenocarcinoma and requires further CT examination or surgical biopsy; if the subject's serum lung adenocarcinoma diagnostic model score is less than or equal to 0.0368, it indicates that the subject does not have a risk of developing lung adenocarcinoma and no further examination is required.

[0105] Example 2

[0106] Based on the training and testing sets in Example 1, this example also employs a deep learning-based neural network to construct a single-modal diagnostic model and a dual-modal diagnostic model, specifically as follows:

[0107] (1) The training set was trained using a neural network to obtain a single-modal initial diagnostic model: the serum metabolic fingerprint was extracted through six feature extraction blocks. Each feature extraction block contained a fully connected layer with 1024 hidden units. After Dropout operation and LeakyReLU activation, the diagnostic score was calculated through a fully connected layer.

[0108] (2) Test the test set in the single-modal initial diagnostic model to obtain the single-modal diagnostic model.

[0109] (3) In the training set, the score of the single-modal diagnostic model obtained in step (2) of this embodiment and the serum CEA content of the sample are used as joint inputs. A fully connected layer is used to calculate the final probability and obtain the dual-modal initial diagnostic model.

[0110] (4) Test the test set in the bimodal initial diagnostic model to obtain the bimodal diagnostic model.

[0111] In this embodiment, binary cross-entropy is used as the loss function to guide gradient descent in model construction. The initial learning rates of the Adam optimizer are 0.0001, 0.9β1, and 0.999β2. The entire model training process is carried out for 1000 epochs on an Nvidia GeForce RTX2070 GPU (Nvidia Corporation, California, USA).

[0112] The diagnostic performance of the test set on the training model is shown in Table 3. Compared with the traditional method for detecting CEA content, the diagnostic model constructed in this embodiment not only significantly improves the AUC value to 0.782-0.843, but also achieves excellent results in each group of samples (0-IV stage, 0-II stage).

[0113] Table 3 Comparison of diagnostic performance of CEA and Example 2 diagnostic models on the test set.

[0114]

[0115]

[0116] In summary, this invention provides a multi-omics diagnostic model for lung adenocarcinoma based on serum metabolic fingerprinting and its construction method. The diagnostic model constructed by this method acquires serum metabolic fingerprinting and CEA content information of samples, realizing dual-modal analysis of metabolomics and the protein tumor marker CEA. Compared with traditional single-modal analysis methods, the diagnostic model of this invention improves the sensitivity, accuracy and throughput of lung adenocarcinoma screening. Moreover, the model construction method is simple, convenient and fast, greatly reducing the workload and cost of screening, making it suitable for large-scale screening and easy to promote and apply in clinical practice.

[0117] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

Claims

1. A method for constructing a multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprinting, characterized in that, Includes the following steps: S1, using LDI-MS technology based on nanomatrix materials to perform metabolic detection on diseased serum samples and control serum samples, and obtain the original metabolic fingerprint profiles of the two samples; S2, perform preprocessing on the original metabolic fingerprint to obtain the serum metabolic fingerprint of the sample; S3, A single-modality diagnostic system is constructed by training the serum metabolic fingerprint using machine learning methods; specifically including: S3.1, the serum metabolic fingerprints of the diseased serum samples and control serum samples obtained in step S2 are divided into corresponding training sets and test sets; S3.2, Principal component analysis and Pearson correlation analysis are used sequentially to select features from serum metabolic fingerprints on the training set; and support vector machine is further used to build a model on the training set to obtain a single-modality initial diagnostic system; S3.3, Test the single-modal initial diagnostic system on the test set to obtain the single-modal diagnostic system; S4, Obtain the CEA protein content of the diseased serum sample and the control serum sample; S5, combining the score of the single-modal diagnostic system in step S3 and the CEA content obtained in step S4 as input, a dual-modal diagnostic system is constructed using machine learning methods, thus obtaining the aforementioned multi-omics diagnostic system for lung adenocarcinoma, specifically including: S5.1, Using the score of the single-modal diagnostic system in step S3 and the CEA content obtained in step S4 as input, the Gaussian Naive Bayes algorithm is used to construct the dual-modal initial diagnostic system on the training set. S5.2, Test the dual-modal initial diagnostic system on the test set to obtain the dual-modal diagnostic system.

2. The method for constructing a multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprinting as described in claim 1, characterized in that, In step S2, the spectral preprocessing includes: noise reduction, curve smoothing, baseline correction, and peak extraction.

3. The method for constructing a multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprinting as described in claim 1, characterized in that, Step S1 specifically includes: S1.1, Collect serum samples from patients with lung adenocarcinoma and non-lung adenocarcinoma controls, and prepare nanomatrix materials; S1.2, The serum sample and the nanomatrix material were diluted with deionized water to prepare the serum sample and nanomatrix suspension to be tested; S1.3, Spot the serum sample to be tested on the LDI-MS mass spectrometry target plate, dry it at room temperature, then spot the matrix suspension and dry it at room temperature; S1.4, Detect the serum sample to be tested in LDI-TOF-MS to obtain the original metabolic fingerprint of the serum sample.

4. The method for constructing a multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprinting as described in claim 3, characterized in that, In step S1.2, the serum sample was diluted 10 times, and the concentration of the nanomatrix suspension was 1 mg / mL.

5. A multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprint, constructed by the method according to any one of claims 1-4.

6. A method of using the multi-omics diagnostic system for lung adenocarcinoma based on serum metabolic fingerprinting as described in claim 5, characterized in that, Includes the following steps: (1) Take the serum sample to be tested and analyze it using LDI-MS technology based on nanomatrix materials to obtain the original metabolic fingerprint of the serum sample to be tested; at the same time, detect the CEA content in the serum sample to be tested; (2) Perform chromatographic preprocessing on the original metabolic fingerprint to obtain the serum metabolic fingerprint of the sample; (3) Input the CEA content and serum metabolic fingerprint of the sample into the lung adenocarcinoma multi-omics diagnostic system. The system gives a score of 0-1 based on the probability of malignancy.

7. The method of use as described in claim 6, characterized in that, If the system score is greater than the cutoff value, it indicates that the subject has a risk of developing lung adenocarcinoma and needs further CT examination or surgical biopsy; if the system score is not higher than the cutoff value, it indicates that the subject does not have a risk of developing lung adenocarcinoma and no further examination is required.