A combination of breast cancer metabolic biomarkers and a method for constructing a fingerprint model thereof and applications

The use of LDI-MS and machine learning constructs a serum metabolic fingerprint model for breast cancer diagnosis and prognosis, addressing inefficiencies in current methods by providing rapid, accurate, and reproducible results.

CN114813908BActive Publication Date: 2025-07-15SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210129411.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-07-15
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

Existing breast cancer diagnosis methods rely on large instruments and time-consuming imaging tools, and lack efficient and accurate early diagnosis and prognosis prediction methods.

Method used

Nanoparticle-enhanced laser desorption/ionization mass spectrometry technology combined with machine learning and bioinformatics methods to construct a breast cancer serum metabolic fingerprint map, screen out highly specific metabolic biomarkers, and achieve rapid and accurate diagnosis and prognosis prediction.

Benefits of technology

It has achieved efficient, accurate and rapid diagnosis of breast cancer, high detection reproducibility, fast analysis speed, low sample consumption, and the area AUC under the diagnostic curve is 0.948. The prognostic scoring system effectively predicts the survival of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114813908B_ABST
    Figure CN114813908B_ABST
Patent Text Reader

Abstract

The present invention discloses a breast cancer metabolic biomarker combination and a method for constructing a fingerprint model thereof and applications, relating to the technical field of breast cancer detection, including steps such as preprocessing serum samples and inorganic nanoparticles, performing sample preparation and matrix preparation on a mass spectrometry target plate, collecting serum metabolic fingerprints in a nanoparticle-enhanced laser desorption / ionization mass spectrometer, and performing statistical analysis on the serum metabolic fingerprints, etc., having effects such as high detection reproducibility, fast analysis speed, and minimal sample consumption, and achieving high-efficiency diagnostic performance for breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of breast cancer detection, and particularly to a method for constructing a serum metabolic fingerprint map of breast cancer and its application. Background Art

[0002] Breast cancer is the most common cancer in the world currently, with more than 13 million new cases and 450,000 deaths worldwide every year. High-performance analytical chemistry methods are crucial for the early and effective diagnosis and treatment of breast cancer. Currently, the clinical diagnosis of breast cancer still relies on conventional histopathological classification methods and imaging tools, such as mammography, magnetic resonance imaging, and ultrasound. However, these methods require large instruments or strict operations, are time-consuming, and the results are unreliable.

[0003] Recently, more and more breast cancer studies have emphasized the value of metabolomics in its early diagnosis, treatment, and prognosis prediction. Analytical techniques such as nuclear magnetic resonance spectroscopy (NMR) and mass spectrometry (MS) have been developed for the comprehensive screening of cancer metabolomes. However, the application of MS in liquid and gas phase detection depends on sample purification and metabolite enrichment by chromatography, severely limiting the analysis speed and throughput. Laser desorption / ionization (LDI)-MS uses nanoparticle matrices to replace chromatography for solid-phase detection, which helps to promote the analysis of breast cancer metabolomes. In addition, liquid biopsy using blood samples has the advantages of easy operation, non-invasiveness, and high throughput, and can be used as an effective tool to promote early detection, predict the potential for metastasis, and select treatment methods.

[0004] Therefore, the LDI-MS technology combined with bioinformatics methods has the potential to construct a serum metabolic fingerprint map of breast cancer, screen out breast cancer metabolic biomarkers with strong specificity, and achieve efficient, accurate, and rapid diagnosis and prognosis prediction of breast cancer. Summary of the Invention

[0005] In view of the above-mentioned defects of the prior art, the present invention provides a combination of breast cancer metabolic biomarkers, a method for constructing a fingerprint model thereof, and an application, aiming to achieve efficient, accurate, and rapid diagnosis and prognosis prediction of breast cancer.

[0006] In an embodiment of the present invention, a method for constructing a breast cancer serum metabolic fingerprint model includes the following steps:

[0007] Step 1, pre-treat the serum sample;

[0008] Step 2, pre-treat the inorganic nanoparticles;

[0009] Step 3, perform sample preparation on a mass spectrometry target plate, spot 1 μL of each diluted serum sample, and dry it at room temperature;

[0010] Step 4: Prepare the matrix on the mass spectrometry target plate, spot 1 μL of each matrix solution, and dry at room temperature;

[0011] Step 5, collecting serum metabolic fingerprints in a nanoparticle enhanced laser desorption / ionization mass spectrometer;

[0012] Step 6: Perform machine learning on serum metabolic fingerprints.

[0013] Furthermore, step 1 specifically involves diluting the serum sample 10 times with deionized water.

[0014] Furthermore, step 2 specifically comprises preparing the inorganic nanoparticles into a 1 mg / mL matrix solution with deionized water.

[0015] Furthermore, step 6 specifically includes:

[0016] Step 6.1, preprocess the serum metabolic fingerprint collected in step 5 on MATLAB R2020a to obtain the m / z signal;

[0017] Step 6.2, divide the fingerprint into a training set and a test set;

[0018] Step 6.3: Use the elastic network algorithm on MATLAB R2020a to perform feature selection on the m / z signal on the training set, and set the threshold of the frequency parameter in the elastic network algorithm to obtain the m / z feature;

[0019] Step 6.4: Use the neural network to train the model on the training set on Orange 3.25.0 to obtain the diagnostic performance of the model on the training set.

[0020] Step 6.5: Use the model trained in step 6.4 to make predictions on the test set on Orange 3.25.0 to obtain the diagnostic performance of the neural network on the test set.

[0021] Furthermore, the training set and the test set in step 6.2 both include serum metabolic fingerprints of the breast cancer patient group and the control group.

[0022] Preferably, the threshold value in step 6.3 is 95%.

[0023] The present invention also provides an application of the method for constructing the above-mentioned breast cancer serum metabolic fingerprint model in prognosis prediction, which is characterized in that a multi-factor regression analysis is performed using a proportional hazard regression model on SPSS, specifically comprising the following steps:

[0024] S1: serum metabolic fingerprints of serum samples with complete follow-up information were collected;

[0025] S2: Preprocess the serum metabolic fingerprint on MATLAB R2020a to obtain the m / z signal;

[0026] S3: Divide the serum samples in step 1 into a training set and a test set;

[0027] S4: Use the proportional hazards regression model on SPSS to perform multivariate regression analysis on the m / z signal in the training set, and set the threshold of the p value parameter to obtain the m / z features and their corresponding coefficients. A metabolic prognosis scoring system is constructed by linearly summing up the features and coefficients;

[0028] S5: Score all samples using the scoring system obtained in step 4, and divide all samples into a high-risk group and a low-risk group;

[0029] S7: Use survival analysis on Orange 3.25.0 to verify the survival status of the high-risk group and the low-risk group in the training set and the test set.

[0030] Preferably, the threshold is 0.05.

[0031] Finally, the present invention provides a combination of breast cancer serum metabolic biomarkers, including glyceric acid, nicotinamide, histamine, uracil, thymine, 3,4-diiodopyrazole, and dehydro-phenylalanine.

[0032] The beneficial effects of the present invention compared with other breast cancer detection methods are as follows:

[0033] 1. It has high detection reproducibility (the coefficient of variation (CV) of the feature intensity in serum samples is < 30% for about 95%), fast analysis speed (about 30 seconds per sample), and minimal sample consumption (about 100 nL per sample).

[0034] 2. Through machine learning of the serum metabolic fingerprint, we achieved high-efficiency diagnostic performance for breast cancer, with the area under the diagnostic curve (AUC) being 0.948 (95% confidence interval (CI): 0.922 - 0.973). In addition, the metabolic prognosis scoring system constructed from the serum metabolic fingerprint can effectively predict the prognosis and survival of patients (p < 0.005).

[0035] 3. Developed a diagnostic model based on a combination of 7 metabolic biomarkers, with accurate breast cancer diagnostic efficiency, and the AUC is 0.865 (95% CI: 0.820 - 0.911).

[0036] The following will further illustrate the technical solutions and the resulting technical effects of the present invention in conjunction with the accompanying drawings to fully understand the purpose and effects of the present invention. Description of the Drawings

[0037] Figure 1 are the representative metabolic fingerprints of the sera of breast cancer patients and the control group;

[0038] Figure 2 is the heat map corresponding to 301 m / z signals after preprocessing of the serum metabolic fingerprints of 325 samples;

[0039] Figure 3 is the performance comparison chart of the biomarker combination and individual biomarkers in diagnosing breast cancer;

[0040] Figure 4 is the schematic diagram of the prediction performance of the metabolic prognosis scoring system in the training set. Specific Embodiments

[0041] The preferred embodiments of the present invention are described below with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0042] Embodiment 1. Using nanoparticle enhanced laser desorption / ionization mass spectrometry to collect serum metabolic fingerprints:

[0043] Preparation of instruments and reagents: Nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometry, serum samples, deionized water, matrix (inorganic nanoparticles);

[0044] Step 1: Dilute the serum sample 10 times with deionized water;

[0045] Step 2: Prepare a matrix solution of 1 mg / mL of inorganic nanoparticles with deionized water;

[0046] Step 3: Perform sample preparation on the mass spectrometry target plate, spot 1 μL of each diluted serum sample, and dry at room temperature;

[0047] Step 4: Perform matrix preparation on the mass spectrometry target plate, spot 1 μL of each matrix solution, and dry at room temperature;

[0048] Step 5: Collect serum metabolic fingerprints in a nanoparticle enhanced laser desorption / ionization mass spectrometer;

[0049] Step 6: Obtain serum metabolic fingerprints for subsequent analysis.

[0050] Embodiment 2. Perform machine learning on the serum metabolic fingerprints of breast cancer patients and the control group

[0051] Preparation of instruments and reagents: MATLAB R2020a, Orange 3.25.0;

[0052] Step 1: Through Example 1, serum metabolic fingerprints were collected from 325 serum samples (156 control groups and 169 breast cancer patients), and the results are as shown in the appendix Figure 1 as follows;

[0053] Step 2: Preprocess the 325 serum metabolic fingerprints on MATLAB R2020a, including spectral line smoothing, baseline correction, and spectral peak alignment, to obtain 301 m / z signals, and the results are as shown in the appendix Figure 2 as follows;

[0054] Step 3: Divide the 325 patients into a training set (260 samples, including 135 breast cancer patients and 125 control groups) and a test set (65 samples, including 34 breast cancer patients and 31 control groups);

[0055] Step 4: Use the elastic net algorithm on MATLAB R2020a to perform feature selection on the 301 m / z signals in the training set, and set the threshold of the frequency parameter in the elastic net algorithm to 95% to obtain 36 m / z features;

[0056] Step 5: Use a neural network on Orange 3.25.0 to perform model training on the training set to obtain the diagnostic performance of the model on the training set (Table 1);

[0057] Step 6: Use the trained model in Step 6 on Orange 3.25.0 to perform prediction on the test set to obtain the diagnostic performance of the neural network on the test set (Table 1).

[0058]

[0059] Example 3: Identify breast cancer biomarkers in serum and achieve precise diagnosis of breast cancer through a biomarker combination:

[0060] Preparation of instruments and reagents: Nanoparticle-enhanced laser desorption / ionization time-of-flight mass spectrometry, nanoparticle-enhanced laser desorption / ionization Fourier transform ion cyclotron resonance mass spectrometry, serum samples, deionized water, matrix (inorganic nanoparticles);

[0061] Step 1: Dilute the serum samples 10-fold with deionized water;

[0062] Step 2: Prepare a matrix solution of 1 mg / mL of inorganic nanoparticles with deionized water;

[0063] Step 3: Perform sample preparation on the mass spectrometry target plate, spot 1 μL of each diluted aqueous humor sample, and dry at room temperature;

[0064] Step 4: Perform matrix preparation on the mass spectrometry target plate, spot 1 μL of each matrix solution, and dry at room temperature;

[0065] Step 5: Obtain the accurate molecular weight of breast cancer biomarkers in serum using a matrix-assisted laser desorption / ionization Fourier transform ion cyclotron resonance mass spectrometer;

[0066] Step 6: Obtain the secondary mass spectrometry spectra of breast cancer biomarkers in serum using a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer;

[0067] Step 7: Analyze the mass spectrometry detection results to identify 7 metabolic biomarkers (Table 2);

[0068]

[0069]

[0070] Step 8: Compared with individual biomarkers, the combination of 7 biomarkers shows the best performance in breast cancer diagnosis, and the results are as shown in the appendix Figure 3 as follows.

[0071] Example 4: Perform a proportional hazards regression model analysis on the serum metabolic fingerprints of breast cancer patients and the control group to achieve accurate prognosis prediction of breast cancer:

[0072] Preparation of instruments and reagents: MATLAB R2020a, SPSS 22.0;

[0073] Step 1: Through Example 1, collect serum metabolic fingerprints of 143 serum samples with complete follow-up information;

[0074] Step 2: Preprocess the 143 serum metabolic fingerprints on MATLAB R2020a, including spectral line smoothing, baseline correction, and spectral peak alignment, to obtain 301 m / z signals;

[0075] Step 3: Divide the 143 patients into a training set (96 samples) and a test set (47 samples);

[0076] Step 4: Use the proportional hazards regression model on SPSS to perform a multivariate regression analysis on the 301 m / z signals in the training set, and set the threshold of the p value parameter to 0.05 to obtain 4 m / z features and their corresponding coefficients (Table 3), and a metabolic prognosis scoring system is formed by the linear summation of the features and coefficients.

[0077]

[0078] Step 5: Score all samples using the scoring system obtained in Step 4.

[0079] Step 6: Use the median sample score in the training set as the threshold to divide all samples into a high-risk group (scores higher than the threshold) and a low-risk group (scores lower than the threshold).

[0080] Step 7: Use survival analysis on Orange 3.25.0 to verify that the survival of the high-risk group in both the training set and the test set is worse than that of the low-risk group (p < 0.05), indicating that the metabolic prognosis scoring system can effectively predict the prognosis of breast cancer. The results are as shown Figure 4 below.

[0081] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. Application of a method for constructing a serum metabolic fingerprint model of breast cancer in prognosis prediction, characterized in that, The following steps are involved: Step 1, dilute the serum sample 10 times with deionized water; Step 2, preparing the inorganic nanoparticles into a 1 mg / mL matrix solution with deionized water; Step 3: Prepare the sample on the mass spectrometry target plate, spot 1 μL of each diluted serum sample and dry at room temperature; Step 4: Prepare the matrix on the mass spectrometry target plate, spot 1 μL of each matrix solution, and dry at room temperature; Step 5, collecting serum metabolic fingerprints in a nanoparticle enhanced laser desorption / ionization mass spectrometer; Step 6: Perform machine learning on the serum metabolic fingerprint to build a breast cancer serum metabolic fingerprint model; Step 7, preprocessing the serum metabolic fingerprint of the collected serum samples with complete follow-up information on MATLAB R2020a, including spectral line smoothing, baseline correction and spectral peak matching, to obtain m / z signals; Step 8, dividing the serum samples with complete follow-up information in step 7 into a training set and a test set; Step 9: Perform multivariate regression analysis on the m / z signal in step 7 on the training set using a proportional hazard regression model in SPSS, and set a threshold for the parameter p value to obtain the m / z feature and its corresponding coefficient, and the linear addition of the feature and the coefficient constitutes a metabolic prognostic scoring system; Step 10: Use the scoring system obtained in step 9 to score all samples in step 7 and divide the samples into high-risk group and low-risk group; Step 11: Use survival analysis on Orange 3.25.0 to verify the survival of the high-risk group and the low-risk group in the training set and the test set in step 8.

2. The application according to claim 1, characterized in that, The step 6 specifically includes: Step 6.1, preprocess the serum metabolic fingerprint collected in step 5 on MATLAB R2020a to obtain the m / z signal; Step 6.2, divide the fingerprint into a training set and a test set; Step 6.3: Use the elastic network algorithm on MATLAB R2020a to perform feature selection on the m / z signal on the training set, and set the threshold of the frequency parameter in the elastic network algorithm to obtain the m / z feature; Step 6.4: Use the neural network to train the model on the training set on Orange 3.25.0 to obtain the diagnostic performance of the model on the training set. Step 6.5: Use the model trained in step 6.4 to make predictions on the test set on Orange 3.25.0 to obtain the diagnostic performance of the neural network on the test set.

3. The application according to claim 2, characterized in that The training set and the test set described in step 6.2 both include serum metabolic fingerprints of the breast cancer patient group and the control group.

4. The application according to claim 2, characterized in that, The threshold value in step 6.3 is 95%.

5. The application according to claim 1, characterized in that, The threshold in step 9 is 0.

05.

6. A breast cancer serum metabolic biomarker combination, characterized in that, These include glyceric acid, niacinamide, histamine, uracil, thymine, 3,4-diiodopyrazole, and dehydrophenylalanine.

Citation Information

Patent Citations

  • Ovarian tumor serum marker

    CN101832977A

  • Method for improving mass spectrum grouping stability based on deep learning technology

    CN113433206A