Metabolic marker combination for distinguishing benign mammary disease from breast cancer and application of metabolic marker combination

The diagnostic model constructed by combining metabolic biomarkers screened by metabolomics and machine learning algorithms solves the problem of insufficient specificity in breast cancer diagnosis in existing technologies, and achieves early non-invasive diagnosis with high sensitivity and specificity.

CN121899290APending Publication Date: 2026-04-21HARBIN METANOTITIA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN METANOTITIA INC
Filing Date
2025-12-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current technologies lack specificity in distinguishing between benign breast diseases and breast cancer, resulting in a high false positive rate and making it difficult to achieve non-invasive and accurate early diagnosis.

Method used

A combination of metabolic biomarkers, including arginine, triglycerides 56:5, glycoursodeoxycholic acid, and L-prolyl-L-threonine, was screened using metabolomics technology. A diagnostic model was constructed by combining high-throughput analysis and machine learning algorithms to differentiate between benign breast diseases and breast cancer.

Benefits of technology

It provides a highly sensitive and specific non-invasive diagnostic method that can accurately identify breast cancer in its early stages, reduce false positive rates, improve diagnostic efficiency, and is suitable for large-scale population screening and monitoring of high-risk individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121899290A_ABST
    Figure CN121899290A_ABST
Patent Text Reader

Abstract

The invention discloses a metabolic marker combination for distinguishing benign mammary gland diseases and breast cancer and application thereof, and relates to the technical field of biomedicine, the metabolic marker combination for distinguishing benign mammary gland diseases and breast cancer comprises arginine, triglyceride 56: 5, glycine ursodeoxycholic acid and L-prolyl-L-threonine. The metabolic marker combination provided by the invention has extremely high sensitivity and specificity when being used for distinguishing benign breast diseases and breast cancer, is simple to operate, convenient in sample acquisition, relatively low in cost, non-invasive and high in subject compliance, and can realize accurate screening of breast cancer; and early warning and diagnosis on a functional level can be realized in an earlier stage of breast cancer occurrence, and a basis is provided for clinical early intervention. The method effectively solves the problems that the specificity is insufficient and the false positive rate is high when benign and malignant lesions of mammary glands are distinguished in the existing iconography examination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to a combination of metabolic markers for distinguishing between benign breast diseases and breast cancer, and their applications. Background Technology

[0002] Breast cancer (BC) is the most common malignant tumor among women worldwide, seriously threatening their lives and health. Currently, the clinical diagnosis of breast cancer mainly relies on imaging examinations and histopathological biopsies. While imaging examinations (such as mammography / mastography and ultrasound) are widely used as initial screening methods, their specificity is limited, making it difficult to accurately distinguish between benign breast lesions (such as breast hyperplasia and fibroadenoma) and early malignant tumors, often leading to false positive results. This not only causes enormous psychological stress and anxiety for patients but may also trigger unnecessary subsequent invasive examinations. Histopathological biopsy is the "gold standard" for diagnosing breast cancer, but its inherent invasiveness means it cannot be used as a routine monitoring method, and a diagnosis is usually made only after morphological changes have occurred, resulting in a lag in reflecting the real-time functional status of the tumor. Therefore, there is an urgent clinical need for a new, non-invasive, highly sensitive, and highly specific diagnostic method to achieve early detection of breast cancer, optimize clinical decision-making pathways, reduce unnecessary biopsies, and effectively monitor high-risk individuals.

[0003] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of this invention is to provide a combination of metabolic biomarkers for distinguishing between benign breast diseases and breast cancer and their application, aiming to provide metabolic biomarkers with high sensitivity and specificity for distinguishing between benign breast diseases and breast cancer, so as to achieve non-invasive diagnosis and early detection of breast cancer.

[0005] The technical solution of the present invention is as follows: In a first aspect, the present invention provides a combination of metabolic markers for distinguishing between benign breast diseases and breast cancer, wherein the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, and L-prolyl-L-threonine.

[0006] Optionally, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of triglyceride 56:7, triglyceride 54:2, glycine deoxycholic acid, and triglyceride 52:1.

[0007] Optionally, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine.

[0008] Optionally, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:2.

[0009] Optionally, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of dehydroepiandrosterone sulfate, triglyceride 46:2, L-glutamyl-L-serine, triglyceride 44:0, and L-piperidinic acid.

[0010] A second aspect of the present invention provides the use of the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer as described above in the preparation of a product for distinguishing between benign breast diseases and breast cancer.

[0011] Optionally, the samples used by the product in distinguishing between benign breast diseases and breast cancer include at least one of serum, plasma, blood, and dried blood smears.

[0012] Optionally, the product may include reagents or kits.

[0013] Optionally, the kit includes metabolic marker controls and / or metabolic marker standards.

[0014] A third aspect of the present invention provides the application of the combination of metabolic biomarkers described above for distinguishing between benign breast diseases and breast cancer in the construction of a diagnostic model for distinguishing between benign breast diseases and breast cancer.

[0015] Beneficial Effects: The metabolic biomarker combination provided by this invention exhibits extremely high sensitivity and specificity in differentiating benign breast diseases from breast cancer, accurately determining whether a patient has breast cancer. Furthermore, this metabolic biomarker combination is simple to operate, easy to obtain samples, low in cost, non-invasive, and has high subject compliance, enabling precise screening for breast cancer. It can provide functional-level early warning and diagnosis at an earlier stage of breast cancer development, providing a basis for early clinical intervention. This invention effectively solves the problems of insufficient specificity and high false-positive rate in existing imaging examinations when differentiating between benign and malignant breast lesions. Attached Figure Description

[0016] Figure 1 The ROC curve for distinguishing between benign breast diseases and breast cancer based on a combination of 21 metabolic markers in the validation group of Example 2.

[0017] Figure 2 The ROC curve for distinguishing between benign breast diseases and breast cancer based on a combination of 16 metabolic markers in the validation group of Example 3.

[0018] Figure 3 The ROC curve for distinguishing between benign breast diseases and breast cancer based on a combination of 12 metabolic markers in the validation group of Example 4.

[0019] Figure 4 The ROC curve for distinguishing between benign breast diseases and breast cancer based on a combination of eight metabolic markers in the validation group of Example 5.

[0020] Figure 5 The ROC curve for distinguishing between benign breast diseases and breast cancer based on a combination of four metabolic markers in the validation group of Example 6. Detailed Implementation

[0021] This invention provides a combination of metabolic biomarkers for distinguishing between benign breast diseases and breast cancer, and its application. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0023] The following explanation covers some of the metabolic markers mentioned in the embodiments and implementation examples.

[0024] The 56:7 in triglyceride 56:7 refers to the fact that the total number of carbon atoms in the fatty acid chain of the glycerol backbone is 56, containing 7 unsaturated double bonds.

[0025] The 56:5 in triglyceride 56:5 refers to the fact that the total number of carbon atoms in the fatty acid chain of the glycerol backbone is 56, containing 5 unsaturated double bonds.

[0026] The 54:2 ratio in triglyceride 54:2 refers to the fact that the total number of carbon atoms in the fatty acid chain of the glycerol backbone is 54, containing 2 unsaturated double bonds.

[0027] The 52:1 ratio in the triglyceride ratio 52:1 refers to the fact that the total number of carbon atoms in the fatty acid chain of the glycerol backbone is 52, and it contains one unsaturated double bond.

[0028] The 46:2 ratio in triglyceride 46:2 refers to the fact that the total number of carbon atoms in the fatty acid chain of the glycerol backbone is 46, containing 2 unsaturated double bonds.

[0029] The 44:0 in triglyceride 44:0 means that the total number of carbon atoms in the fatty acid chain on the glycerol backbone is 44, and it does not contain any unsaturated double bonds, that is, it is completely saturated.

[0030] The 19:2 in phosphatidylcholine 19:2 refers to the fact that the total number of carbon atoms in the fatty acid chain of the phosphatidylcholine molecule is 19, and it contains 2 unsaturated double bonds.

[0031] In phosphatidylethanolamine 36:6e, 36:6 ​​refers to the total number of carbon atoms in the fatty acid chain of the phosphatidylethanolamine molecule being 36, containing 6 unsaturated double bonds, and 'e' indicates that the lipid is an ether-bonded lipid, meaning that at least one fatty acid chain in the molecule is connected to the glycerol backbone through an ether bond (rather than an ester bond).

[0032] The pathogenesis of breast cancer is biologically a continuous lineage progression. The transformation from benign to malignant is a multi-step, multi-gene-involved, gradual evolution, following a model of "normal breast epithelium → proliferative lesions → dysplasia → carcinoma in situ → invasive carcinoma." Not all benign lesions will progress to cancer, but dysplasia is considered a key precancerous lesion, where cells exhibit atypia, significantly increasing the risk of developing cancer to 4-5 times that of the general population. A critical window for intervention and diagnosis exists during the transition from "dysplasia" to "carcinoma in situ / invasive carcinoma." Current technologies struggle to effectively identify lesions based on morphological changes at this stage, highlighting the urgent need to develop novel diagnostic tools capable of early warning and accurate identification.

[0033] Metabolomics, an important branch of systems biology, reveals the overall physiological and pathological state of an organism by qualitatively and quantitatively analyzing the dynamic changes of all small-molecule metabolites (typically <1000 Da) under specific physiological or pathological conditions. Metabolites are the final products of gene expression and protein function, directly reflecting the functional phenotype of an organism. The occurrence and development of tumors are accompanied by the reprogramming of cellular metabolism (such as the Wahlberg effect), and these alterations in metabolic pathways directly lead to significant changes in the endogenous metabolite profile in bodily fluids such as blood and urine. Therefore, metabolomics can sensitively capture early functional changes during tumor development. Metabolic disorders often occur before morphological changes. By detecting minute fluctuations in metabolites, it is hoped that early warnings can be issued in the very early stages of tumors or even precancerous lesions. Furthermore, metabolomics research samples (such as blood and urine) are easy to obtain, the sampling process is non-invasive or minimally invasive, patient compliance is good, and it is suitable for large-scale population screening, regular follow-up of high-risk groups, and dynamic monitoring of treatment efficacy. Most importantly, metabolomics analyzes metabolic networks holistically from a systems biology perspective, which helps to discover combinations of metabolic biomarkers with diagnostic value. Such combinations typically have higher diagnostic accuracy and reliability than single biomarkers.

[0034] In this invention, metabolomics technology is applied to the differential diagnosis of breast cancer, effectively compensating for the shortcomings of existing diagnostic techniques. By combining high-throughput analysis with multivariate statistical analysis and machine learning algorithms, a set of characteristic metabolic biomarkers that can significantly distinguish between benign breast diseases and breast cancer can be screened. Diagnostic models or kits developed based on this biomarker combination can provide clinicians with an objective and quantitative auxiliary diagnostic tool, helping to reduce unnecessary biopsies, improve diagnostic efficiency and accuracy, and provide new and important evidence for individualized treatment management and prognostic assessment of patients. Specifically, this invention provides a combination of metabolic biomarkers for distinguishing between benign breast diseases and breast cancer, wherein the combination of metabolic biomarkers for distinguishing between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, and L-prolyl-L-threonine.

[0035] The metabolic biomarker combination provided by this invention exhibits extremely high sensitivity and specificity in differentiating benign breast diseases from breast cancer, accurately determining whether a patient has breast cancer. Furthermore, this combination is simple to operate, easy to obtain samples, low in cost, non-invasive, and has high subject compliance, enabling precise screening for breast cancer. It allows for functional-level early warning and diagnosis at an earlier stage of breast cancer development, providing a basis for early clinical intervention. This invention effectively solves the problems of insufficient specificity and high false-positive rates in existing imaging examinations when differentiating between benign and malignant breast lesions.

[0036] Specifically, the combination of metabolic biomarkers provided by this invention has the following advantages: (1) High accuracy and reliability: This invention is based on a set of screened metabolic markers to determine the benign and malignant lesions of the breast, rather than relying on a single indicator. It can reflect the tumor-specific metabolic disorders at the system level, thereby significantly improving the sensitivity and specificity of differential diagnosis.

[0037] (2) Non-invasive and high compliance: The detection method of the present invention is based on easily obtainable biological samples such as blood. It is non-invasive, greatly reduces the pain and psychological burden of patients, and is more easily accepted. It is suitable for large-scale population screening and repeated dynamic monitoring of high-risk individuals.

[0038] (3) Prospective and early warning capabilities: Since metabolic reprogramming is an early event of cancer, this invention can detect metabolic abnormalities before morphological changes are obvious, thereby achieving an earlier diagnosis than imaging and winning valuable treatment time for patients.

[0039] (4) Easy to standardize and promote: Based on modern high-throughput analysis technology (such as mass spectrometry), the detection method of the present invention is easy to automate and standardize, which is conducive to its promotion and application in clinical laboratories and the formation of stable and reliable diagnostic products.

[0040] In some embodiments, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of triglycerides 56:7, triglycerides 54:2, glycosaminoglycans, and triglycerides 52:1. That is, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycosaminoglycans, and L-prolyl-L-threonine, and at least one of triglycerides 56:7, triglycerides 54:2, glycosaminoglycans, and triglycerides 52:1.

[0041] The combination of metabolic markers provided by this invention has extremely high sensitivity and specificity in distinguishing between benign breast diseases and breast cancer, and can accurately determine whether a patient has breast cancer.

[0042] In some embodiments, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine. That is, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, and triglycerides 52:1, and at least one of 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine.

[0043] The combination of metabolic markers provided by this invention has extremely high sensitivity and specificity in distinguishing between benign breast diseases and breast cancer, and can accurately determine whether a patient has breast cancer.

[0044] In some embodiments, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:2.

[0045] In other words, the combination of metabolic markers used to distinguish between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, triglycerides 52:1, 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine, as well as at least one of uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:2.

[0046] The combination of metabolic markers provided by this invention has extremely high sensitivity and specificity in distinguishing between benign breast diseases and breast cancer, and can accurately determine whether a patient has breast cancer.

[0047] In some embodiments, the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer further includes at least one of dehydroepiandrosterone sulfate, triglyceride 46:2, L-glutamyl-L-serine, triglyceride 44:0, and L-piperidinic acid.

[0048] In other words, the combination of metabolic markers used to distinguish between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, triglycerides 52:1, 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, hypoxanthine, uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:2, as well as at least one of dehydroepiandrosterone sulfate, triglycerides 46:2, L-glutamyl-L-serine, triglycerides 44:0, and L-piperidinic acid.

[0049] The combination of metabolic markers provided by this invention has extremely high sensitivity and specificity in distinguishing between benign breast diseases and breast cancer, and can accurately determine whether a patient has breast cancer.

[0050] This invention also provides an application of the combination of metabolic markers described above for distinguishing between benign breast diseases and breast cancer in the preparation of products for distinguishing between benign breast diseases and breast cancer.

[0051] The combination of metabolic markers provided by this invention can effectively distinguish between benign breast diseases and breast cancer, with high accuracy, high sensitivity and high specificity. Therefore, it can be used to prepare products for distinguishing between benign breast diseases and breast cancer.

[0052] In some embodiments, the samples used by the product to differentiate between benign breast diseases and breast cancer include at least one of serum, plasma, blood, and dried blood smears.

[0053] In some embodiments, the product includes reagents or kits.

[0054] In some embodiments, the kit includes metabolic marker controls and / or metabolic marker standards.

[0055] This invention also provides an application of the combination of metabolic biomarkers described above for distinguishing between benign breast diseases and breast cancer in constructing a diagnostic model for distinguishing between benign breast diseases and breast cancer.

[0056] In some implementations, the method for constructing a diagnostic model to distinguish between benign breast diseases and breast cancer includes the following steps: Using the relative abundance values ​​of biomarkers as input and the disease type (benign breast disease or breast cancer) as output, the model is trained using the Support Vector Machine (SVM) machine learning algorithm to obtain a diagnostic model for distinguishing between benign breast diseases and breast cancer.

[0057] The model training process utilizes sample data from the modeling group, randomly divided into training and test sets. During training, 1000 randomized cyclic cross-validations are performed; in each cycle, the training and test sets are randomly re-split, and the model training and prediction process is repeated. This repeated validation effectively reduces the randomness of the random splits, improving the model's robustness and reliability. In each cross-validation, the model's classification accuracy is calculated based on the prediction results on the test set. The average accuracy results from 1000 randomized cyclic cross-validations are used as a comprehensive evaluation index of the diagnostic model's overall discriminative ability.

[0058] In some implementations, the constructed diagnostic model is validated by inputting validation group sample data that has never participated in any training or iterative validation process into the constructed diagnostic model, and calculating its discrimination sensitivity and specificity, among other indicators. These indicators are unbiased estimates of the diagnostic model's ability to predict new external samples. The invention will be further illustrated below with specific embodiments.

[0059] Example 1 1. Subject Information 1) The inclusion and exclusion criteria for breast cancer patients are as follows: Inclusion criteria (all of the following conditions must be met for inclusion): (1) Females aged ≥18 years; (2) Read and fully understand the information, sign the informed consent form, and be able to provide a blood sample for metabolomics testing; (3) Primary breast cancer is diagnosed by biopsy or postoperative pathology or by comprehensive evaluation by a clinician.

[0060] Exclusion criteria (meeting any of the following criteria will result in exclusion): (1) During pregnancy or lactation; (2) Emergency room visit or resuscitation required; (3) History of blood transfusion within 7 days prior to sampling; (4) Has received an organ transplant or allogeneic bone marrow or stem cell transplant; (5) History of other malignant tumors within 5 years, or having received any anti-tumor treatment before sampling; (6) Simultaneous co-occurrence of multiple primary malignant tumors.

[0061] 2) Inclusion and exclusion criteria for patients with benign breast diseases: Inclusion criteria (all of the following conditions must be met for inclusion): (1) Females aged ≥18 years; (2) Read and fully understand the information, sign the informed consent form, and be able to provide a blood sample for metabolomics testing; (3) After biopsy or postoperative pathology or comprehensive evaluation by clinicians, the diagnosis is confirmed as benign breast disease (including but not limited to breast hyperplasia, breast cysts, breast duct ectasia, fibroadenoma, etc.), and breast cancer is excluded.

[0062] Exclusion criteria (meeting any of the following criteria will result in exclusion): (1) During pregnancy or lactation; (2) Emergency room visit or resuscitation required; (3) History of blood transfusion within 7 days prior to sampling; (4) Has received an organ transplant or allogeneic bone marrow or stem cell transplant; (5) Any history of malignant tumors or any anti-tumor treatment prior to sampling.

[0063] 2. Sample Information This study collected plasma samples from 205 participants across three medical centers, including 136 participants with breast cancer (BC) and 69 participants with benign breast disease (BBD).

[0064] Plasma samples from 136 breast cancer patients were randomly assigned to a modeling group and a validation group, and plasma samples from 69 patients with benign breast diseases were also randomly assigned to a modeling group and a validation group (i.e., the samples from the modeling group and the validation group were different). The modeling group consisted of plasma samples from 102 breast cancer patients and 52 benign breast disease patients; the validation group consisted of plasma samples from 34 breast cancer patients and 17 benign breast disease patients (see Table 1).

[0065] Table 1. Sample Information

[0066] 3. Plasma sample pretreatment and metabolite detection (1) Reagents: Methanol, acetonitrile, water, acetic acid, isopropanol, and methyl tert-butyl ether of mass spectrometry grade, and formic acid and ammonium acetate of chromatographic (HPLC) grade were all purchased from Sigma-Aldrich, USA.

[0067] (2) Sample preprocessing: After being removed from the -80°C freezer, the plasma samples were thawed on ice and vortexed for 10 seconds. Then, 100 μL of plasma was added to a pre-cooled 1000 μL mixed solution (composed of methyl tert-butyl ether and methanol in a 3:1 volume ratio) and vortexed to obtain the sample extract. Next, 500 μL of the mixed solution (composed of methanol and water in a 3:1 volume ratio) was added to the sample extract, followed by sonication, standing, vortexing, and centrifugation to separate the layers. The upper layer was the organic phase, and the lower layer was the aqueous phase.

[0068] Organic phase processing: After sample separation, take the upper 500 μL organic phase into a centrifuge tube, dry it using a vacuum concentrator (Speed-Vac), add 200 μL of a mixed solution (composed of acetonitrile and isopropanol in a volume ratio of 3:1) to reconstitute, and incubate at room temperature for 15 minutes; vortex the centrifuge tube to mix, sonicate for 5 minutes, and centrifuge at room temperature for 5 minutes (12000 rpm); after centrifugation, take 180 μL of the supernatant into a 2 mL glass vial, which is used as organic phase metabolite (or lipid metabolite) for detection by LC-MS (liquid chromatography-mass spectrometry).

[0069] Aqueous phase processing: After sample separation, the lower 400 μL aqueous phase was transferred to a centrifuge tube, and 1100 μL of ice-cold methanol was added to precipitate proteins. After vortexing and centrifugation, 1000 μL of the supernatant was transferred to a centrifuge tube, concentrated and dried, and then reconstituted with 200 μL of mass spectrometry-grade water at room temperature for 15 min. After incubation, the mixture was vortexed and sonicated for 5 min, followed by centrifugation at 12000 rpm for 5 min at room temperature. Finally, 180 μL of the supernatant was transferred to a 2 mL glass vial for analysis as an aqueous metabolite using LC-MS.

[0070] (3) High-resolution liquid chromatography-mass spectrometry (UHPLC-MS) detection: Chromatographic parameters of organic phase metabolites: Organic phase metabolites were separated into small molecules using a Waters ACQUTTY UPLC® BEH C8 (2.1 mm × 100 mm, 1.7 µm) column at a column temperature of 60 °C. Liquid chromatography and mass spectrometry were performed using an ACQUITY UPLC I-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific), respectively.

[0071] The composition of the organic phase metabolite mobile phase is as follows: Mobile phase A was an aqueous solution containing 0.1% (w / w) acetic acid and 10 mmol / L ammonium acetate; mobile phase B was a mixed solution of acetonitrile and isopropanol containing 0.1% (w / w) acetic acid and 10 mmol / L ammonium acetate (the volume ratio of acetonitrile to isopropanol was 7:3), and the flow rate was 0.4 mL / min.

[0072] The elution gradient for organic phase metabolite separation is as follows: 0-1 min, 55% mobile phase B; 1-4 min, 55-75% mobile phase B; 4-12 min, 75-89% mobile phase B; 12-15 min, 89-100% mobile phase B; 15-19.5 min, 100% mobile phase B; 19.5-19.51 min, 100-55% mobile phase B; 19.51 min-24 min, 55% mobile phase B; injection volume 2 µL, autosampler temperature 10 °C.

[0073] Aqueous phase metabolite chromatographic parameters: Aqueous metabolite analysis was performed using a Waters ACQUTTY UPLC® HSS T3 column (2.1 mm × 100 mm, 1.8 µm) for small molecule separation at a column temperature of 40 °C; liquid chromatography and mass spectrometry were performed using an ACQUITY UPLC I-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific), respectively.

[0074] The composition of the aqueous metabolite mobile phase is as follows: Mobile phase A is an aqueous solution containing 0.1% (mass content) formic acid; mobile phase B is an acetonitrile solution containing 0.1% (mass content) formic acid, and the flow rate is 0.4 mL / min.

[0075] The elution gradient for aqueous metabolite separation is as follows: 0-1 min, 1% mobile phase B; 1-11 min, 1-40% mobile phase B; 11-13 min, 40-70% mobile phase B; 13-15 min, 70-99% mobile phase B; 15-18 min, 99% mobile phase B; 18-19 min, 99-1% mobile phase B; 19-22 min, 1% mobile phase B; sample injection volume is 3 µL, autosampler temperature is 10 °C.

[0076] LC-MS mass spectrometry parameters: Both aqueous and organic metabolites were acquired using full scan and data-dependent acquisition (DDA) for both MS1 and MS2 mass spectrometry. The full scan mass spectrometry range was 100–1500 Da. The MS2 scan ranges (Full MS / dd-MS2) were 100–310 Da, 300–710 Da, and 700–1500 Da. The MS instrument used was an Orbitrap high-resolution mass spectrometer (Thermo Fisher) equipped with an electrospray ionization (ESI) source, employing both positive and negative ionization modes for data acquisition. Specific parameters were as follows: In full scan mode, the automatic gain control (AGC) was 3 × 10⁻⁶. 6 The maximum ion implantation time (Maximum IT) is 200 ms, and the full scan resolution is 70,000 FWHM (Full Width at Half Maximum) (@200 m / z). In the second-stage scan mode (Full MS / dd-MS2), the resolution of the second-stage mass spectrometer is 17,500 FWHM, the quadrupole isolation window width is 1.5 m / z, and the AGC is 1 × 10⁻⁶. 5 The maximum ion implantation time is 50 ms, and the relative collision energy (HCD) is 30%. In addition, the ion spray voltage is 3500 V in positive mode and 3000 V in negative mode; the atomizing gas pressure is 20 psi, the sheath gas temperature is 400 °C, and the sheath gas flow rate is 10 L / min.

[0077] 4. Metabolomics data preprocessing and metabolite identification 1) Metabolomics data preprocessing: Metabolomics data preprocessing includes steps such as peak extraction, peak alignment, peak filtering, and missing value imputation, specifically including: (1) Convert the raw mass spectrometry data into an mzXML format file and use OpenMS software to perform peak extraction, peak alignment and other processing to convert the mass spectrometry data into a data matrix.

[0078] (2) Apply different filtering criteria to remove the following interfering peaks: (i) isotope peaks; (ii) fragments caused by the ionization method within the analyte source; (iii) redundant peaks, such as additional low-intensity adducts of the same analyte and redundant derivatives, to ensure the quality of the analytable dataset and remove characteristic peaks with a missing rate of >50% in all samples.

[0079] (3) Missing values ​​were filled using the MICE Forest chain equation. To remove batch effects, a normalization autoencoder (NormAE) was used for normalization.

[0080] 2) Identification of metabolites: After analyzing the raw data using software, the spectral information of the parent ion and secondary fragment ions of the compound was obtained, such as the mass-to-charge ratio (m / z) of the primary mass spectrometer, the fragments of the secondary ions, and the retention time. The metabolites were then qualitatively identified by comparing the spectral information with that in public databases. Commonly used public metabolite databases include HMDB (www.hmdb.ca), PubChem (https: / / pubchem.ncbi.nlm.nih.gov), MassBank (http: / / www.massbank.jp), MassBank of North America (https: / / massbank.us), and lipid databases such as Lipidmap (https: / / www.lipidmaps.org) and Lipidblast (https: / / fiehnlab.ucdavis.edu / projects / LipidBlast). Based on the initially identified metabolites, final validation was performed using the retention times, MS1 and MS2 mass spectrometry data of standards separated under the same chromatographic column and mass spectrometry conditions. The criteria for metabolite identification are: retention time within 0.1 min, and deviation between theoretical and measured molecular weight of metabolite less than 10 ppm.

[0081] 5. Data Analysis Results Metabolic biomarker screening: The Least Absolute Shrinkage and Selection Operator (Lasso) regression feature screening method was used to analyze plasma samples from subjects in the benign breast disease group and the breast cancer group. Finally, 21 differential metabolites were screened as metabolic biomarkers to distinguish between benign breast disease and breast cancer (see Table 2).

[0082] Table 2. 21 metabolic markers used to differentiate between benign breast diseases and breast cancer

[0083] Example 2: A diagnostic model for differentiating benign breast diseases from breast cancer was constructed using a combination of 21 metabolic biomarkers. To evaluate the diagnostic efficacy of the 21 metabolic biomarker combinations selected in Example 1 in distinguishing between benign breast diseases and breast cancer, multivariate ROC curve analysis was performed on these 21 combinations in the modeling group of Example 1 to construct a diagnostic model for differentiating between benign breast diseases and breast cancer. Specifically, three-quarters of the sample data from the benign breast disease group and three-quarters of the sample data from the breast cancer group in the modeling group of Example 1 were randomly selected as the training set to construct and train the machine learning classification model; while the remaining one-quarter of the sample data from the benign breast disease group and the remaining one-quarter of the sample data from the breast cancer group were used as the test set to verify the discriminative ability of the trained model.

[0084] The relative abundance values ​​of 21 metabolic biomarkers in the training set were used as input features, and the corresponding disease group labels (benign breast disease or breast cancer) were used as supervision labels. An SVM algorithm was used to learn the comprehensive discriminative weights of each metabolic biomarker in distinguishing between the two disease states, thus establishing a classification model. During model training, 1000 randomized cyclic cross-validations were used. In each cycle, the training and test sets were randomly re-split, and the model training and prediction process was repeated. Multiple repeated validations effectively reduced the randomness of random partitioning and improved the robustness and reliability of the model. In each cross-validation, the classification accuracy of the model was calculated based on the prediction results of the test set samples. The average accuracy results obtained from 1000 randomized cyclic cross-validations were used as a comprehensive evaluation index of the model's overall discriminative ability. The model output is the classification result of whether the sample in the test set belongs to breast cancer or benign breast disease, and also outputs the predicted probability value of the sample being classified as breast cancer. A multivariate ROC curve was constructed, and the area under the curve (AUC) was calculated to evaluate the diagnostic efficacy of the combination of 21 metabolic biomarkers in distinguishing between benign breast disease and breast cancer.

[0085] ROC curve (Receiver Operating Characteristic) analysis is based on a series of different binary classification methods (cutoff values ​​or decision thresholds), plotting a curve with sensitivity (true positive rate) on the ordinate and 1-specificity (false positive rate) on the x-axis. The closer the ROC curve is to the upper left corner, the higher the diagnostic accuracy of the biomarker. The area under the ROC curve (AUC) of the metabolic biomarker combination can be used to determine the diagnostic value of potential biomarkers. The AUC value is a key indicator of model performance. When the AUC is greater than 0.5, the closer it is to 1, the better the model performance and the better the judgment effect. If it is less than 0.5, it indicates that the model performance is close to random guessing, with poor discrimination ability and poor accuracy. In addition to AUC, the performance evaluation of ROC classification models also includes indicators such as sensitivity and specificity.

[0086] The formula for calculating sensitivity is:

[0087] The specificity calculation formula is as follows:

[0088] in: TP (True Positive): The number of samples that are actually positive but were correctly predicted as positive. TN (True Negative): The number of samples that are actually negative but are correctly predicted as negative. FP (False Positive): The number of samples that are actually negative but are incorrectly predicted as positive. FN (False Negative): The number of samples that are actually positive but are incorrectly predicted as negative.

[0089] The results showed that the diagnostic model constructed based on the combination of 21 metabolic biomarkers had an AUC of 0.885 (sensitivity: 76.9%, specificity: 76.9%), indicating that the constructed diagnostic model has high diagnostic efficacy.

[0090] To further validate the predictive performance of the diagnostic model for distinguishing between benign breast diseases and breast cancer, built based on the modeling group, the validation set dataset from Example 1 was used as unknown samples and placed into the diagnostic model constructed using the modeling group. The independent validation performance of the diagnostic model on unknown datasets other than the modeling set dataset was evaluated. Each sample in the validation set outputs a diagnostic threshold, resulting in a confusion matrix (including true positive, true negative, false positive, and false negative). Sensitivity and specificity can be calculated using formulas. The matrix is ​​plotted with sensitivity as the ordinate, and 1... Specificity (1) In a ROC analysis graph with specificity on the x-axis, the true positive rate and false positive rate values ​​corresponding to each sample in the ROC curve are connected to form the final ROC curve. The closer the ROC curve is to the upper left corner, the higher the diagnostic accuracy of the metabolic marker. The point with the best sensitivity and specificity is selected and is called the optimal diagnostic threshold.

[0091] In the predictive model constructed based on a combination of 21 metabolic biomarkers, the diagnostic threshold was 0.6651 (above this threshold, the diagnosis was breast cancer; below this threshold, the diagnosis was benign breast disease). Table 3 shows the confusion matrix results, confirming the model's good discriminative ability. Among 34 breast cancer patients, 26 were diagnosed with breast cancer, and 8 were misdiagnosed as having benign breast disease; among 17 patients with benign breast disease, 14 were diagnosed with benign breast disease, and 3 were misdiagnosed as having breast cancer. The ROC analysis results of the diagnostic model in the validation group are as follows: Figure 1 As shown, sensitivity and specificity were calculated based on the confusion matrix results, with an AUC of 0.856 (sensitivity: 76.5%, specificity: 82.4%). These results indicate that the constructed diagnostic model for distinguishing between benign breast diseases and breast cancer also demonstrated good diagnostic performance in the validation group.

[0092] Example 3: Constructing a diagnostic model to differentiate between benign breast diseases and breast cancer using a combination of 16 metabolic biomarkers. Using the modeling group data from Example 1, a diagnostic model was constructed using a combination of 16 metabolic biomarkers. The only difference from Example 2 was that the diagnostic model was constructed using a combination of 16 metabolic biomarkers: arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, triglycerides 52:1, 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, hypoxanthine, uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:2.

[0093] The ROC results showed an AUC of 0.861 (sensitivity: 72%, specificity: 84.6%). This indicates that the constructed diagnostic model has high diagnostic efficacy.

[0094] To further verify the effectiveness of the diagnostic model for distinguishing between benign breast diseases and breast cancer constructed based on the modeling group data, validation was performed in the validation group of Example 1. The ROC results are as follows: Figure 2 As shown, the results indicate that in the validation group, AUC = 0.830 (sensitivity: 70.6%, specificity: 77.8%).

[0095] The results above indicate that the diagnostic model constructed to distinguish between benign breast diseases and breast cancer also has good diagnostic performance in the validation group.

[0096] Example 4: A diagnostic model for differentiating benign breast diseases from breast cancer was constructed using a combination of 12 metabolic biomarkers. Using the modeling group data from Example 1, a diagnostic model was constructed using a combination of 12 metabolic biomarkers. The only difference from Example 2 was that the diagnostic model was constructed using a combination of 12 metabolic biomarkers: arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, triglycerides 52:1, 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine.

[0097] The ROC results showed an AUC of 0.83 (sensitivity: 76%, specificity: 76.9%). This indicates that the constructed diagnostic model has high diagnostic efficacy.

[0098] To further verify the effectiveness of the diagnostic model for distinguishing between benign breast diseases and breast cancer constructed based on the modeling group data, validation was performed in the validation group of Example 1. The ROC results are as follows: Figure 3 As shown, the results indicate that in the validation group, AUC = 0.811 (sensitivity: 73.5%, specificity: 82.4%).

[0099] The results above indicate that the diagnostic model constructed to distinguish between benign breast diseases and breast cancer also has good diagnostic performance in the validation group.

[0100] Example 5: A diagnostic model for differentiating benign breast diseases from breast cancer was constructed using a combination of eight metabolic biomarkers. Using the modeling group data from Example 1, a diagnostic model was constructed using a combination of 8 metabolic biomarkers. The only difference from Example 2 was that the diagnostic model was constructed using a combination of 8 metabolic biomarkers: arginine, triglycerides 56:5, glycoursodeoxycholic acid, L-prolyl-L-threonine, triglycerides 56:7, triglycerides 54:2, glycoursodeoxycholic acid, and triglycerides 52:1.

[0101] The ROC results showed an AUC of 0.849 (sensitivity: 80.8%, specificity: 76.9%). This indicates that the constructed diagnostic model has high diagnostic efficacy.

[0102] To further verify the effectiveness of the diagnostic model for distinguishing between benign breast diseases and breast cancer constructed based on the modeling group data, validation was performed in the validation group of Example 1. The ROC results are as follows: Figure 4 As shown, the results indicate that in the validation group, AUC = 0.810 (sensitivity: 76.5%, specificity: 70.6%).

[0103] The results above indicate that the diagnostic model constructed to distinguish between benign breast diseases and breast cancer also has good diagnostic performance in the validation group.

[0104] Example 6: A diagnostic model for differentiating benign breast diseases from breast cancer was constructed using a combination of four metabolic biomarkers. Using the modeling group data from Example 1, a diagnostic model was constructed using a combination of four metabolic biomarkers. The only difference from Example 2 was that the diagnostic model was constructed using a combination of four metabolic biomarkers: arginine, triglycerides 56:5, glycoursodeoxycholic acid, and L-prolyl-L-threonine.

[0105] The ROC results showed an AUC of 0.843 (sensitivity: 76.9%, specificity: 75%). This indicates that the constructed diagnostic model has high diagnostic efficacy.

[0106] To further verify the effectiveness of the diagnostic model for distinguishing between benign breast diseases and breast cancer constructed based on the modeling group data, validation was performed in the validation group of Example 1. The ROC results are as follows: Figure 5 As shown, the results indicate that in the validation group, AUC = 0.820 (sensitivity: 76.5%, specificity: 76.5%).

[0107] The results above indicate that the diagnostic model constructed to distinguish between benign breast diseases and breast cancer also has good diagnostic performance in the validation group.

[0108] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A combination of metabolic markers for differentiating benign breast diseases from breast cancer, characterized in that, The combination of metabolic markers used to distinguish between benign breast diseases and breast cancer includes arginine, triglycerides 56:5, glycoursodeoxycholic acid, and L-prolyl-L-threonine.

2. The combination of metabolic markers for distinguishing between benign breast diseases and breast cancer according to claim 1, characterized in that, The combination of metabolic markers used to distinguish between benign breast diseases and breast cancer also includes at least one of triglyceride 56:7, triglyceride 54:2, glycine deoxycholic acid, and triglyceride 52:

1.

3. The combination of metabolic markers for distinguishing between benign breast diseases and breast cancer according to claim 2, characterized in that, The combination of metabolic markers used to distinguish between benign breast diseases and breast cancer also includes at least one of 3,6-dehydro-D-galactose, linolenic acid, phosphatidylethanolamine 36:6e, and hypoxanthine.

4. The combination of metabolic markers for distinguishing between benign breast diseases and breast cancer according to claim 3, characterized in that, The combination of metabolic markers used to distinguish between benign breast diseases and breast cancer also includes at least one of uracil, octyl-L-carnitine, cholic acid, and phosphatidylcholine 19:

2.

5. The combination of metabolic markers for distinguishing between benign breast diseases and breast cancer according to claim 4, characterized in that, The combination of metabolic markers used to distinguish between benign breast diseases and breast cancer also includes at least one of dehydroepiandrosterone sulfate, triglyceride 46:2, L-glutamyl-L-serine, triglyceride 44:0, and L-piperidinic acid.

6. The use of the combination of metabolic markers for distinguishing between benign breast diseases and breast cancer as described in any one of claims 1-5 in the preparation of a product for distinguishing between benign breast diseases and breast cancer.

7. The application according to claim 6, characterized in that, The samples used by the product to differentiate between benign breast diseases and breast cancer include at least one of serum, plasma, blood, and dried blood smears.

8. The application according to claim 7, characterized in that, The products include reagents or kits.

9. The application according to claim 8, characterized in that, The kit includes metabolic biomarker controls and / or metabolic biomarker standards.

10. The application of a combination of metabolic biomarkers for distinguishing between benign breast diseases and breast cancer as described in any one of claims 1-5 in constructing a diagnostic model for distinguishing between benign breast diseases and breast cancer.