A marker panel for breast cancer screening, diagnosis and uses thereof

By constructing a combination of peptide biomarkers and combining them with mass spectrometry and machine learning techniques, the problems of insufficient sensitivity and specificity in breast cancer screening and diagnosis have been solved, achieving more efficient breast cancer detection.

CN121186360BActive Publication Date: 2026-04-14长兴固容生物科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing breast cancer screening and diagnosis methods suffer from insufficient sensitivity and specificity. Liquid biopsy testing is costly and inefficient. In particular, mammography is prone to missed diagnoses in high-density breast tissue. Traditional biomarkers such as CEA and CA125 have low specificity, and ctDNA extraction is easily affected by interference.

Method used

A biomarker set, including peptides with specific sequences, is used to detect combinations of peptide biomarkers in serum using mass spectrometry. A classification model is constructed for auxiliary diagnosis and early screening, and the detection results are optimized by combining machine learning algorithms.

Benefits of technology

It improves the sensitivity and specificity of breast cancer auxiliary diagnosis, which is superior to existing methods, especially digital breast tomography fusion imaging and traditional biomarkers, while reducing detection costs and improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121186360B_ABST
    Figure CN121186360B_ABST
Patent Text Reader

Abstract

The application discloses a marker group for breast cancer screening and diagnosis and application thereof, and belongs to the field of molecular biology technology.The marker group comprises at least four polypeptides shown in sequences SEQ ID NO.1-36.The marker group constructed by the application has more excellent sensitivity and specificity when used for the auxiliary diagnosis and early screening of breast cancer.The auxiliary diagnosis result is better than that of the currently recognized breast screening method, digital breast tomosynthesis (DBT) (sensitivity and specificity are 97.8% and 66.7% respectively).The early screening result is better than that of the traditional breast cancer marker CEA detection (sensitivity is 63.33%, and specificity is 86%) and CA125 detection (sensitivity is 66.67%, and specificity is 90%).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular biology technology, specifically relating to a biomarker group for breast cancer screening and diagnosis and its uses. Background Technology

[0002] Early-stage breast cancer patients often have no obvious symptoms. However, as tumor cells proliferate, they can form lumps inside the breast. In the middle and late stages, breast tumor cells can metastasize to other organs or tissues throughout the body via lymphatic vessels and blood vessels, seriously threatening the patient's life. Currently, international guidelines recommend mammography as the primary screening method for breast cancer. However, mammography has limitations in processing high-density breast tissue, and relying solely on X-ray screening may result in a missed diagnosis in 30% of cases.

[0003] Currently, tumor markers such as carcinoembryonic antigen (CEA), carbohydrate antigen 125 (CA125), and cytokeratin 19 fragment 21-1 (CYFRA21-1) are increasingly used clinically to aid in the diagnosis of breast cancer. However, the value of combined detection of these markers remains inconclusive. CEA is a proteoglycan complex and is often used clinically as a tumor marker for various adenocarcinomas; however, its specificity is low when detected alone, and it often needs to be combined with other markers for diagnosis. Studies have shown that serum CA125 levels can be significantly elevated in patients with breast cancer and other cancers, suggesting that CA125 has high sensitivity but poor specificity in diagnosing tumors.

[0004] Liquid biopsy technology also demonstrates unique value in long-term monitoring after breast cancer surgery. This method allows patients to detect signs of metastasis at least four years earlier, facilitating timely identification of metastatic tumors and enabling effective intervention in their early stages, thus significantly improving patient survival rates. Studies have shown that detecting ctDNA levels in peripheral blood can diagnose breast cancer up to five months earlier than clinical imaging and radiological techniques. Compared to patients with benign breast lesions, patients with early-stage breast cancer have higher ctDNA levels, which decrease after surgery. However, due to individual differences and the extremely low concentration in peripheral blood, ctDNA extraction is easily interfered with by cfDNA. Furthermore, the rarity and heterogeneity of the detection targets in liquid biopsy lead to significant variations depending on disease stage, location, and detection method. This results in high testing costs and low efficiency for current liquid biopsy methods. Summary of the Invention

[0005] In view of the above-mentioned shortcomings in the prior art, the present invention provides a set of biomarkers for breast cancer screening and diagnosis and their uses. It has excellent sensitivity and specificity when used for auxiliary diagnosis and early screening of breast cancer, and is expected to be applied to the diagnosis and treatment of breast cancer.

[0006] To achieve the above objectives, the technical solution adopted by the present invention to solve its technical problem is as follows:

[0007] A biomarker set for breast cancer screening and diagnosis includes at least four of the polypeptides shown in SEQ ID NO. 1-36, with specific sequences shown in Table 1.

[0008] Table 1. Peptide Sequences

[0009]

[0010] Among them, the fourth amino acid G in peptide 13 is modified with Phospho; the fourth amino acid G in peptide 15 is modified with Dehydrated; the first amino acid Q in peptide 16 is modified with Gln->pyro-Glu, and the sixth amino acid N is modified with Dehydrated; the second amino acid K in peptide 17 is modified with Acetyl; and the third amino acid G in peptide 27 is modified with Phospho.

[0011] Furthermore, the biomarker set includes the peptides shown in sequences 1 and 2, as well as the following combinations of peptides:

[0012] The polypeptide combination is one of sequence 3 and sequence 4; sequence 5 and sequence 6; sequence 11 and sequence 12; sequence 30 and sequence 36; sequence 3, sequence 4 and sequence 5; sequence 5, sequence 6 and sequence 7; sequence 11, sequence 12 and sequence 20; sequence 9, sequence 30 and sequence 36.

[0013] Furthermore, the biomarker set includes the peptides shown in sequences 7 and 8, as well as the following combinations of peptides:

[0014] The polypeptide combination is one of the following: sequence 3 and sequence 4; sequence 9 and sequence 10; sequence 3, sequence 4 and sequence 9; or sequence 9, sequence 10 and sequence 11.

[0015] Furthermore, the biomarker set includes peptides shown in sequences 1, 10, and 30, as well as peptides shown in sequences 36, 20, 36, and 9, or sequences 20 and 22.

[0016] Furthermore, the biomarker set includes the peptides shown in sequences 5 and 15, as well as the following combinations of peptides:

[0017] The polypeptide combination is one of the following: sequence 25 and sequence 36; sequence 2 and sequence 16; sequence 9, sequence 25 and sequence 36; sequence 2, sequence 16 and sequence 20.

[0018] Furthermore, the biomarker set includes the polypeptides shown in sequences 4, 8, 19, 20 and / or 22.

[0019] Furthermore, the biomarker set includes sequence 7, sequence 10, and the following combination of peptides:

[0020] The polypeptide combination includes sequences 2 and 4; sequences 17 and 18; or sequences 17, 18 and 22.

[0021] Furthermore, the biomarker group includes polypeptides as shown in SEQ ID NO. 1~36.

[0022] The use of the above biomarkers in the preparation of formulations for breast cancer screening and diagnosis.

[0023] The above biomarkers may be used in basic medical research for non-diagnostic / therapeutic purposes.

[0024] Further, basic medical research includes Western blotting, immunohistochemistry, or flow cytometry.

[0025] The beneficial effects of this invention are:

[0026] The biomarker combination constructed in this invention exhibits superior sensitivity and specificity in the auxiliary diagnosis and early screening of breast cancer. Its auxiliary diagnostic results are superior to the currently accepted breast screening method, digital breast computed tomography (DBT) (sensitivity and specificity of 97.8% and 66.7%, respectively). Its early screening results are superior to traditional breast cancer biomarkers CEA detection (sensitivity of 63.33%, specificity of 86%) and CA125 detection (sensitivity of 66.67%, specificity of 90%), and it holds promise for application in the diagnosis and treatment of breast cancer. Attached Figure Description

[0027] Figure 1 ROC curve for marker combination 1;

[0028] Figure 2 ROC curve for marker combination 2;

[0029] Figure 3 ROC curve for marker combination 3;

[0030] Figure 4 ROC curve for marker combination 4;

[0031] Figure 5 ROC curve for marker combination 5;

[0032] Figure 6 ROC curve for marker combination 6;

[0033] Figure 7 ROC curve for marker combination 7;

[0034] Figure 8 ROC curve for marker combination 8;

[0035] Figure 9 ROC curve for marker combination 9;

[0036] Figure 10 ROC curve for marker combination 10;

[0037] Figure 11 ROC curve for marker combination 11;

[0038] Figure 12 ROC curve for marker combination 12;

[0039] Figure 13 ROC curve for marker combination 13;

[0040] Figure 14 ROC curve for marker combination 14;

[0041] Figure 15 ROC curve for marker combination 15;

[0042] Figure 16 ROC curve for marker combination 16;

[0043] Figure 17 ROC curve for marker combination 17;

[0044] Figure 18 ROC curve for marker combination 18;

[0045] Figure 19 ROC curve for marker combination 19;

[0046] Figure 20 ROC curve for marker combination 20;

[0047] Figure 21 ROC curve for marker combination 21;

[0048] Figure 22 ROC curve for marker combination 22;

[0049] Figure 23 ROC curve for marker combination 23;

[0050] Figure 24 ROC curve for marker combination 24;

[0051] Figure 25 ROC curve for marker combination 25;

[0052] Figure 26 ROC curve for marker combination 26. Detailed Implementation

[0053] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0054] The patient samples used in this invention are all from Zhongshan Hospital affiliated with Fudan University and have passed ethical review.

[0055] The experimental methods used in this invention are as follows:

[0056] I. Serum Sample Collection

[0057] 1) Sample type: serum.

[0058] 2) Collection requirements: Fasting is required. Use a coagulation tube to draw 5 mL of venous blood, let it stand for 30 min, centrifuge at 3000 rpm for 15 min, and take out about 1 mL of serum and put it into a cryopreservation tube.

[0059] 3) Sample storage:

[0060] Use on the same day; store at 2-8℃.

[0061] If not used on the same day, store at -20℃ for up to 30 days;

[0062] If stored for an extended period (more than one month), it should be kept at -80°C.

[0063] The freeze-thaw cycle should not exceed 3 times.

[0064] II. Extraction of analytes from serum

[0065] 1) After calibrating the mass spectrometer, turn on the Solid Bio Fully Automated Sample Analysis System SPS1000 / SPS4000, and put in the consumables, matching reagent kits and the sample to be tested;

[0066] Select the procedure method "Concentrated loading";

[0067] Run the program:

[0068] a. Opening a hole;

[0069] b. Take at least 10µL of serum sample and activation reagent, mix them in a 1:1 ratio, and place them in the G-row pre-reserved well for later use;

[0070] c. Clean the custom pipette tip in cleaning reagent 1 and cleaning reagent 2 in sequence. Each time, aspirate at least 10µL of liquid and repeat the aspiration and dispensing process at least 3 times.

[0071] d. Process the serum mixture in the G-row wells using the cleaned custom pipette tips. Aspirate at least 10 µL of solution each time, repeating the process at least three times.

[0072] e. Clean the custom pipette tip after adsorbing the serum mixture using cleaning reagent 3. During cleaning, aspirate at least 10µL of liquid each time, repeating the aspiration and dispensing process at least 3 times.

[0073] f. Transfer no less than 10µL of buffer reagent into the H-row pre-reserved hole, and place the customized pipette tip after using cleaning reagent 3 into the liquid to draw no less than 10µL of liquid. Repeat the suction and aspiration at least 3 times.

[0074] g. Transfer at least 10µL of sample matrix solution into the H-row pre-reserved well to complete sample processing;

[0075] h. Spot 2.0 µL of the solution from well H onto the hydrophobic-coated biochip (Wuxi Pimo Technology Co., Ltd.).

[0076] i. Vacuum drying for 240 seconds.

[0077] The main components of each reagent are shown in Table 2.

[0078] Table 2 Reagent Composition

[0079]

[0080] III. Mass Spectrometry Data Acquisition and Upload

[0081] The hydrophobic coated biochip (Wuxi Pimo Technology Co., Ltd.) was vacuum dried and placed into a mass spectrometer;

[0082] Data acquisition is performed using the pre-defined SP1 voltage (target high voltage), SP2 voltage (pulse high voltage), focusing voltage (lens high voltage), detector voltage (MCP voltage), pulse delay time, acquisition card range, target diameter, laser frequency, calibration method, and laser intensity.

[0083] IV. Quality Control

[0084] 1) After data collection, the data will be uploaded to the "Mass Spectrometry Data Analysis Software";

[0085] 2) The software reads the sample information and signal spectrum, and judges whether the sample and sample pretreatment are qualified according to the quality control model; quality control failure may include a variety of possibilities, including the sample is not a breast lesion sample, the signal spectrum intensity is not up to standard, etc.

[0086] 3) If the quality control fails, adjust the corresponding parameters according to the quality control results and repeat the serum analyte extraction process;

[0087] 4) If the quality control is qualified, proceed to the next process.

[0088] V. Establishment of Positive Criterion Value and Result Analysis

[0089] The study of positive cutoff values ​​used breast cancer samples with clear diagnostic information and normal human samples, with benign breast cancer samples covering non-breast cancer samples such as breast hyperplasia and fibroadenoma. The core algorithm is based on supervised learning of known breast cancer sample atlases. Through a series of processes such as smoothing, noise reduction, and baseline removal, feature peaks are screened, a classification model is constructed, and the similarity between the hormone signal atlas and known hormone signal atlases stored in the software is calculated (Cannataro M, Guzzi PH, Mazza T, et al. Preprocessing, Management, and Analysis of Mass Spectrometry Proteomics Data[J]. 2005.). Finally, the similarity score positive cutoff value of the kit is determined by the Youden index maximization method. When the similarity score < positive cutoff value, the sample test result is negative; when the similarity score ≥ positive cutoff value, the sample test result is positive. When using the maximum similarity score as the positive cutoff value to assist in the diagnosis of breast cancer or to conduct early breast cancer screening, the sensitivity is calculated as: Sensitivity = (Number of true positives / (Number of true positives + Number of false negatives)) × 100%, and Specificity is calculated as: (Number of true negatives / (Number of true negatives + Number of false positives)) × 100%.

[0090] Example 1: Screening and Identification of Biomarkers

[0091] This invention analyzed 400 normal human samples (all female, aged 25 to 65 years, mean age 44.3 ± 11.8 years; specific age distribution: 68 cases aged 25-34, 107 cases aged 35-44, 118 cases aged 45-54, and 107 cases aged 55-65) and 400 breast cancer samples (all female, aged 28 to 75 years, mean age 52.6 ± 10.4 years; specific age distribution: 42 cases aged 28-37, 42 cases aged 38-47, and 42 cases aged 48-47). 98 patients (143 aged 48-57, 86 aged 58-67, and 31 aged 68-75) underwent time-of-flight mass spectrometry (TOF-MS). Through first-level mass spectrometry testing, the relative abundance differences of characteristic peak data in normal individuals and breast cancer patients were comprehensively considered, along with statistical differences (p<0.05, t-test). A machine learning algorithm (random forest) was used to rank the impact factors for feature selection and to assess the matching degree of data in the database. This identified 36 blood peptides with diagnostic capabilities for breast cancer. The mass-to-charge ratio (m / z), relative abundance, and impact factors of these peptides are shown in Table 1. The relative abundance was normalized to ln(mean signal intensity of cancer patient samples / mean signal intensity of normal samples) based on the normal human sample. The feature peak selection in machine learning was based on feature importance assessment using ensemble learning. Multiple decision trees were constructed to quantify the contribution of each mass-to-charge ratio (m / z) peak in classification / prediction. The feature importance score, or influence factor, is obtained by calculating the mean reduction in impurity caused by the feature when splitting across all tree nodes (Biau, G., Scornet, E. A random forestguided tour. TEST 25, 197–227 (2016). https: / / doi.org / 10.1007 / s11749-016-0481-7).

[0092] The specific parameters for first-order mass spectrometry are as follows:

[0093] Ionization method: Matrix-assisted laser desorption / ionization (MALDI), with α-cyano-4-hydroxycinnamic acid (CHCA) as the matrix.

[0094] Quality range: 100-4000 Da.

[0095] Resolution: 20000 (full quality range).

[0096] Laser energy: 30-40%.

[0097] Acquisition mode: Positive ion mode.

[0098] Calibration: External quality calibration was performed using the Bruker Peptide Calibration Standard.

[0099] Subsequently, the sequences of these 36 substances in the clinical serum were confirmed using secondary mass spectrometry (MS / MS or TOF / TOF is a peptide identification method recommended by the guidelines of the China Food and Drug Administration and the U.S. Food and Drug Administration). The secondary mass spectrometry data analysis process is as follows:

[0100] Data analysis methods: Mascot software (version 2.8) was used for database searching, with the UniProt Human Proteome Database (released in 2023) as the target database. Search parameters: Enzyme was set to "no digestion", parent ion mass error was allowed ±0.5 Da, fragment ion mass error was allowed ±0.3 Da, fixed modification was cysteine ​​urea methylation, and variable modification was methionine oxidation.

[0101] Sequence confirmation criteria: The confirmation of peptide sequences is based on the matching of fragment ion spectra (b- and y- ions) with theoretical spectra. A Mascot score higher than 30 (p<0.05) is considered significant.

[0102] False positive exclusion: Specificity was verified through reverse database search, and the false positive rate was controlled to below 1%.

[0103] Secondary mass spectrometry parameters:

[0104] Collision-induced dissociation (CID).

[0105] Collision energy: 30 eV.

[0106] Fragment ion mass range: 100-3500 Da.

[0107] Data acquisition: Each sample is scanned at least 1000 times with lasers to improve the signal-to-noise ratio.

[0108] The sequences and specificity of these 36 biomarkers were confirmed by secondary mass spectrometry, ruling out false positives. The specific sequences are shown in Table 1.

[0109] Example 2: Validation of Marker Combinations

[0110] Based on the 36 polypeptide biomarkers identified and confirmed in Example 1 (sequences and mass-to-charge ratios are shown in Table 1), different biomarker combinations were formed as shown in Table 3, and the sensitivity and specificity of the biomarker combinations in the auxiliary diagnosis and early screening of breast cancer were verified. The analysis process is as follows: For all validation cohort samples, the same MALDI-TOF MS platform and parameters as in Example 1 were used for detection to obtain the mass spectrometry peak intensity data of all biomarkers in Table 1. The entire detection process was blinded, meaning that the experimental operators were unaware of the sample grouping information. The raw mass spectrometry data were processed by baseline correction, smoothing, and normalization (based on the internal standard peak intensity), and then the peak area or intensity value of each biomarker was extracted. Statistical analysis was performed using R software (version 4.0.2). The preprocessed biomarker intensity data were input into a logistic regression model. For auxiliary diagnostic validation, ten-fold cross-validation was performed using all samples from this cohort (Sun T, Liu J, Yuan H, Li X, Yan H. Construction of a risk prediction model for lung infection after chemotherapy in lung cancer patients based on the machine learning algorithm. Front Oncol. 2024 Aug 9;14:1403392. doi: 10.3389 / fonc.2024.1403392. PMID:39184040; PMCID: PMC11341396.). For early screening validation, given the imbalanced sample size, oversampling (SMOTE) was used to process the data before model training and testing (van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B. The harm of class imbalance corrections for riskprediction models: illustration and simulation using logistic regression. JAm Med Inform Assoc. 2022 Aug 16;29(9):1525-1534. doi: 10.1093 / jamia / ocac093.PMID: 35686364; PMCID: PMC9382395.).The data analysis process referenced general guidelines for constructing clinical predictive models (Zweig MH, Campbell G. Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine. Clin Chem. 1993 Apr;39(4):561-77. Erratum in: Clin Chem 1993 Aug;39(8):1589. PMID: 8472349.). Sensitivity, specificity, and area under the receiver operating characteristic (AUC) curve were calculated for each biomarker combination. The specific calculation results for the performance indicators are shown in Table 3, and the corresponding ROC curves are shown in [Table data missing]. Figures 1-26 .

[0111] 1) The use of various biomarker combinations in the auxiliary diagnosis of breast cancer was validated using a sample of 500 breast cancer patients (all female, aged 30 to 72 years, with a mean age of 51.8 ± 10.2 years; specifically, 48 patients aged 30-39, 127 patients aged 40-49, 158 patients aged 50-59, 132 patients aged 60-69, and 35 patients aged 70-72) and 500 healthy individuals (all female, aged 25 to 68 years, with a mean age of 43.7 ± 11.6 years; specifically, 73 patients aged 25-34, 128 patients aged 35-44, 146 patients aged 45-54, 115 patients aged 55-64, and 38 patients aged 65-68).

[0112] 2) The efficacy of various biomarker combinations for early breast cancer screening was validated using a sample of 350 breast cancer patients (all female, aged 32 to 70 years, mean age 50.3 ± 10.8 years; specifically, 43 patients aged 32-41, 108 patients aged 42-51, 126 patients aged 52-61, and 73 patients aged 62-70) and 3500 healthy individuals (all female, aged 25 to 75 years, mean age 45.2 ± 12.4 years; specifically, 423 patients aged 25-34, 687 patients aged 35-44, 812 patients aged 45-54, 798 patients aged 55-64, and 780 patients aged 65-75).

[0113] Table 3 Sensitivity and specificity of biomarker combinations for assisted diagnosis and early screening

[0114]

[0115] (Continued from the table above)

[0116]

[0117] (Continued from the table above)

[0118]

[0119] According to Table 3 and Figures 1-26 The test results show that the auxiliary diagnostic results of the biomarker combination selected in this invention are superior to the currently recognized breast screening method, digital breast computed tomography (DBT) (sensitivity and specificity are 97.8% and 66.7%, respectively). The early screening results are superior to traditional breast cancer biomarkers CEA detection (sensitivity 63.33%, specificity 86%) and CA125 detection (sensitivity 66.67%, specificity 90%).

[0120] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A biomarker set for breast cancer screening and diagnosis, characterized in that, The biomarker group is selected from one of the following combinations of peptides: A) Sequence 1, Sequence 2, Sequence 3 and Sequence 4; B) Sequences 1, 2, 5, and 6; C) Sequence 1, Sequence 2, Sequence 11 and Sequence 12; D) Sequences 1, 2, 30, and 36; E) Sequence 1, Sequence 2, Sequence 3, Sequence 4 and Sequence 5; F) Sequence 1, Sequence 2, Sequence 5, Sequence 6 and Sequence 7; G) Sequence 1, Sequence 2, Sequence 11, Sequence 12 and Sequence 20; H) Sequence 1, Sequence 2, Sequence 9, Sequence 30 and Sequence 36; I) Sequences 2, 5, 15, and 16; J) Sequence 2, Sequence 5, Sequence 15, Sequence 16 and Sequence 20; K) Sequence 2, Sequence 7, Sequence 10 and Sequence 4; L) Sequences 1 to 36; The amino acid sequences of sequences 1 to 36 are shown in SEQ ID NO. 1 to 36; Specifically, amino acid G at position 4 in sequence 13 has a Phospho modification; amino acid G at position 4 in sequence 15 has a Dehydrated modification; amino acid Q at position 1 in sequence 16 has a Gln->pyro-Glu modification, and amino acid N at position 6 has a Dehydrated modification; amino acid K at position 2 in sequence 17 has an Acetyl modification; and amino acid G at position 3 in sequence 27 has a Phospho modification.

2. Use of the reagent for detecting the biomarker group of claim 1 in the preparation of preparations for breast cancer screening and diagnosis.

Citation Information

Patent Citations

  • Specific polypeptide and use of such specific polypeptide in breast cancer early diagnosis

    CN102558328A

  • Serum biomarkers

    CN112639474A