Markers for serum protein-based breast tumor screening and uses thereof

By screening 23 serum proteins and neutrophil ratios using LC-MS mass spectrometry, and combining this with machine learning algorithms to construct a breast tumor classifier, the problem of difficulty in early diagnosis of breast cancer in existing technologies has been solved, achieving high accuracy and high specificity in the diagnosis of breast lesions.

CN117517660BActive Publication Date: 2026-07-24GUANGDONG GENERAL HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG GENERAL HOSPITAL
Filing Date
2023-11-28
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Current technologies lack serum protein biomarkers that can be used to distinguish between healthy individuals, benign breast nodules, and breast cancer. Furthermore, previous studies have failed to effectively combine the intersecting protein profiles of tissues and serum, making early diagnosis of breast cancer difficult.

Method used

Twenty-three serum proteins and neutrophil ratios were screened using LC-MS mass spectrometry. A breast tumor classifier was constructed by combining random forest algorithm and extreme value gradient enhancement algorithm to distinguish between healthy individuals, benign breast lesions, and malignant breast tumors.

Benefits of technology

It achieved high accuracy and specificity in the diagnosis of breast lesions, with an accuracy of 0.87, precision of 0.93, and area under the ROC curve of 0.96 in the independent validation cohort, significantly improving the efficacy of early diagnosis of breast cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004580181570000011
    Figure HDA0004580181570000011
  • Figure HDA0004580181570000012
    Figure HDA0004580181570000012
  • Figure HDA0004580181570000021
    Figure HDA0004580181570000021
Patent Text Reader

Abstract

The present application relates to the field of biological medical technology, in particular to a serum protein-based breast tumor screening marker and application. The present application constructs a serum protein-based breast tumor marker, which can be used for serum detection of patients suspected of having breast tumors, effectively distinguishing healthy people, breast nodule patients and breast cancer patients. The classifier has an Accuracy of 0.87, a Precision of 0.93, a Recall of 0.82 and an F1-score of 0.86 in an independent verification queue. The area under the ROC curve for benign nodules is 0.96, the area under the ROC curve for breast cancer is 0.96, and the area under the ROC curve for healthy people is 0.96. The present application has high accuracy and good specificity, and has a good application prospect in the clinical diagnosis of breast lesions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to biomarkers and applications for breast tumor screening based on serum proteins. Background Technology

[0002] Breast cancer is one of the most common malignant tumors among women worldwide. Due to the lack of non-invasive, highly sensitive, specific and stable early diagnostic biomarkers, the situation of early diagnosis and prevention of breast cancer in my country is severe.

[0003] Blood, due to its easy availability, high homogeneity, and rich content of various substances, has become the preferred subject for research on novel tumor markers. Blood proteins, in particular, are favored in early diagnosis due to their stability and compatibility with various methodologies. M Toi et al. screened the serum of 90 healthy individuals and 54 breast cancer patients using mass spectrometry, finding that RALGAPA2, PKG1, TJP2, and NFX1 proteins may serve as diagnostic markers for breast cancer, with a receiver operating characteristic (AUC) of 0.874. Furthermore, researchers discovered that CDH5 may serve as a serum marker for metastatic breast cancer, and its expression level can significantly distinguish between recurrent and non-recurrent breast cancer, with a specificity of 90%. Meanwhile, Ladd's team found that serum breast cancer proteins DUSP9, EED, EFNA5, ITGB1, and PPMT1 can serve as indicators of early triple-negative breast cancer.

[0004] However, current research on breast cancer protein biomarkers lacks cohort studies of benign breast lesions and the entire course of breast cancer, resulting in a lack of serum protein biomarker profiles that can differentiate between healthy individuals, benign breast nodules, and breast cancer. Furthermore, previous studies have screened protein biomarkers based either on tumor tissue or on the serum of cancer patients, without investigating the protein profiles where these two approaches overlap. Therefore, it is crucial to develop a serum protein biomarker profile based on patient serum and protein biomarkers that can differentiate between healthy individuals, benign breast nodules, and breast cancer. Summary of the Invention

[0005] In view of this, the technical problem to be solved by the present invention is to provide biomarkers for breast tumor screening based on serum proteins and their applications.

[0006] This invention provides biomarkers for breast disease screening, including: OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3, and neutrophil ratio GR.

[0007] Furthermore,

[0008] The breast diseases mentioned include: benign breast tumors and / or malignant breast tumors; the benign breast tumors include fibroadenomas or papillomas; the malignant breast tumors include breast cancer. The biomarkers described in this invention include 23 serum proteins and neutrophil ratio. The detection results of this invention show that including the neutrophil ratio in the classification detection spectrum significantly improves diagnostic efficacy compared to not including the neutrophil ratio.

[0009] This invention provides a method for preparing the aforementioned marker, which includes the following steps:

[0010] Step 1: After preprocessing and mass spectrometry detection of serum samples from healthy individuals, benign breast tumors, and malignant breast tumors, protein proteomic data of the three types of serum samples were obtained through protein identification and quantification.

[0011] Step 2: The proteomic data of the three serum samples were subjected to quality control, preprocessing, correction and screening to obtain differential proteomic data 1;

[0012] Step 3: The differential proteome data 1 is combined with blood routine indicators and classified, and the biomarkers are obtained by screening based on the training set and the test set.

[0013] The pretreatment includes the following steps: after removing the five most abundant serum proteins from the serum using the High-Select Top 14 Protein Removal Kit, the serum is treated for 17 hours at a trypsin:protein mass ratio of 1:25.

[0014] The pretreatment process also includes a drying step; the dried sample is then subjected to mass spectrometry detection using an aqueous solution containing 0.1% formic acid by volume.

[0015] The false discovery rate (FDR) for the protein identification was 1%.

[0016] The quantitative parameters were set as follows: Precursor FDR of 1%; Log lev of 1; quality precision of 20 ppm; MS1 precision of 10 ppm; scanning window of 30°; implicit proteome as genes; and robust LC (high precision) as the quantitative strategy.

[0017] In step 2, the proteomic data from the three serum samples undergo a quality control step before preprocessing. This quality control involves pooling all three serum samples into a serum pool as a QC standard. The QC standard is analyzed using the same methods and conditions as the serum cohort. The Pearson correlation coefficient of the QC standard is calculated. The average correlation coefficient of the QC standard is 0.85. The minimum correlation coefficient is 0.82, and the maximum correlation coefficient is 0.89, demonstrating the stability of the mass spectrometry analysis platform.

[0018] The threshold for the preprocessing is:

[0019] In the three serum samples, the overlap of protein types in each serum sample with the other two serum samples is ≥50%.

[0020] In this invention, the threshold setting of the preprocessing will affect the reliability or reproducibility of protein identification, and further affect the diagnostic performance. In this invention, the threshold of the preprocessing has been optimized. In some embodiments of this invention, the threshold is set to 30%, which significantly reduces the diagnostic performance.

[0021] The correction involves performing KNN imputation on each sample type of data using impute to imputate missing values, with the imputation parameter set to 10. -5 .

[0022] The screening criteria are as follows: among the three serum samples, the expression value of the specific serum protein in any diseased serum sample is ≥ 1.5 times the expression value of the other two groups; and the specific serum protein is among the top 10% of the differentially expressed proteins in multiple groups with the greatest difference.

[0023] The arbitrary lesion serum sample includes serum samples of benign breast tumors and serum samples of malignant breast tumors.

[0024] This invention optimizes the screening criteria. Combined with a pretreatment threshold, experimental results show that only when the pretreatment threshold is set to 50%, and the expression value of the specific serum protein in any diseased serum sample is ≥1.5 times the expression value of the other two groups, and this is combined with the neutrophil ratio in routine blood tests, does the diagnostic effect for breast lesions significantly outperform other conditions. Furthermore, the number of differentially expressed proteins also affects the diagnostic efficacy of the screened biomarkers. Comparing the top 10% of differentially expressed proteins in each group with the two groups showing the greatest difference in multiple differentially expressed proteins, selecting the top 10% of the groups with the greatest difference in multiple differentially expressed proteins is more suitable for screening biomarkers with high specificity and accuracy according to this invention.

[0025] The classification methods include random forest algorithm, support vector machine algorithm, adaptive boosting algorithm, extreme gradient boosting algorithm, etc. In a specific embodiment of the present invention, random forest algorithm and extreme gradient boosting algorithm are used for classification. The results show that extreme gradient boosting algorithm is more suitable for screening the markers of the present invention.

[0026] Furthermore, in the preparation method described in this invention,

[0027] The parameters for mass spectrometry detection include:

[0028] Column: C18, particle size 1.9 μm; inner diameter 150 μm, length 15 cm; pore size

[0029] The chromatographic conditions are as follows:

[0030] Mobile phase A: 0.1% formic acid and water (volume fraction);

[0031] Mobile phase B: 0.1% formic acid and acetonitrile (volume fraction);

[0032] The flow rate is 600 nL / min;

[0033] The gradient range and duration are: 15%–30% mobile phase B, separation for 75 min.

[0034] Mass spectrometry conditions: MS1 scan from 300–1, 400 m / z at 60 kΩ resolution (AGC target 4e5 or 50 ms). Then, 30 DIA fragments were acquired at 15 kΩ resolution with an AGC target of 5e4 or a maximum injection time of 22 ms. The "Implant ions for all available parallel times" setting was enabled. High-energy collisional dissociation (HCD) fragmentation was set to 30% of the normalized collision energy. Spectra were recorded in profile mode. The default charge status for MS2 was set to 3.

[0035] The biomarkers described in this invention are accurate diagnostic markers for breast lesions, obtained under optimized conditions with multiple parameters. Compared to other methods that can only diagnose tumors and non-tumor conditions, the biomarkers of this invention can diagnose healthy individuals, benign breast lesions, and malignant breast tumors with high accuracy and specificity. The various parameters in the preparation method of the biomarkers described in this invention work together. The preprocessing threshold is: in three serum samples, the overlap of protein types in each serum sample with the other two serum samples is ≥50%; the screening criteria are: in any of the three serum samples, the expression value of the specific serum protein in any lesion serum sample is ≥1.5 times the expression value of the other two groups; and the specific serum protein is among the top 10% of differentially expressed proteins; only by employing an extreme value gradient enhancement algorithm can the highly accurate and specific diagnostic biomarkers of this invention be obtained.

[0036] This invention provides the application of the described biomarker or the biomarker prepared by the described method in the preparation of breast lesion diagnostic products or the establishment of breast disease screening models.

[0037] This invention provides a breast tumor diagnostic product, which includes or is prepared using the biomarkers described in this invention or the preparation method described in this invention.

[0038] This invention constructs a serum protein-based breast tumor marker, which can be used for serum testing in patients suspected of having breast tumors. It effectively distinguishes between healthy individuals, patients with breast nodules, and breast cancer patients. In an independent validation cohort, this classifier achieved an accuracy of 0.87, a precision of 0.93, a recall of 0.82, and an F1-score of 0.86. The area under the ROC curve (AUC) for benign nodules, breast cancer, and healthy individuals was 0.96. It exhibits high accuracy and specificity, showing promising application prospects in the clinical diagnosis of breast lesions. Attached Figure Description

[0039] Figure 1 It is an analysis flowchart;

[0040] Figure 2 This is a schematic diagram of proteins specifically expressed by healthy individuals, fibroadenomas, papillomas, and breast cancer.

[0041] Figure 3 This is the number of proteomes in all patients when the false discovery rate (FDR) at the protein and peptide levels is 1%.

[0042] Figure 4 These are proteins specifically expressed in healthy controls, fibroadenomas, papillomas, and breast cancers. A represents the variation of breast-related proteins among the groups; B represents the variation of serum proteins among the groups; C represents the variation of related proteins with drug targets among the groups; D represents the variation of proteins with FDA-approved drugs among the groups; E represents the variation of membrane surface protein markers among the groups; and F represents the variation of tumor-related proteins among the groups.

[0043] Figure 5Principal component analysis and expression differential volcano plots are used for healthy controls (HC), fibroadenomas, papillomas, and breast cancers. A represents principal component analysis, and B represents expression differential volcano plots.

[0044] Figure 6 This is a list of tumor-specific serum proteins identified according to the standards for tumor-specific serum proteins. A shows the expression heatmap of specific proteins across different groups; B shows the cluster analysis of specific proteins in each group.

[0045] Figure 7 This represents the classification results of the protein spectrum classifier in the training queue;

[0046] Figure 8 It is a heatmap of 24 protein discovery cohorts;

[0047] Figure 9 This represents the classification results of the protein spectrum classifier in the independent validation queue;

[0048] Figure 10 It is a heatmap of 24 protein validation cohorts;

[0049] Figure 11 It is a high-abundance protein removal kit for screening;

[0050] Figure 12 This refers to the optimization of pretreatment conditions, where A is the ratio of enzyme to protein and B is the pancreatic enzyme treatment time.

[0051] Figure 13 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 30%. When the sample fold was greater than 1.5 times, the top 2 specifically expressed proteins in each group were selected as the components of the protein profile. Random forest method was used for classification. Figure A is the ROC curve of the classification results, and Figure B is the heatmap of the classification results.

[0052] Figure 14 The classification results were obtained by using 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of this cohort was used as the validation cohort. Figure A shows the ROC curve of the classification results, and Figure B shows the heatmap of the classification results.

[0053] Figure 15The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold was greater than 1.5 times, the top 2 specifically expressed proteins in each group were selected as components of the protein profile. Random forest method was used for classification. Figure A is the ROC curve of the classification results, and Figure B is the heatmap of the classification results.

[0054] Figure 16 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50%, and the sample Fold was greater than 1.5 times, the top 2 specifically expressed proteins in each group were selected as components of the proteomic profile. Random forest was used for classification. Here, A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0055] Figure 17 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.0, the top 2 specifically expressed proteins in each group were selected as components of the protein profile. Random forest method was used for classification. Here, A is the ROC curve of the classification results and B is the heatmap of the classification results.

[0056] Figure 18 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold>1.0 was selected, the top 2 specifically expressed proteins in each group were selected as components of the proteomic profile. Random forest method was used for classification. Here, A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0057] Figure 19 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 2.0, the top 2 specifically expressed proteins in each group were selected as components of the protein profile. Random forest method was used for classification. Here, A is the ROC curve of the classification results and B is the heatmap of the classification results.

[0058] Figure 20Using 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens, 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 2.0, the top 2 specifically expressed proteins in each group were selected as components of the proteomic profile. Random forest method was used for classification. Here, A is the ROC curve of the classification results and B is the heatmap of the classification results.

[0059] Figure 21 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 2 specifically expressed proteins in each group were selected as components of the protein profile. Random forest method was used for classification. Here, A is the ROC curve of the classification results and B is the heatmap of the classification results.

[0060] Figure 22 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 2 specifically expressed proteins in each group were selected as components of the proteomic profile. Random forest was used for classification. Here, A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0061] Figure 23 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 2 specifically expressed proteins of each group were selected as components of the protein profile. Neutrophil ratio was added as a classification indicator. Random forest method was used for classification. A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0062] Figure 24 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 2 specifically expressed proteins in each group were selected as components of the protein profile. Neutrophil ratio was added as a classification indicator. Random forest method was used for classification. Here, A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0063] Figure 25The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 10% of the specifically expressed proteins in each group were selected as components of the protein profile. Neutrophil ratio was added as a classification indicator. Random forest method was used for classification. Here, A is the ROC curve of the classification results and B is the heatmap of the classification results.

[0064] Figure 26 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 10% of the specifically expressed proteins in each group were selected as components of the protein profile. Neutrophil ratio was added as a classification indicator. Random forest method was used for classification. Here, A is the ROC curve of the classification results, and B is the heatmap of the classification results.

[0065] Figure 27 A new independent sample cohort was used, consisting of 57 breast cancer samples, 27 healthy individuals, and 29 benign breast nodule samples. The classifier obtained under these conditions (with data preprocessing parameters set to 50%, sample Fold > 1.5, the top 10% of specifically expressed proteins in each group selected as protein profile components, and neutrophil ratio added as a classification indicator, and classification performed using the random forest method) was validated, and the classification results were obtained.

[0066] Figure 28 A new independent sample cohort was used, consisting of 57 breast cancer samples, 27 healthy individuals, and 29 benign breast nodule samples. The classifier obtained under these conditions (with data preprocessing parameters set to 50%, sample Fold > 1.5, the top 10% of specifically expressed proteins in each group selected as components of the proteome, and neutrophil ratio added as a classification indicator, and classification performed using the random forest method) was validated, and the resulting proteomic heatmap was obtained.

[0067] Figure 29 The importance of each indicator in the classified protein spectrum is displayed by the classifier under the following conditions (when the data preprocessing parameters are set to 50%, the screening sample Fold>1.5, the top 10% of the specifically expressed proteins in each group are selected as components of the protein spectrum, and the neutrophil ratio is added as a classification indicator, and the random forest method is used for classification).

[0068] Figure 30The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 70% of the samples in this cohort were used as the training cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 10% of the specifically expressed proteins in each group were selected as components of the protein profile. Neutrophil ratio was also added as a classification indicator, and the extreme value gradient enhancement algorithm was used for classification.

[0069] Figure 31 The study used 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. 30% of the samples in this cohort were used as a validation cohort. When the data preprocessing parameters were set to 50% and the sample Fold > 1.5, the top 10% of the specifically expressed proteins in each group were selected as components of the protein profile. Neutrophil ratio was also added as a classification indicator, and the extreme value gradient enhancement algorithm was used for classification.

[0070] Figure 32 A new independent sample cohort was used, consisting of 57 breast cancer samples, 27 healthy individuals, and 29 benign breast nodule samples. The classifier obtained under these conditions (with data preprocessing parameters set to 50%, sample Fold > 1.5, the top 10% of specifically expressed proteins in each group selected as protein profile components, and neutrophil ratio added as a classification indicator, and classification performed using an extreme value gradient enhancement algorithm) was validated, and the classification results were obtained.

[0071] Figure 33 The proteomic heatmap was obtained by using 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy specimens. Under these conditions (data preprocessing parameters were set to 50%, specimen Fold > 1.5, the top 10% of specifically expressed proteins in each group were selected as components of the proteomic profile, and neutrophil ratio was added as a classification indicator, and the extreme value gradient enhancement algorithm was used for classification), the classifier was validated.

[0072] Figure 34 A new independent sample cohort was used, consisting of 57 breast cancer samples, 27 healthy individuals, and 29 benign breast nodule samples. The classifier obtained under these conditions (with data preprocessing parameters set to 50%, sample Fold > 1.5, the top 10% of specifically expressed proteins in each group selected as components of the proteome, and neutrophil ratio added as a classification indicator, and classification performed using an extreme value gradient enhancement algorithm) was validated, and the resulting proteomic heatmap was obtained.

[0073] Figure 35The importance of each indicator in the classified protein spectrum is displayed by the classifier under the following conditions (when the data preprocessing parameters are set to 50%, the sample Fold>1.5, the top 10% of the specifically expressed proteins in each group are selected as components of the protein spectrum, and the neutrophil ratio is added as a classification indicator, and the extreme value gradient enhancement algorithm is used for classification). Detailed Implementation

[0074] This invention provides a breast tumor classifier based on serum proteins. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to implement it. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can obviously make modifications or appropriate alterations and combinations to the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.

[0075] Currently, due to the lack of studies on benign breast lesions and the entire course of breast cancer cohorts in previous research on breast cancer protein biomarkers, there is a lack of serum protein biomarker profiles suitable for differentiating between healthy individuals, benign breast nodules, and breast cancer. Furthermore, previous studies have either screened protein biomarkers based on tumor tissue or serum from cancer patients, without investigating the overlapping protein profiles between the two. Therefore, this study analyzed the serum of 56 healthy individuals, 112 individuals with benign breast lesions, and 154 individuals with breast cancer (carcinoma in situ, microinvasive, invasive, lymph node metastasis, and distant metastasis), and identified 24... The serum characteristics include 23 proteins and 1 blood routine test indicator (OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3, GR), which can effectively distinguish between healthy people, patients with breast nodules, and patients with breast cancer.

[0076] We used LC-MS mass spectrometry to detect serum samples from healthy individuals, patients with benign breast nodules, and breast cancer patients, screening for proteins that are specifically expressed in both tissues and serum at different stages of disease development. We also established a serum protein classifier (24 characteristic profiles) to differentiate between benign and malignant breast tumors.

[0077] The reagents and consumables used in this invention are all commercially available products that can be purchased on the market.

[0078] The present invention will be further illustrated below with reference to the embodiments:

[0079] Example 1: Construction of a Breast Tumor Classifier Based on Serum Proteins

[0080] I. Construction of a Breast Tumor Classifier Based on Serum Proteins

[0081] 1. Serum sample pretreatment

[0082] Serum samples from 56 healthy individuals, 112 individuals with benign breast lesions, and 154 individuals with breast cancer (carcinoma in situ, microinvasive, invasive, lymph node metastasis, and distant metastasis) were pre-processed.

[0083] First, the five most abundant serum proteins (LGKC, IGHG1, ALB, IGHG2, and APOA1) were removed using a commercial kit (Thermo Fisher, A36369) according to the manufacturer's instructions, followed by inactivation at 85°C for 10 minutes. Then, the serum sample, after removing the high-abundance serum proteins, was treated with trypsin at 37°C for 17 hours (trypsin:protein (mass) = 1:25), followed by peptide extraction and drying.

[0084] 2. Mass spectrometry detection

[0085] Samples were measured using an LC-MS instrument comprised of an EASY-nLC 1200 ultra-high pressure system (Thermo Fisher Scientific), coupled with a nano-electrospray ionization source (Thermo Fisher Scientific) and a Fusion Lumos Orbitrap for plasma samples (Thermo Fisher Scientific). Serum samples were dissolved in 12 μL of loading buffer (0.1 (v / v)% formic acid aqueous solution), and 5 μL of the solution was loaded onto a 100 μL id×2.5 cm, C18 trapping column at a maximum pressure of 280 bar using 14 μL of solvent a (0.1 (v / v)% formic acid aqueous solution). Peptides were then analyzed on a 150 μm id×15 cm column (C18, 1.9 μm). Separation was performed on a mobile phase (Dr. Maisch GmbH) with linear mobile phase B (CAN acetonitrile and 0.1 (v / v)% formic acid) at 15%–30%, at a flow rate of 600 nL / min for 75 min. Mass spectrometry analysis was performed using the data-independent method (DIA). The DIA method consisted of an MS1 ​​scan from 300–1400 m / z at 60 kV resolution (AGC target 4e5 or 50 ms). Thirty DIA fragments were then obtained at 15 kV resolution with an AGC target of 5e4 or a maximum injection time of 22 ms. The "Implant ions for all available parallel times" setting was enabled. High-energy collisional dissociation (HCD) fragmentation was set to 30% of the normalized collision energy. Spectra were recorded in profile mode. The default charge status for MS2 was set to 3.

[0086] 3. Protein identification and quantification methods

[0087] All data were processed using Firmiana. DIA data were retrieved from the UniProt human protein database (updated 2019.12.17, 20406 entries) using FragPipe (v12.1) and MSFragger (2.2)22. The precursor mass tolerance was 20 ppm, and the product ion mass tolerance was 50 mmu. A maximum of two isoforms were allowed to be omitted. The search engine used cysteine ​​aminomethylation as a fixed modification and N-acetylation and methionine oxidation as variable modifications. Precursor ion fractional charge limits were +2, +3, and +4. Data were also retrieved from a decoy database so that protein identification was accepted with a 1% false discovery rate (FDR), and the results of the DDA data were incorporated into the spectral library. A total of 327 spectral libraries were used as reference libraries.

[0088] DIA data were analyzed using DIA-NN (v1.7.0). The default settings for DIA-NN were (Precursor FDR: 1%, Log lev: 1, Quality Precision: 20 ppm, MS1 Precision: 10 ppm, Scan Window: 30, Implicit Proteome: Genes, Quantification Strategy: Robust LC (High Precision)). Identification peptides were quantified by averaging the ion peak areas of all chromatographic fragments in the reference libraries. Label-free protein quantification was calculated using the label-free, intensity-based absolute quantification (iBAQ) method. Peak area values ​​for the corresponding proteins were calculated. The Total Quantification (FOT) was used to represent the normalized abundance of a specific protein in the sample. FOT was defined as the protein's iBAQ divided by the total iBAQ of all identified proteins in the sample. For ease of representation, the ft value was multiplied by 10. 5 and with 10 -5 Estimate missing values.

[0089] 4. Proteomics Data Preprocessing

[0090] (1) Quality Control of Mass Spectrometry Platform: To control the mass spectrometry performance during serum sample testing, all 322 samples were pooled into a serum pool as QC standards. The QC standards were analyzed using the same methods and conditions as the serum cohort. The Pearson correlation coefficients of the QC standards were calculated. The average correlation coefficient of the QC standards was 0.85. The minimum correlation coefficient was 0.82, and the maximum correlation coefficient was 0.89, demonstrating the stability of the mass spectrometry analysis platform.

[0091] (2) DIA proteomics data preprocessing and batch correction

[0092] To balance confidence levels in protein identification with sample heterogeneity, we selected and estimated proteins using specific thresholds. First, our analysis focused on proteins identified in more than 50% of samples across each sample type (two tumor subtypes and one healthy control). (The percentage of proteins co-identified across the three serum samples was ≥50%, i.e., the preprocessing threshold was set to 50%). Second, we performed KNN imputation on the data for each sample type separately using the "impute" function from the "impute" R package. Then, we merged the input data from all three sample types: serum samples from healthy individuals, benign breast nodules, and a breast cancer patient cohort. The number of proteomes in each sample ranged from 1723 to 2214, with a median of 1895. Figure 2 (A and B in the original text). Label-free quantification was performed on all patient samples, identifying a total of 8,944 proteomes, with a false discovery rate (FDR) of 1% at the protein and peptide levels. Figure 2 (B in the text). The dynamic range of the identified proteins spanned eight orders of magnitude (…). Figure 3 A total of 8944 proteomes were detected in 322 serum samples. Among them, 5290 proteomes were found in cancer patients and healthy controls. Specific proteomes were found in benign nodules (fibroadenomas, papillomas), breast cancer patients, and healthy controls, with 313, 70, 611, and 208 proteomes respectively. Figure 4 Since the proteins missing differ for each tumor subtype, we used 10... -5 Fill in the blanks. Finally, we use the R tool Combat to remove batch effects by treating tumor type as a covariate.

[0093] Further analysis revealed a relatively significant separation between BC samples and non-BC samples (including fibroadenomas, papillomas, and healthy controls) based on principal component analysis (PCA) of 3000 proteins. Figure 5 The analysis revealed molecular differences between them. PCA analysis showed significant differences in the proteomes of breast and non-breast cancer patients, indicating that the protein differences between the two samples exceeded the differences between individuals. 1,923 and 2,055 upregulated proteins were identified in non-breast cancer and breast cancer, respectively (Fold change > 2, meaning the difference in protein expression between breast and non-breast cancer was at least 2-fold). Notably, 853 specific proteins were identified in non-breast cancer samples (Student's t-test, P < 0.05, n = 391) and in breast cancer samples (Student's t-test, P < 0.05, n = 447). Figure 5 (B in the middle).

[0094] 5. Method for establishing a serum protein profile classifier to differentiate between healthy individuals, breast nodules, and breast cancer patients.

[0095] To elucidate the serum proteomic expression patterns between two benign breast tumors (fibroadenoma and papilloma) and a malignant tumor (breast cancer), we established criteria for defining tumor-specific serum proteins: the expression level of a tumor-specific serum protein in one tumor group should be at least 1.5 times higher than that in either of the other two groups (i.e., Fold > 1.5; Kruskal-Willis test, p < 0.05). We identified 725 tumor-specific serum proteins. We then used differentially expressed proteins and routine blood indicators (such as...) Figure 6 We discovered that proteins associated with neutrophil degranulation were specifically expressed in the malignant tumor group, so we tentatively incorporated the neutrophil ratio into the classifier. We selected the top 10% of differentially expressed proteins from each group and used an extreme value gradient enhancement algorithm (ANOVA test, p<0.05) to determine the protein subset that distinguishes between breast cancer, breast nodules, and normal serum. Secondly, to train and subsequently test the serum proteomic classifier, we divided the samples based on sample type (i.e., healthy individuals (Normal), benign nodules (BBT), and breast cancer (BC)). 70% and 30% of all samples were used as the training and testing sets, respectively. For our selected serum proteomic classifier, composed of 23 proteins and the neutrophil ratio (GR), the 23 proteins are shown in Table 1 below.

[0096] Table 1. Serum protein profile classifier

[0097] 1 P55809 3-ketoacid coenzyme A transferase 1 OXCT1 2 Q9Y3E1 Hepatoma-derived growth factor-related protein 3 HDGFL3 3 Q13838 Spliceosome RNA helicase DDX39B DDX39B 4 P21399 Cytoplasmic aconitate hydratase ACO1 5 Q15020 Squamous cell carcinoma antigen recognized by T-cells 3 SART3 6 Q8WTW3 Conserved oligomeric Golgi complex subunit 1 COG1 7 P61160 Actin-related protein 2 ACTR2 8 P09417 Dihydropteridine reductase QDPR 9 P07911 Uromodulin UMOD 10 Q5VZ89 DENN domain-containing protein 4C DENND4C 11 P01040 Cystatin-A CSTA 12 O75106 Retina-specific copper amine oxidase AOC2 13 Q15493 Regucalcin RGN 14 Q14764 Major Vault Protein MVP 15 Q12931 Heat shock protein TRAP1 16 A0A1B0GUS4 Ubiquitin-conjugating enzyme E2 L5 UBE2L5 17 Q9P2B2 Prostaglandin F2 receptor negative regulator PTGFRN 18 Q8TAQ2 SWI / SNF complex subunit SMARCC2 SMARCC2 19 Q5T1M5 FK506-binding protein 15 FKBP15 20 O95747 Serine / threonine-protein kinase OSR1 OXSR1 21 O43157 Plexin-B1 PLXNB1 22 P02766 Transthyretin TTR 23 P49720 Proteasome subunit beta type-3 PSMB3

[0098] Using 10-fold cross-validation, the serum protein profile classifier achieved an accuracy of 1 and a precision of 1 on the training set. When applied to 30% of the test samples, the mean area under the receiver operating characteristic (ROC) curve (AUC) was 1. Figure 7 , Figure 8 The accuracy is 0.91 and the precision is 0.94.

[0099] 6. Sensitivity and specificity of serum protein profile classifiers in independent cohorts to differentiate between healthy individuals, breast nodules, and breast cancer patients.

[0100] To evaluate the accuracy of the serum proteomic classifier in identifying predictive features of breast tumors, we further evaluated the performance of our model by quantifying proteomics from 103 patients with breast tumors and healthy individuals (benign nodules, n=29; breast cancer, n=57; healthy individuals, n=27) using DIA quantitative proteomics. The AUC was 0.95, the accuracy was 0.87, and the precision was 0.93. Furthermore, heatmaps showed clear separations between breast cancer, benign breast nodules, and healthy individuals in the new independent cohort. Figure 9 , Figure 10 ).

[0101] Example 2: Some optimizations in the construction of a serum protein-based breast tumor classifier

[0102] 1. Serum Sample Pretreatment

[0103] This invention optimizes the sample pretreatment method, the ratio of trypsin to protein, and the trypsin treatment time, with the following results: Figure 11 and Figure 12 The High-Select Top 14 Protein Removal Kit (Thermo Fisher) showed the best results in purifying the five most abundant serum proteins; the ratio of trypsin to protein (mass) was 1:25, and treatment for 17 hours yielded the best results.

[0104] II. Optimization of DIA Proteomics Data Preprocessing Parameters

[0105] We optimized the parameters related to DIA proteomics data preprocessing. First, we adjusted the threshold for the percentage of proteins commonly identified across all groups. This threshold is typically set to 30% or 50%. Our previous experience suggests that a higher threshold percentage requires greater reproducibility of the identified proteins in other samples. Therefore, we tried setting the threshold at 30% and 50% respectively, comparing the results with the same training and validation cohorts. The subsequent protein screening conditions (the expression value of tumor-specific serum proteins in one tumor group must be at least a fold higher than any of the other two groups; when the fold > 1.5, the top two specifically expressed proteins in each group are selected as components of the proteome) remained consistent, and random forest classification was used for both. The results showed that when the parameter was 30%, the AUC of the test cohort was 0.65, and when the parameter was 50%, the AUC was 0.78. Therefore, setting the parameter to 50% clearly yielded more ideal results. The results are as follows... Figures 13-16 (Class 1, Class 2 and Class 3 represent HC (Healthy control), BBT (Benign nodule) and BC (Breast cancer) respectively).

[0106] Secondly, there are many parameters that can be optimized when screening for specific proteins. The first is the fold relative number of proteins identified in each group (the fold by which the expression value of a tumor-specific serum protein in one tumor group is at least higher than that in any of the other two groups). Generally, values ​​of 1.0, 1.5, and 2.0 are chosen (i.e., setting fold > 1.0, fold > 1.5, or fold > 2.0). Therefore, we still used the same training and validation cohorts for comparison, keeping other parameters unchanged. When the fold is > 1.0, we selected the top two specifically expressed proteins in each group as components of the proteome and used a random forest method for classification. The AUC of the test cohort was 0.59. When the fold is > 2.0, the AUC of the test cohort was 0.64, and when the fold is > 1.5, the AUC of the test cohort was 0.78. Therefore, it is clear that the results are more ideal when the fold is > 1.5. The results are as follows... Figures 17-21 (Class 1, Class 2, and Class 3 represent HC (Healthy control), BBT (Benign nodule), and BC (Breast Cancer), respectively).

[0107] In addition to proteomic analysis, our preliminary proteomic analysis of the cohort samples revealed a significant increase in proteins associated with neutrophil degranulation in the serum of breast cancer patients compared to other groups. Therefore, we hypothesized that the neutrophil ratio could be incorporated into the proteomic analysis as a classification indicator for healthy individuals, benign breast tumors, and breast cancer. We compared the diagnostic efficacy of including and not including the neutrophil ratio in the classification analysis. Under the same cohort and with other parameters unchanged, the AUC of the test cohort was 0.78 without the neutrophil ratio as a classification indicator; the AUC of the test cohort was 0.87 with the neutrophil ratio as a classification indicator. Figures 21-24 (Class 1, Class 2 and Class 3 represent HC (Healthycontrol), BBT (Benign nodule) and BC (Breast Cancer) respectively).

[0108] Meanwhile, the number of differentially expressed proteins selected may also affect the diagnostic efficacy of the classifier. Therefore, while keeping other parameters constant, we tried adjusting the criteria for selecting differentially expressed proteins in the classifier. We changed the previous criterion of selecting the top two differentially expressed proteins in each group to selecting the top 10% of the differentially expressed proteins in each group as components of the protein profile, and also incorporated the neutrophil ratio. Using a random forest method for classification, the AUC in the 30% test cohort was 0.99. However, in the new independent validation cohort, its sensitivity and specificity were only 70%. (See attached results). Figures 25-29To further improve the diagnostic performance of the classifier, we changed the machine learning algorithm, using the extreme gradient boosting method for classification. The AUC for the 30% test queue was 1, the AUC for the validation queue was 0.95, and the accuracy was 0.87. Figures 30-35 .

[0109] Comparative Example 1: The protein classifier obtained by this invention compared with other protein classifiers.

[0110] The biggest difference between this classifier and other classifiers is that it does not require the extraction of exosomes and can be detected directly in serum; it can significantly distinguish between healthy people, benign breast tumors, and breast cancer.

[0111] 1. Potential breast cancer tumor biomarkers were screened in serum samples using SELDI mass spectrometry. Proteins with peak values ​​of 4.3 kDa, 8.1 kDa, and 8.9 kDa were selected as potential biomarkers to construct a classifier that distinguished between breast cancer patients and non-cancer controls. The sensitivity was 93% and the specificity was 91%. The AUC was 0.972 (Li J, Zhang Z, Rosenzweig J, et al. Proteomics and bioinformatics approaches for identification of serum biomarkers to detect breast cancer[J]. Clinical chemistry, 2002, 48(8):1296~1304). It could only distinguish between tumors and non-tumors, and could not distinguish between benign tumors.

[0112] 2. Using SELDI-TOF mass spectrometry (MS) to identify differentially expressed proteins in the serum of breast cancer patients and healthy volunteers, a novel prognostic biomarker panel consisting of five serum proteins was constructed. The AUC value was 0.939, with a sensitivity and specificity of 86.6% and 92.4%, respectively, and an overall accuracy of 89.4% (Chung, Li, et al. "Novel serum protein biomarker panel revealed by mass spectrometry and its prognostic value in breast cancer." Breast cancer research 16.3(2014):1-12). This panel can only indicate the patient's prognosis and cannot distinguish between tumors and non-tumor cells, nor can it differentiate between benign tumors.

[0113] 3. MALDI-ToF mass spectrometry was used to analyze serum samples from patients with stage I and II breast cancer and age-matched healthy controls. A classifier constructed from three spectral components was established to distinguish between the control group and early-stage cancer patients, with a sensitivity of 83% and a specificity of 85% (Pietrowska, Monika, et al. Mass spectrometry-based serum proteome pattern analysis in molecular diagnostics of early stage breast cancer. Journal of Translational Medicine 7(2009):1-13). However, it could only distinguish between tumors and non-tumor cells and could not distinguish between benign tumors.

[0114] 4. Using proteomics methods to screen serum samples from 45 breast cancer patients and 46 healthy women, a set of protein peaks consisting of 14 biomarkers was identified that could distinguish breast cancer from non-cancer controls. The sensitivity was 89%, the specificity was 67%, and the area under the receiver operating curve (AUC) was 0.8. This method can detect the difference between breast cancer patients and non-cancer controls. Daniel, et al. "Serum proteome profiling of primary breast cancer indicates a specific biomarker profile." Oncology reports 26.5 (2011): 1051-1056). It can only distinguish between tumors and non-tumors, and cannot distinguish between benign tumors.

[0115] 5. High-resolution SELDI-TOF mass spectrometry analysis was performed on blood samples from patients with negative breast examination results and stage 1 invasive ductal carcinoma. A discriminant spectrum consisting of seven ion peaks was constructed. The sensitivity and specificity in the training set were 95.6% and 86.5%, respectively. In the validation set, the sensitivity was 96.5% and the specificity was 85.7% (Belluco, Claudio, et al. Serum proteomic analysis identifies a highly sensitive and specific discriminatory pattern in stage 1 breast cancer. Annals of Surgical Oncology 14 (2007): 2470-2476). It could only distinguish between tumors and non-tumors, and could not distinguish between benign tumors.

[0116] 6. Serum proteomics maps were analyzed using SELDITOF-MS mass spectrometry. A classification model consisting of apolipoprotein C1, the C-terminal truncated form of C3a, and complement component C3a was constructed to distinguish breast cancer patients from non-cancer controls. The sensitivity and specificity of this model were 96.45% and 94.87%, respectively (Fan, Yuxia et al. "Detection and identification of potential biomarkers of breast cancer." Journal of Cancer Research and Clinical Oncology 136(2010):1243-1254). However, it could only distinguish between tumors and non-tumors and could not distinguish between benign tumors.

[0117] 7. Proteins in the serum of 54 patients with negative axillary lymph nodes, 47 patients with positive axillary lymph nodes and 101 healthy controls were detected by mass spectrometry (MS). A diagnostic model was constructed using four proteins with m / z peaks at 3979, 5643, 6437, and 8929. The sensitivity for identifying healthy controls and breast cancer was 96.4%, the specificity was 88.12%, and the accuracy was 92.08%. A diagnostic model was also constructed using four proteins with m / z peaks at 5643, 4651, 2377, and 2240. This model identified breast cancer with and without axillary lymph node metastasis with a sensitivity of 87.04%, a specificity of 87.23%, and an accuracy of 87.13% (Wang, Liang, et al. "Primary study of lymph node metastasis-related serum biomarkers in breast cancer." The Anatomical Record: Advances in Integrative Anatomy and Evolutionary Biology 294.11(2011):1818–1824).

[0118] 8. Patent 202210107535.3, a set of exosome markers and applications for early diagnosis of breast cancer, although it also constructs a classifier to distinguish between benign tumors, healthy tissues and breast cancer, it requires the extraction of exosomes and subsequent protein detection. The range of proteins detected is completely different, which increases the difficulty of clinical detection. Moreover, the AUC area in the independent validation set is only 0.87, which is significantly lower than our 0.96.

[0119] 9. The article "Proteomic analysis of circulating extracellular vesicles identifies potential markers of breast cancer progression, recurrence, and response" also constructed a breast disease classifier composed of exosome vesicle proteins. It can only distinguish between breast cancer and healthy individuals, but cannot distinguish between benign tumors.

[0120] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. Biomarkers for breast disease screening, characterized in that, include: OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3, and neutrophil ratio; The breast diseases include: benign breast tumors and / or malignant breast tumors; the benign breast tumors include fibroadenomas or papillomas; The malignant breast tumors include breast cancer.

2. The method for screening markers according to claim 1, characterized in that, Includes the following steps: Step 1: After preprocessing and mass spectrometry detection of serum samples from healthy individuals, benign breast tumors, and malignant breast tumors, protein proteome data of the three types of serum samples were obtained through protein identification and quantification. Step 2: The proteomic data of the three serum samples were subjected to quality control, preprocessing, correction and screening to obtain differential proteomic data 1; Step 3: The differential proteome data 1 is combined with blood routine indicators and classified, and the biomarkers are obtained by screening based on the training set and the test set.

3. The screening method according to claim 2, characterized in that, The threshold for the preprocessing is: The percentage of proteins commonly identified in the three serum samples is ≥50%.

4. The screening method according to claim 2, characterized in that, The screening criteria are as follows: In the three serum samples, the expression value of the specific serum protein in any diseased serum sample was ≥ 1.5 times that in the other two serum samples; and the specific serum protein was among the top 10% of the differentially expressed proteins in the multiple groups with the greatest difference. The arbitrary lesion serum sample includes serum samples of benign breast tumors and serum samples of malignant breast tumors.

5. The screening method according to claim 2, characterized in that, The preprocessing includes the following steps: After removing the five most abundant serum proteins from the serum using the High-Select Top 14 Protein Removal Kit, the serum was treated for 17 h at a trypsin:protein mass ratio of 1:

25.

6. The screening method according to claim 2, characterized in that, The classification method is the extreme value gradient enhancement algorithm.

7. The screening method according to claim 2, characterized in that, The conditions for the mass spectrometry detection are as follows: Chromatographic column: C18, particle size 1.9 μm; inner diameter 150 μm, length 15 cm; pore size 120 Å; The chromatographic conditions are as follows: Mobile phase A: 0.1% formic acid and water (volume fraction); Mobile phase B: 0.1% formic acid and acetonitrile (volume fraction); The flow rate is 600 nL / min; The gradient range and duration are: 15%~30% mobile phase B, separation for 75 min.

8. The application of the biomarker as described in claim 1 or the biomarker obtained by the screening method described in any one of claims 2 to 7 in the preparation of breast lesion diagnostic products or the establishment of breast disease screening models; The breast diseases mentioned include: Benign breast tumors and / or malignant breast tumors; the benign breast tumors include fibroadenomas or papillomas; The malignant breast tumors include breast cancer.

9. A breast tumor diagnostic product, characterized in that, This includes markers obtained by screening using the markers described in claim 1 or the screening method described in any one of claims 2 to 7.