Breast tumor screening marker based on serum protein, and application
By performing mass spectrometry detection and data analysis on the serum samples of breast cancer patients, 24 serum characteristics were screened out and breast tumor classifiers were constructed, which solved the problem of lack of serum protein marker profiles in the prior art that distinguished healthy people, benign breast nodules and breast cancer, and achieved high accuracy and high specificity of breast tumor diagnosis.
Patent Information
- Application Number
- PCT/CN2023/142250
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art lacks serum protein marker profiles that can be used to distinguish healthy people, benign breast nodules and breast cancer, and previous studies are mainly based on serum from tumor tissue or patients, and no protein spectrum with intersection between the two has been studied.
By mass spectrometry and data analysis were performed on serum samples from 56 healthy people, 112 benign breast lesions and 154 breast cancer patients, 24 serum characteristics, including 23 proteins and 1 conventional blood test indicator, and a breast tumor classifier based on serum protein was constructed.
It has achieved effective distinction between healthy people, breast nodules patients and breast cancer patients. The classifier has high accuracy and strong specificity in the independent verification cohort, and the area under the ROC curve reaches 0.96, which has good clinical application prospects.
Smart Images

Figure CN2023142250_05062025_PF_FP_ABST
Abstract
Description
Serum protein-based markers for breast tumor screening and their applications
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311613239.1 and invention name “Markers and Applications for Serum Protein-Based Breast Tumor Screening”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of biomedical technology, and in particular to serum protein-based markers for breast tumor screening and their applications. Background Art
[0003] Breast cancer is one of the most common malignant tumors in women worldwide. Due to the lack of non-invasive, highly sensitive, specific and stable early diagnostic markers, the situation of early diagnosis and prevention of breast cancer in my country is grim.
[0004] Blood, due to its easy access, high homogeneity, and rich content, has become a prime candidate for research on novel tumor markers. Proteins in blood, however, are particularly attractive for early diagnosis due to their stability and methodological compatibility. Using mass spectrometry, M. Toi et al. screened the sera of 90 healthy individuals and 54 breast cancer patients and identified RALGAPA2, PKG1, TJP2, and NFX1 as potential diagnostic markers for breast cancer, with an area under the receiver operating characteristic curve (AUC) of 0.874. Furthermore, researchers identified CDH5 as a potential serum marker for metastatic breast cancer, with its expression level significantly differentiating recurrent from non-recurrent breast cancer with a specificity of 90%. Meanwhile, Ladd's team discovered that the serum breast cancer proteins DUSP9, EED, EFNA5, ITGB1, and PPMT1 may serve as early-stage markers for triple-negative breast cancer.
[0005] However, due to the lack of benign breast lesions and full-course breast cancer cohorts in previous studies on breast cancer protein markers, there is currently a lack of serum protein marker spectra that can be used to distinguish healthy people, benign breast nodules, and breast cancer. In addition, previous scholars screened protein markers based on either tumor tissue or the serum of tumor patients, and did not study the protein spectra that overlap between the two. Therefore, it is crucial to develop a serum protein marker spectra based on patient serum and protein markers that can be used to distinguish healthy people, benign breast nodules, and breast cancer.
[0006] Summary of the Invention
[0007] In view of this, the technical problem to be solved by the present invention is to provide markers for breast tumor screening based on serum proteins and their applications.
[0008] The present invention provides markers for breast disease screening, which include: OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3 and neutrophil ratio GR.
[0009] Further,
[0010] Breast diseases include benign breast lesions and / or malignant breast tumors; benign breast tumors include fibroadenomas or papillomas; and malignant breast tumors include breast cancer. The markers described in the present invention include 23 serum proteins and a neutrophil ratio. The test results of the present invention show that the diagnostic efficacy of the classification test spectrum is significantly improved when the neutrophil ratio is added compared to when it is not added.
[0011] The present invention provides a method for preparing the marker, which comprises the following steps:
[0012] Step 1: Pre-processing and mass spectrometry analysis of serum samples from healthy individuals, serum samples from benign breast lesions, and serum samples from malignant breast tumors, followed by protein identification and quantification, to obtain proteomic data for the three serum samples;
[0013] Step 2: The proteomic data of the three serum samples are subjected to quality control, preprocessing, correction and screening to obtain differential proteomic data 1;
[0014] Step 3: The differential proteomic data 1 is combined with the blood routine indexes and then classified, and the markers are screened and obtained based on the training set and the test set.
[0015] The pretreatment includes the following steps: after removing the five most abundant serum proteins in the serum using a High-Select Top 14 protein removal kit, the cells were treated for 17 hours at a mass ratio of trypsin to protein of 1:25.
[0016] The method further comprises a drying step after the pretreatment; the dried sample is subjected to mass spectrometry detection using an aqueous solution containing 0.1% formic acid by volume.
[0017] The false discovery rate (FDR) for the protein identification was 1%;
[0018] The quantitative parameters were set as follows: Precursor FDR was 1%; Log lev was 1; mass accuracy was 20 ppm; MS1 accuracy was 10 ppm; scanning window was 30; the latent proteome was genes; and the quantitative strategy was robust LC (high precision).
[0019] In step 2, the proteomic data of the three serum samples were pre-processed and a quality control step was also included. The quality control step was to mix all the three serum samples into a serum pool as a QC standard. The QC standard was analyzed using the same method and conditions as the serum cohort. The Pearson correlation coefficient of the QC standard was calculated. The average correlation coefficient of the QC standard was 0.85. The minimum correlation coefficient was 0.82, and the maximum correlation coefficient was 0.89, demonstrating the stability of the mass spectrometry analysis platform.
[0020] The threshold value of the preprocessing is:
[0021] Among the three serum samples, the protein species in each serum sample overlapped with those in the other two serum samples by ≥50%.
[0022] In the present invention, the threshold setting of the preprocessing will affect the credibility or reproducibility of protein identification, and further affect the diagnostic performance; in the present invention, the threshold of the preprocessing is optimized; in some embodiments of the present invention, the threshold is set to 30%; then the diagnostic performance is significantly reduced.
[0023] The correction is to use impute to perform KNN imputation on the data of each sample type to fill the missing values, and the filling parameter is set to 10 -5 .
[0024] The screening criteria are: among the three serum samples, the expression value of the specific serum protein in any diseased serum sample is ≥ 1.5 times the expression value of the other two groups; and the specific serum protein is in the top 10% with the largest difference among the multiple groups of differentially expressed proteins;
[0025] The arbitrary lesion serum samples include benign breast tumor lesion serum samples and malignant breast tumor serum samples.
[0026] The present invention optimizes the screening criteria. Experimental results, combined with the pretreatment threshold, show that only when the pretreatment threshold is set to 50% and the expression value of the specific serum protein in any lesion serum sample is ≥1.5 times the expression value of the other two groups, combined with the neutrophil ratio in the blood routine index, can the diagnosis of breast lesions be significantly better than other conditions. Furthermore, the number of differentially expressed proteins also affects the diagnostic efficacy of the screened markers. Comparing the top 10% of differentially expressed proteins in each group with the top two groups with the largest differences in differentially expressed proteins across multiple groups, selecting the top 10% with the largest differences in differentially expressed proteins across multiple groups is more suitable for screening markers with high specificity and accuracy according to the present invention.
[0027] The classification methods include random forest algorithm, support vector machine algorithm, adaptive boosting algorithm, extreme gradient boosting algorithm, etc. In a specific embodiment of the present invention, random forest algorithm and extreme gradient boosting algorithm are used for classification. The results show that the extreme gradient boosting algorithm is more suitable for the screening of the markers of the present invention.
[0028] Furthermore, in the preparation method of the present invention,
[0029] The parameters of the mass spectrometry detection include:
[0030] Chromatographic column: C18, particle size 1.9 μm; inner diameter 150 μm, length 15 cm; pore size
[0031] The chromatographic conditions are:
[0032] Mobile phase A: 0.1% formic acid and water;
[0033] Mobile phase B: 0.1% formic acid and acetonitrile;
[0034] The flow rate was 600 nL / min;
[0035] The gradient range and duration were: 15% to 30% mobile phase B, separation time 75 min.
[0036] Mass spectrometry conditions: MS1 scans were performed from 300-1,400 m / z at 60 k resolution (AGC target 4e5 or 50 ms). Subsequently, 30 DIA fragments were acquired at 15 k resolution with an AGC target of 5e4 or a maximum injection time of 22 ms. The "Inject ions for all available parallel time" setting was enabled. High-energy collisional dissociation (HCD) fragmentation was set to 30% of the normalized collision energy. Spectra were recorded in profile mode. The default charge state for log-mean MS2 was set to 3.
[0037] The markers described in the present invention are obtained through the optimization of multiple parameters and are capable of accurately diagnosing breast lesions. Compared to other methods that can only diagnose tumors and non-tumors, the markers described in the present invention can diagnose healthy individuals, benign breast lesions, and malignant breast tumors with high diagnostic accuracy and specificity. The various parameters in the marker preparation method of the present invention are coordinated with each other. The pre-processing threshold is: the overlap of protein types in each serum sample with the other two serum samples is ≥50%; the screening criteria are: the expression value of a specific serum protein in any lesion serum sample among the three serum samples is ≥1.5 times the expression value of the other two groups; and the specific serum protein is in the top 10% of the differentially expressed proteins with the greatest difference among the multiple groups. The highly accurate and specific diagnostic markers described in the present invention can only be obtained when the extreme gradient boosting algorithm is used.
[0038] The present invention provides the use of the marker or the marker prepared by the preparation method in preparing a breast lesion diagnosis product or establishing a breast disease screening model.
[0039] The present invention provides a breast tumor diagnostic product, which includes or utilizes the marker of the present invention or the marker prepared by the preparation method.
[0040] The present invention constructs a serum protein-based breast tumor marker. Serum testing can be performed on patients suspected of breast tumors to effectively distinguish healthy people, patients with breast nodules, and breast cancer patients. In an independent validation cohort, the classifier achieved an accuracy of 0.87, a precision of 0.93, a recall of 0.82, and an F1-score of 0.86. The area under the ROC curve for benign nodules was 0.96, the area under the ROC curve for breast cancer was 0.96, and the area under the ROC curve for healthy people was 0.96. With high accuracy and good specificity, the classifier has good application prospects in the clinical diagnosis of breast lesions. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is an analysis flow chart;
[0042] FIG2 is a schematic diagram of proteins specifically expressed in healthy controls, fibroadenomas, papillomas, and breast cancer;
[0043] Figure 3 shows the number of proteomes of all patients at a false discovery rate (FDR) of 1% at the protein and peptide levels;
[0044] Figure 4 shows the specific expression of proteins in healthy controls, fibroadenomas, papillomas, and breast cancer. A shows the distribution of breast-related proteins among the groups; B shows the distribution of serum proteins among the groups; C shows the distribution of related proteins with drug targets among the groups; D shows the distribution of FDA-approved drugs among the groups; E shows the distribution of membrane surface protein markers among the groups; and F shows the distribution of tumor-related proteins among the groups.
[0045] Figure 5 shows the principal component analysis and expression difference volcano plot of healthy control (HC), fibroadenoma (Fibroma), papilloma (Papillomas) and breast cancer (breast cancer), where A is the principal component analysis; B is the expression difference volcano plot;
[0046] FIG6 shows the tumor-specific serum proteins identified according to the tumor-specific serum protein standard, wherein A is the expression heat map of each group of specific proteins in different groups; B is the cluster analysis of each group of specific proteins;
[0047] Figure 7 shows the classification results of the protein spectrum classifier in the training cohort;
[0048] FIG8 is a heat map of the discovery cohort of 24 protein profiles;
[0049] Figure 9 shows the classification results of the protein spectrum classifier in an independent validation cohort;
[0050] FIG10 is a heat map of the validation cohort of 24 protein profiles;
[0051] Figure 11 is a screening of a high abundance protein removal kit;
[0052] Figure 12 shows the optimization of pretreatment conditions, where A is the ratio of enzyme to protein and B is the trypsin treatment time;
[0053] Figure 13 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 30%. When the Fold of the screening specimens was greater than 1.5 times, the top 2 specifically expressed proteins in each group were selected as the components of the protein spectrum. Classification was performed using the random forest method. Figure A is the ROC curve of the classification results, and Figure B is the heat map of the classification results.
[0054] Figure 14 shows the classification results of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens, with 30% of the samples in this cohort used as the validation cohort. Figure A shows the receiver operating characteristic (ROC) curve for the classification results, and Figure B shows the heat map of the classification results.
[0055] Figure 15 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and when the sample Fold > 1.5 was screened, the top 2 specifically expressed proteins in each group were selected as protein spectrum components. Classification was performed using the random forest method. Figure A is the ROC curve of the classification results, and Figure B is the heat map of the classification results.
[0056] Figure 16 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 30% of the samples in this cohort were used as the validation cohort. When the data preprocessing parameter was set to 50%, the top two specifically expressed proteins in each group were selected as protein spectrum components when the Fold ratio was greater than 1.5. Classification was performed using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0057] Figure 17 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and specimens were screened for Fold > 1.0. The top two specifically expressed proteins in each group were selected as protein profile components for classification using the random forest method. Figure A is the ROC curve of the classification results, and Figure B is the heat map of the classification results.
[0058] Figure 18 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 30% of the samples in this cohort were used as the validation cohort. The data preprocessing parameter was set to 50%, and the specimens were screened for Fold>1.0. The top two specifically expressed proteins in each group were selected as protein spectrum components and classified using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0059] Figure 19 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and specimens were screened for Fold > 2.0. The top 2 specifically expressed proteins in each group were selected as protein profile components for classification using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0060] Figure 20 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 30% of the samples in this cohort were used as the validation cohort. The data preprocessing parameter was set to 50%, and the specimens were screened for Fold>2.0. The top 2 specifically expressed proteins in each group were selected as the components of the protein spectrum and classified using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0061] Figure 21 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and specimens with Fold > 1.5 were screened. The top two specifically expressed proteins in each group were selected as protein profile components and classified using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0062] Figure 22 shows 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens. 30% of the samples in this cohort were used as the validation cohort. The data preprocessing parameter was set to 50%, and specimens were screened for Fold > 1.5. The top two specifically expressed proteins in each group were selected as protein profile components and classified using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0063] Figure 23 shows a sample of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and samples were screened for Fold > 1.5. The top two specifically expressed proteins in each group were selected as protein profile components, and the neutrophil ratio was added as a classification indicator. Classification was performed using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0064] Figure 24 shows a validation cohort of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 30% of the samples were used as the validation cohort. The data preprocessing parameter was set to 50%, and samples were screened for Fold > 1.5. The top two specifically expressed proteins in each group were selected as protein profile components, and the neutrophil ratio was added as a classification indicator. Classification was performed using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0065] Figure 25 shows a sample of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and samples were screened for Fold > 1.5. The top 10% of specifically expressed proteins in each group were selected as protein profile components. The neutrophil ratio was added as a classification indicator, and classification was performed using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0066] Figure 26 shows a validation cohort of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 30% of the samples were used as the validation cohort. The data preprocessing parameter was set to 50%, and samples were screened for Fold > 1.5. The top 10% of specifically expressed proteins in each group were selected as protein profile components. Neutrophil ratio was added as a classification indicator, and classification was performed using the random forest method. A is the ROC curve of the classification results, and B is the heat map of the classification results.
[0067] Figure 27 shows the classification results obtained by validating the classifier using a new independent sample cohort consisting of 57 breast cancer samples, 27 healthy subjects, and 29 benign breast nodule samples under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, and adding the neutrophil ratio as a classification indicator, using the random forest method for classification).
[0068] Figure 28 shows a proteomic heatmap obtained by validating the classifier using a new independent sample cohort consisting of 57 breast cancer samples, 27 healthy subjects, and 29 benign breast nodule samples under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, adding the neutrophil ratio as a classification metric, and using the random forest method for classification).
[0069] Figure 29 shows the importance of each indicator in the protein profile obtained by the classifier under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, adding the neutrophil ratio as a classification indicator, and using the random forest method for classification);
[0070] Figure 30 shows a sample of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 70% of the samples in this cohort were used as the training cohort. The data preprocessing parameter was set to 50%, and samples were screened for Fold > 1.5. The top 10% of specifically expressed proteins in each group were selected as protein profile components. The neutrophil ratio was added as a classification indicator, and classification was performed using the extreme value gradient boosting algorithm.
[0071] Figure 31 shows a validation cohort of 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy subjects. 30% of the samples were used as the validation cohort. The data preprocessing parameter was set to 50%, and when Fold > 1.5 was selected for screening, the top 10% of specifically expressed proteins in each group were selected as protein profile components. Neutrophil ratio was added as a classification indicator, and classification was performed using the extreme value gradient boosting algorithm.
[0072] Figure 32 shows the classification results obtained by validating the classifier obtained under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, adding the neutrophil ratio as a classification indicator, and using the extreme gradient boosting algorithm for classification) using a new independent sample cohort consisting of 57 breast cancer samples, 27 healthy subjects, and 29 benign breast nodule samples.
[0073] Figure 33 shows a proteomic heatmap obtained by validating the classifier obtained under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, adding the neutrophil ratio as a classification metric, and using the extreme value gradient boosting algorithm for classification) using 151 breast cancer specimens, 105 benign breast nodule specimens, and 55 healthy human specimens.
[0074] Figure 34 shows a proteomic heatmap obtained by validating the classifier using a new independent sample cohort consisting of 57 breast cancer samples, 27 healthy subjects, and 29 benign breast nodule samples under these conditions (data preprocessing parameters set to 50%, screening samples with Fold>1.5, selecting the top 10% of specifically expressed proteins in each group as protein profile components, adding the neutrophil ratio as a classification metric, and using the extreme gradient boosting algorithm for classification).
[0075] Figure 35 shows the importance of each indicator of the classified protein spectrum obtained by the classifier under this condition (when the data preprocessing parameters are set to 50%, when the screening specimen Fold>1.5, the TOP10% of the specifically expressed proteins in each group are selected as the components of the protein spectrum, and the neutrophil ratio is added to the classification indicators, and classification is performed using the extreme gradient enhancement algorithm). DETAILED DESCRIPTION
[0076] The present invention provides a breast tumor classifier based on serum protein. Those skilled in the art can learn from the content of this article and appropriately improve the process parameters to achieve it. It is particularly important to point out that all similar replacements and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications herein without departing from the content, spirit and scope of the present invention to achieve and apply the technology of the present invention.
[0077] At present, due to the lack of benign breast lesions and breast cancer full-course cohorts in previous studies on breast cancer protein markers, there is currently a lack of serum protein marker spectra suitable for distinguishing healthy people, benign breast nodules, and breast cancer. In addition, previous scholars screened protein markers based on tumor tissue or serum of tumor patients, and did not study the protein spectra that overlapped between the two. Therefore, this study found 24 serum markers of 56 healthy people, 112 benign breast lesions, and 154 breast cancers (carcinoma in situ, microinvasion, invasive carcinoma, lymph node metastasis, and distant metastasis) by analyzing the serum of 56 healthy people, 112 benign breast lesions, and 154 breast cancers (carcinoma in situ, microinvasion, invasive carcinoma, lymph node metastasis, and distant metastasis). The serum characteristics include 23 proteins and 1 routine blood test indicator (OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3, GR) which can effectively distinguish healthy people, patients with breast nodules, and breast cancer patients.
[0078] We used LC-MS mass spectrometry to test the serum of healthy individuals and cohorts of patients with benign breast nodules and breast cancer, screening for proteins that are specifically expressed in both tissues and serum at various stages of disease progression. At the same time, by establishing a serum protein classifier (24 characteristic spectra), we were able to distinguish between benign and malignant breast tumors.
[0079] The reagents and consumables used in the present invention are all common commercial products and can be purchased in the market.
[0080] The present invention will be further described below in conjunction with the embodiments:
[0081] Example 1 Construction of a serum protein-based breast tumor classifier
[0082] 1. Construction of a serum protein-based breast tumor classifier
[0083] 1. Serum sample pretreatment
[0084] Serum samples from 56 healthy individuals, 112 benign breast lesions, and 154 breast cancer cases (carcinoma in situ, microinvasion, invasive carcinoma, lymph node metastasis, and distant metastasis) were pre-treated;
[0085] First, the five most abundant serum proteins (LGKC, IGHG1, ALB, IGHG2, and APOA1) were removed using a commercial kit (Thermo Fisher, A36369) according to the manufacturer's instructions and then inactivated at 85°C for 10 minutes. The serum samples, free of high-abundance serum proteins, were then treated with trypsin at 37°C for 17 hours (trypsin:protein (mass) = 1:25), and peptides were extracted and dried.
[0086] 2. Mass spectrometry
[0087] The samples were measured using an LC-MS instrument consisting of an EASY-nLC 1200 ultrahigh pressure system (Thermo Fisher Scientific) coupled to a Fusion Lumos Orbitrap (Thermo Fisher Scientific) for plasma samples via a nanoelectrospray ionization source (Thermo Fisher Scientific). Serum samples were dissolved in 12 μL of loading buffer (0.1 (v / v)% formic acid in water) for peptides, and 5 μL of the sample was loaded onto a 100 μL id×2.5 cm, C18 trapping column with 14 μL of solvent a (0.1 (v / v)% formic acid in water) at a maximum pressure of 280 bar. The peptides were concentrated on a 150 μm id×15 cm column (C18, 1.9 μm, The separation was performed on a HPLC-MS / ...
[0088] 3. Protein identification and quantification methods
[0089] All data were processed using Firmiana. DIA data were searched against the UniProt human protein database (updated 2019.12.17, 20406 entries) using FragPipe (v12.1) and MSFragger (2.2)22. The mass tolerance for precursors was 20 ppm and the mass tolerance for product ions was 50 mmu. A maximum of two isoforms were allowed to be missed. The search engine used cysteine aminomethylation as a fixed modification and N-acetylation and methionine oxidation as variable modifications. The precursor ion fraction charge limit was +2, +3 and +4. The data were also searched against a decoy database so that protein identifications were accepted at a false discovery rate (FDR) of 1%, and the results of the DDA data were merged into the spectral library. A total of 327 spectral libraries were used as reference spectral libraries.
[0090] DIA data were analyzed using DIA-NN (v1.7.0). The default settings of DIA-NN were (Precursor FDR: 1%, Log lev: 1, mass accuracy: 20 ppm, MS1 accuracy: 10 ppm, scan window: 30, implicit proteome: gene, quantification strategy: robust LC (high accuracy)). The quantification of identified peptides was calculated as the average of the peak areas of chromatographic fragment ions in all reference spectral libraries. Label-free protein quantification was calculated using the label-free, intensity-based absolute quantification (iBAQ) method. We calculated the peak area values of the corresponding proteins. The total fraction (FOT) was used to represent the normalized abundance of a specific protein in the sample. The FOT is defined as the iBAQ of a protein divided by the total iBAQ of all identified proteins in the sample. For ease of representation, the ft value was multiplied by 10 5 and use 10 -5 Impute missing values.
[0091] 4. Proteomic Data Preprocessing
[0092] (1) Quality control of the mass spectrometry platform: To perform quality control of the mass spectrometry performance during the serum sample testing process, all 322 samples were mixed into a serum pool as QC standards. The QC standards were analyzed using the same methods and conditions as the serum cohort. The Pearson correlation coefficients of the QC standards were calculated. The average correlation coefficient of the QC standards was 0.85. The minimum correlation coefficient was 0.82, and the maximum correlation coefficient was 0.89, demonstrating the stability of the mass spectrometry analysis platform.
[0093] (2) DIA proteomics data preprocessing and batch correction
[0094] Considering the balance between confidence in protein identification and sample heterogeneity, we selected proteins using specific thresholds for imputation. First, our analysis focused on proteins identified in more than 50% of samples within each sample type (two tumor subtypes and one healthy control) (the percentage of proteins identified across the three serum samples was ≥50%, meaning the preprocessing threshold was set to 50% here). Second, we performed KNN imputation on the data for each sample type separately using the "impute" function from the "impute" R package. We then combined the input data from all three sample types: serum samples from healthy individuals, benign breast nodules, and breast cancer patient cohorts. The number of proteomes per sample ranged from 1723 to 2214, with a median of 1895 (Figure 2, A and B). Label-free quantitative measurements across all patient samples revealed a total of 8,944 proteomes with a false discovery rate (FDR) of 1% at both the protein and peptide levels (Figure 2, B). The dynamic range of identified proteins spanned eight orders of magnitude (Figure 3). A total of 8944 protein groups were detected in 322 serum samples, of which 5290 protein groups were shared by cancer patients and healthy controls, and 313, 70, 611, and 208 protein groups were specific to benign nodules (fibroadenomas, papillomas), breast cancer patients, and healthy controls, respectively (Figure 4). -5 Finally, we used the R tool Combat to remove batch effects by including tumor type as a covariate.
[0095] Further analysis revealed that principal component analysis (PCA) based on 3,000 proteins showed a relatively clear separation between BC samples and non-BC samples (including fibroadenomas, papillomas, and healthy controls) (Figure 5A), revealing the molecular differences between them. PCA analysis showed that the proteomes of breast cancer and non-breast cancer were significantly different, indicating that the protein differences between the two samples exceeded the differences between individuals. 1,923 and 2,055 upregulated proteins (Fold change > 2, the difference in expression between breast cancer proteins and non-breast cancer proteins was at least 2-fold) were identified in non-breast cancer samples and breast cancer samples, respectively. Notably, 853 specific proteins were identified in non-breast cancer samples (Student's t-test, P < 0.05, n = 391) and breast cancer samples (Student's t-test, P < 0.05, n = 447) (Figure 5B).
[0096] 5. Method for establishing a serum protein profile classifier to distinguish healthy individuals, breast nodules, and breast cancer patients
[0097] To elucidate serum proteomic expression patterns between two types of benign breast tumors (fibroadenomas and papillomas) and a malignant tumor (breast cancer), we established criteria for defining tumor-specific serum proteins: the expression of a tumor-specific serum protein in one tumor group should be at least 1.5-fold higher than in either of the other two groups (i.e., Fold > 1.5; Kruskal-Willis test, p < 0.05). We identified 725 tumor-specific serum proteins. We then classified the top 10% of differentially expressed proteins in each group using the extreme gradient boosting algorithm (ANOVA test, p < 0.05) based on differentially expressed proteins and routine blood tests (as shown in Figure 6 , proteins associated with neutrophil degranulation were found to be specifically expressed in the malignant tumor group, so we experimentally added neutrophil proportion to the classifier). We then used the extreme gradient boosting algorithm (ANOVA test, p < 0.05) to identify a subset of proteins that distinguish breast cancer, breast nodules, and normal serum. Secondly, in order to train and subsequently test the serum protein profile classifier, we divided the samples based on sample type (i.e., healthy people (Normal), benign nodules (BBT), breast cancer (BC)), and used 70% and 30% of all samples as training sets and test sets respectively. The serum protein profile classifier we screened consists of 23 proteins and neutrophil ratio (GR), as shown in Table 1 below:
[0098] Table 1. Serum protein profile classifiers
[0099] Using 10-fold cross-validation, the serum protein profile classifier was found to have an accuracy of 1 and a precision of 1 in the training set. When applied to 30% of the test samples, the average area under the receiver operating characteristic (ROC) curve (AUC) was 1 (Figures 7 and 8), with an accuracy of 0.91 and a precision of 0.94.
[0100] 6. Sensitivity and specificity of the serum protein profile classifier for distinguishing healthy individuals, breast nodules, and breast cancer patients in an independent cohort
[0101] To evaluate the accuracy of the discriminative serum protein profile classifier for predictive features of breast tumors, we further evaluated the performance of our model to quantify the proteomics data from 103 breast tumor and healthy patients (benign nodules, N = 29; breast cancer, n = 57; healthy subjects, N = 27). Using DIA quantitative proteomics, we achieved an AUC of 0.95, an accuracy of 0.87, and a precision of 0.93. In addition, heat maps showed that there was a clear separation between breast cancer, benign breast nodules, and healthy subjects in the new independent cohort (Figures 9 and 10).
[0102] Example 2: Some optimizations in the construction of a serum protein-based breast tumor classifier
[0103] 1. Serum sample pretreatment
[0104] The present invention optimizes the sample pretreatment method, the ratio of trypsin to protein, and the trypsin treatment time. The results are shown in Figures 11 and 12. The High-Select Top 14 Protein Removal Kit (Thermo Fisher) was most effective in purifying the five most abundant serum proteins in serum; the best effect was achieved with a trypsin:protein (mass) ratio of 1:25 and a treatment time of 17 hours.
[0105] 2. Optimization of DIA proteomics data preprocessing parameters
[0106] We optimized the parameters for DIA proteomics data preprocessing. First, we adjusted the threshold for the percentage of proteins commonly identified across all groups. This threshold was typically set at 30% or 50%. Our previous experience suggests that a higher threshold increases the requirement for reproducibility of identified proteins across other samples. Therefore, we experimented with thresholds of 30% and 50%, respectively. Comparisons were performed using the same training and validation cohorts. Subsequent protein selection criteria (tumor-specific serum protein expression in one tumor group must be at least a fold higher than in either of the other two groups; if the fold is >1.5, the top two proteins specifically expressed in each group were selected as protein profile components) were used. Random forest classification was used for both cohorts. The results showed that a 30% threshold yielded an AUC of 0.65 for the test cohort, while a 50% threshold yielded an AUC of 0.78. Therefore, a 50% threshold clearly yielded more optimal results. The results are shown in Figures 13 to 16 (class 1, class 2, and class 3 represent HC (Healthy control), BBT (benign nodule), and BC (breast cancer) respectively).
[0107] Secondly, when screening specific proteins, there are also many parameters that can be optimized. The first is the relative number multiple of identified proteins in each group (the expression value of tumor-specific serum protein in one tumor group is at least fold higher than that of any of the other two groups), which is generally selected as 1.0, 1.5, 2.0 (i.e., setting fold>1.0, fold>1.5 or fold>2.0). Therefore, we still use the same set of training cohorts and validation cohorts for comparison, and other parameters remain unchanged. When Fold>1.0 times, the top 2 specifically expressed proteins in each group are selected as the components of the protein spectrum, and the random forest method is used for classification. The AUC of the test cohort is 0.59, and when Fold>2.0 times, the AUC of the test cohort is 0.64, and when Fold>1.5, the AUC of the test cohort is 0.78. Therefore, it is obvious that when Fold>1.5, the result is more ideal. The results are shown in Figures 17 to 21 (class 1, class 2, and class 3 represent HC (Healthy control), BBT (benign nodule), and BC (breast cancer), respectively);
[0108] In addition to protein spectrum analysis, our preliminary proteome analysis of the cohort specimens revealed that neutrophil degranulation-related proteins were significantly elevated in breast cancer serum compared with those in other groups. Therefore, we speculated that the neutrophil ratio could be added to the protein profile as an indicator for classifying healthy individuals, benign breast tumors, and breast cancer. Therefore, we compared the diagnostic efficacy of the profile with and without the neutrophil ratio. In the same cohort, with other parameters remaining unchanged, the AUC for the test cohort was 0.78 when the neutrophil ratio was not included as a classification indicator; when the neutrophil ratio was included as a classification indicator, the AUC for the test cohort was 0.87. The results are shown in Figures 21 to 24 (class 1, class 2, and class 3 represent HC (Healthy Control), BBT (Benign Nodules), and BC (Breast Cancer), respectively).
[0109] At the same time, the number of differentially expressed proteins selected may also affect the diagnostic efficacy of the classifier. Therefore, while keeping other parameters unchanged, we tried to adjust the criteria for differentially expressed proteins to be included in the classifier. We changed the previous criterion of selecting the top 2 differentially expressed proteins in each group to selecting the top 10% of differentially expressed proteins in each group as the protein spectrum components. At the same time, we added the neutrophil ratio and used the random forest method for classification. The AUC for the 30% test cohort was 0.99, but in the new independent validation cohort, its sensitivity and specificity were only 70%. The results are shown in Figures 25 to 29. To further improve the diagnostic efficacy of the classifier, we replaced the machine learning algorithm and used the extreme gradient boosting algorithm for classification. The AUC for the 30% test cohort was 1, the AUC for the validation cohort was 0.95, and the accuracy was 0.87 (Figures 30 to 35).
[0110] Comparative Example 1 Comparison of the protein classifier obtained by the present invention with other protein classifiers
[0111] The biggest difference between this classifier and other classifiers is that it does not require the extraction of exosomes and can be directly detected in the serum; it can significantly distinguish between healthy people, benign breast tumors, and breast cancer.
[0112] 1. Using SELDI mass spectrometry, potential breast cancer biomarkers were screened in serum samples. Proteins with peak values of 4.3 kDa, 8.1 kDa, and 8.9 kDa were selected as potential biomarkers for constructing a classifier that distinguished breast cancer patients from non-cancer controls with a sensitivity of 93% and a specificity of 91%. The AUC was 0.972 (Li J, Zhang Z, Rosenzweig J, et al. Proteomics and bioinformatics approaches for identification of serum biomarkers to detect breast cancer [J]. Clinical chemistry, 2002, 48(8):1296-1304). This classifier can only distinguish between tumors and non-tumors, but cannot distinguish benign tumors.
[0113] 2. SELDI-TOF mass spectrometry (MS) was used to identify differentially expressed proteins in the serum of breast cancer patients and healthy volunteers, and a new prognostic biomarker panel consisting of five serum proteins was constructed with an AUC value of 0.939, a sensitivity and specificity of 86.6% and 92.4%, respectively, and an overall accuracy of 89.4% (Chung, Li, et al. "Novel serum protein biomarker panel revealed by mass spectrometry and its prognostic value in breast cancer." Breast cancer research 16.3(2014):1-12). This panel can only indicate the patient's prognosis, but cannot distinguish between tumors and non-tumors, and cannot distinguish benign tumors.
[0114] 3. MALDI-ToF mass spectrometry was used to detect serum samples from patients with stage I and II breast cancer and age-matched healthy controls. A classifier constructed from three spectral components was established to distinguish between controls and early-stage cancer patients, with a sensitivity of 83% and a specificity of 85% (Pietrowska, Monika, et al. "Mass spectrometry-based serum proteome pattern analysis in molecular diagnostics of early stage breast cancer." Journal of Translational Medicine 7(2009):1-13). However, this classifier could only distinguish between tumors and non-tumors, but not benign tumors.
[0115] 4. Using proteomics methods to screen serum samples from 45 breast cancer patients and 46 healthy women, a group of 14 biomarkers was found that can distinguish between breast cancer and non-cancer controls. The sensitivity was 89%, the specificity was 67%, and the area under the receiver operating characteristic curve was 0.8, which can detect the difference between breast cancer patients and non-cancer controls ( Daniel, et al. "Serum proteome profiling of primary breast cancer indicates a specific biomarker profile." Oncology reports 26.5 (2011): 1051-1056), can only distinguish between tumors and non-tumors, but cannot distinguish benign tumors.
[0116] 5. High-resolution SELDI-TOF mass spectrometry analysis was performed on blood samples from patients with negative breast examination results and stage 1 invasive ductal carcinoma, and a discriminant pattern consisting of seven ion peaks was constructed. The sensitivity and specificity in the training set were 95.6% and 86.5%, respectively; in the validation set, the sensitivity was 96.5% and the specificity was 85.7% (Belluco, Claudio, et al. "Serum proteomic analysis identifies a highly sensitive and specific discriminatory pattern in stage 1 breast cancer." Annals of Surgical Oncology 14 (2007): 2470-2476). This method can only distinguish between tumors and non-tumors, but cannot distinguish benign tumors.
[0117] 6. Using SELDITOF-MS mass spectrometry to analyze serum proteomic profiles, a classification model consisting of apolipoprotein CI, a C-terminal truncated form of C3a, and complement component C3a was constructed to distinguish breast cancer patients from non-cancer controls. The sensitivity and specificity of this model were 96.45% and 94.87%, respectively (Fan, Yuxia et al. "Detection and identification of potential biomarkers of breast cancer." Journal of cancer research and clinical oncology 136 (2010): 1243-1254). However, this model can only distinguish between tumors and non-tumors, but cannot distinguish between benign tumors.
[0118] 7. Proteins in the serum of 54 axillary lymph node-negative and 47 axillary lymph node-positive breast cancer patients and 101 healthy controls were detected by mass spectrometry (MS). Diagnostic model 1 was constructed using four proteins with m / z peaks of 3979, 5643, 6437, and 8929, with a sensitivity of 96.4%, a specificity of 88.12%, and an accuracy of 92.08% for identifying healthy controls and breast cancer; a diagnostic model was constructed using four proteins with m / z peaks of 5643, 4651, 2377, and 2240, with a sensitivity of 87.04%, a specificity of 87.23%, and an accuracy of 87.13% for identifying breast cancer with and without axillary lymph node metastasis (Wang, Liang, et al. "Primary study of lymph node metastasis-related serum biomarkers in breast cancer." The Anatomical Record: Advances in Integrative Anatomy and Evolutionary Biology 294.11(2011):1818-1824).
[0119] 8. Patent 202210107535.3, a set of exosome markers and applications for early diagnosis of breast cancer. Although a classifier was constructed to distinguish between benign tumors, healthy tissues, and breast cancer, it requires the extraction of exosomes and subsequent protein detection. The detected protein range is completely different, which increases the difficulty of clinical detection. In addition, the AUC area in the independent validation set is only 0.87, significantly lower than our 0.96.
[0120] 9. The article "Proteomic analysis of circulating extracellular vesicles identifies potential markers of breast cancer progression, recurrence, and response" also constructed a breast disease classifier composed of exosome vesicle proteins. It can only distinguish breast cancer from healthy people, but cannot distinguish benign tumors.
[0121] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. Markers for breast disease screening, It is characterized in that include: OXCT1, HDGFL3, DDX39B, ACO1, SART3, COG1, ACTR2, QDPR, UMOD, DENND4C, CSTA, AOC2, RGN, MVP, TRAP1, UBE2L5, PTGFRN, SMARCC2, FKBP15, OXSR1, PLXNB1, TTR, PSMB3, and neutrophil ratio.
2. The marker according to claim 1, It is characterized in that The breast diseases include: benign breast tumor lesions and / or breast malignant tumors; the benign breast tumor lesions include fibroadenoma or papilloma; the malignant breast tumor includes breast cancer.
3. A method for preparing a marker according to claim 1 or 2, It is characterized in that The steps include: Step 1: Pre-processing and mass spectrometry detection of healthy human serum samples, benign breast tumor serum samples and malignant breast tumor serum samples, and then performing protein identification and quantification to obtain proteomic data of three serum samples respectively; Step 2, the proteomic data of the three serum samples are subjected to quality control, preprocessing, correction and screening to obtain differential proteomic data 1; Step 3: The differential proteomic data 1 is combined with the blood routine indexes and then classified, and the markers are screened based on the training set and the test set.
4. The preparation method according to claim 3, It is characterized in that The threshold value of the preprocessing is: The percentage of proteins commonly identified in the three serum samples was ≥50%.
5. The preparation method according to claim 3, It is characterized in that The screening criteria are: Among the three serum samples, the expression value of the specific serum protein in any lesion serum sample is ≥ 1.5 times the expression value of the other two serum samples; and the specific serum protein is in the top 10% with the largest difference among multiple groups of differentially expressed proteins; The arbitrary lesion serum samples include benign breast tumor lesion serum samples and malignant breast tumor serum samples.
6. The preparation method according to claim 3, It is characterized in that The pre-processing steps include: The five most abundant serum proteins in serum were removed using the High-Select Top 14 Protein Removal Kit and then treated for 17 h at a trypsin:protein mass ratio of 1:
25.
7. The preparation method according to claim 3, It is characterized in that The classification method is an extreme value gradient enhancement algorithm.
8. The preparation method according to claim 3, It is characterized in that The conditions for the quality test are: Chromatographic column: C18, particle size 1.9 μm; inner diameter 150 μm, length 15 cm; pore size The chromatographic conditions are: Mobile phase A: 0.1% formic acid and water; Mobile phase B: 0.1% formic acid and acetonitrile; Flow rate: 600 nL / min; The gradient range and duration were: 15% to 30% mobile phase B, separation 75 min.
9. Use of the marker according to claim 1 or 2 or the marker prepared by the preparation method according to any one of claims 3 to 8 in preparing a breast lesion diagnosis product or establishing a breast disease screening model.
10. Breast tumor diagnosis products, It is characterized in that A marker comprising or using the marker described in claim 1 or 2 or a marker prepared by the preparation method described in any one of claims 3 to 8.
Citation Information
Patent Citations
Mammary gland non-malignant tumor and malignant tumor blood serum special protein and uses thereof
CN101354393A
Serum protein marker for early screening and diagnosis of breast cancer, kit and detection method
CN110716043A
Exosome markers for diagnosing lymph node metastasis of invasive breast cancer and application of exosome markers
CN114428173A
Cancer-related biological materials in microvesicles
US20140045915A1
Methods of assessing breast cancer using circulating hormone receptor transcripts
US20220042109A1
Cited By
Abdominal aortic aneurysm serum protein fingerprint detection and analysis method and system based on LASSO regression algorithm
CN121142060A
Protein marker, prediction method and system for predicting iodine uptake capability of metastasis of cancer A
CN121768676A