Protein combination for breast cancer diagnosis and application thereof

CN122793992APending Publication Date: 2026-09-22FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610982640.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

这些方法虽然有效,但存在局限性:影像学检查可能漏诊早期病变或产生假阳性;组织活检是侵入性的,且需要较长时间获取结果

Benefits of technology

(1)模型完全可书面固化:逻辑回归存在明确线性数学公式,所有系数、截距固定,无黑箱算法,专利权利要求、试剂盒配套算法可直接完整落地;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122793992A_ABST
    Figure CN122793992A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of breast cancer diagnostic markers. The present application provides a protein combination for breast cancer diagnosis and its application, which comprises one or more of the corresponding UniProt numbers P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, P02679. The present application uses an interpretable weighted logistic regression model to complete non-invasive screening of breast cancer, which has significant technical effects, and has significant advantages in high diagnostic accuracy, non-invasive detection, early diagnosis ability, rapid detection process, etc., and has wide clinical application value. The method of the present application not only improves the efficiency and accuracy of breast cancer diagnosis, but also provides important support for the development of personalized treatment and precision medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of breast cancer diagnostic biomarkers, and more particularly to a protein combination for breast cancer diagnosis and its application. Background Technology

[0002] Breast cancer is one of the most common cancers among women worldwide, and early diagnosis is crucial for improving cure and survival rates. Traditional diagnostic methods include imaging studies (such as mammography, ultrasound, and MRI) and biopsies (pathological examination). While effective, these methods have limitations: imaging studies may miss early lesions or produce false positives; and biopsies are invasive and require a long time to obtain results.

[0003] Biomarkers (such as peptides, proteins, genes, and metabolites) play a crucial role in cancer diagnosis. Proteins, as executors of gene function, directly reflect cellular state. Protein expression levels are closely related to the occurrence and development of cancer. Previous studies have shown that certain characteristic peptides corresponding to specific proteins are abnormally expressed in breast cancer patients and can serve as diagnostic biomarkers.

[0004] Peptidomics is the study of peptide expression, modification, and function. In recent years, the development of high-throughput mass spectrometry has made large-scale screening of cancer-specific peptides possible. By comparing peptide expression differences between breast cancer patients and healthy individuals, specific diagnostic biomarkers can be screened. Single peptide diagnostic efficacy is limited, with insufficient sensitivity or specificity; combined statistical models of peptide segments can significantly improve discriminative power. Existing tree-based machine learning models have a black-box structure, making it difficult to formalize patents and solidify reagent kit algorithms. Logistic regression, on the other hand, possesses a complete, writable linear mathematical formula with fixed parameters, facilitating translation and implementation. Therefore, developing a non-invasive, highly discriminative, and fully documented early breast cancer screening program has significant clinical value. Summary of the Invention

[0005] The purpose of this invention is to provide a protein combination for breast cancer diagnosis and its application. By using a weighted logistic regression model with non-collinear characteristic peptides corresponding to 8 UniProt numbers, non-invasive early screening of breast cancer can be achieved. The model formula can be fully incorporated into the patent and supporting analysis software, making up for the shortcomings of existing technology models that are not interpretable and have insufficient screening accuracy.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution: This invention provides a protein combination for breast cancer diagnosis, the protein combination including one or more of the corresponding UniProt numbers P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679.

[0007] The present invention also provides a breast cancer diagnostic model based on the aforementioned breast cancer early screening diagnostic biomarkers. The breast cancer diagnostic model uses the absolute quantitative mass spectrometry value of the peptide corresponding to each UniProt number as input features and outputs the probability of breast cancer risk in the range of 0 to 1. The formula for calculating the linear logit of the model is as follows: logit(P)=-3.164881+0.113829×P05546- 0.036906×P0DJI8+0.007089×P19823+1.011715×P35858-1.449817×Q9Y6Z7+1.599488×P33908-0.596567×P00738+0.215522×P02679; Formula for converting disease probability: P=1 / (1+exp(-logit(P))) The optimal diagnostic threshold is 0.142, and the judgment rule is: if the model output probability P≥0.142, it is judged as breast cancer positive, and if P<0.142, it is judged as healthy negative; the model training uses a binary classification logistic loss function, with 100 iterations and a sample balancing weight of 1.

[0008] The present invention also provides the application of the aforementioned protein combination as a diagnostic biomarker for breast cancer.

[0009] The present invention also provides the application of the aforementioned protein combination in the preparation of products for breast cancer diagnosis.

[0010] Preferably, the product is a reagent kit.

[0011] The present invention also provides a kit for the diagnosis of breast cancer, the kit comprising a targeted mass spectrometry quantitative reagent for detecting the corresponding characteristic peptides of the eight UniProt.

[0012] Beneficial effects: (1) The model can be fully documented: Logistic regression has a clear linear mathematical formula, all coefficients and intercepts are fixed, there is no black box algorithm, and the patent claims and the matching algorithm of the reagent kit can be directly and completely implemented; (2) Stable performance: AUC=0.758, accuracy 0.688, sensitivity 0.738, specificity 0.675 on independent external test set, with better levels of missed diagnosis and misdiagnosis than single protein / peptide markers; (3) Good generalization: The difference between the AUC of the training set and the test set is only 0.073, which is within the stable range without overfitting, and the results of different clinical sample cohorts fluctuate little; (4) Non-invasive testing: Only serum samples are required, without the need for invasive / radiation examinations such as breast biopsy and mammography, resulting in high compliance among the population for screening; (5) Elimination of feature interference: Before modeling, redundant peptides with a Pierre correlation coefficient > 0.6 were removed. The peptides corresponding to the 8 UniProt peptides do not have collinearity cancellation, and the combined efficacy is stable. (6) It is compatible with standardized mass spectrometry platforms. The kit and its supporting process can be automated for batch detection, making it suitable for large-scale physical examinations and early screening.

[0013] When using this patented model for detection, the serum peptide data is first uniformly QC cleaned, and the quantitative values ​​of the corresponding peptides P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679 are extracted. These values ​​are then substituted into the complete LR formula mentioned above to calculate the probability of disease. The screening results are output based on a threshold of 0.142. The entire process does not require retraining the model and can be used directly in batches. Attached Figure Description

[0014] Figure 1 This is the combined ROC curve of the training set and independent test set for the optimal weighted logistic regression model of this invention. Detailed Implementation

[0015] This invention provides a protein combination for breast cancer diagnosis, the protein combination including one or more of the corresponding UniProt numbers P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679.

[0016] The present invention also provides a breast cancer diagnostic model based on the aforementioned breast cancer early screening diagnostic biomarkers. The breast cancer diagnostic model uses the absolute quantitative mass spectrometry value of the peptide corresponding to each UniProt number as input features and outputs the probability of breast cancer risk in the range of 0 to 1. The formula for calculating the linear logit of the model is as follows: logit(P)=-3.164881+0.113829×P05546-0.036906×P0DJI8+0.007089×P19823+1.0117 15×P35858-1.449817×Q9Y6Z7+1.599488×P33908-0.596567×P00738+0.215522×P02679; Formula for converting disease probability: P=1 / (1+exp(-logit(P))) The optimal diagnostic threshold is 0.142, and the judgment rule is: if the model output probability P≥0.142, it is judged as breast cancer positive, and if P<0.142, it is judged as healthy negative; the model training uses a binary classification logistic loss function, with 100 iterations and a sample balancing weight of 1.

[0017] The present invention also provides the application of the aforementioned protein combination as a diagnostic biomarker for breast cancer.

[0018] The present invention also provides the application of the aforementioned protein combination in the preparation of products for breast cancer diagnosis.

[0019] In this invention, the product is preferably a reagent kit, and more preferably a liquid chromatography-tandem mass spectrometry (LC-MS / MS) quantitative reagent kit.

[0020] The present invention also provides a kit for the diagnosis of breast cancer, the kit comprising a targeted mass spectrometry quantitative reagent for detecting the corresponding characteristic peptides of the eight UniProt.

[0021] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0022] Example 1: Screening and Validation of Protein Biomarkers

[0023] Training set: 241 healthy individuals, 63 breast cancer patients; Independent validation test set: 160 healthy individuals, 42 breast cancer patients; The AUC of the training set was calculated for each single peptide (corresponding to a single UniProt protein) and sorted in descending order of AUC. High collinearity redundant peptides with a Pierre correlation coefficient ≥0.6 were progressively removed. Modeling was performed by traversing different combinations of peptide numbers and sample weights. Finally, P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, P02679, with a weight of 1 and a threshold of 0.142, were determined to be the optimal biomarker combination of this invention.

[0024] Sample collection: A total of 506 serum samples were collected from breast cancer patients and healthy controls. Among them, 241 serum samples from healthy controls and 63 serum samples from breast cancer patients were used as the training set for model training, and another 160 serum samples from healthy controls and 42 serum samples from breast cancer patients were used as an independent prediction set to evaluate the model performance.

[0025] Biomarker Validation: Independent clinical validation samples (including disease group and control group) were collected. Total protein was extracted from the samples, and peptides were obtained through enzymatic digestion, desalted, and purified for later use. Subsequently, corresponding stable isotope-labeled relabeled peptides were added to each sample peptide as internal standards to correct for systematic errors in sample pretreatment and subsequent detection. A 6500 QTRAP mass spectrometer was used. After optimizing detection conditions, the spectrometer was used to collect ion response signals of eight UniProt target peptides and their corresponding relabeled peptides. A standard curve was plotted using the peak area ratio of the lightly labeled peptide to the heavily labeled peptide as the standard, and the absolute concentrations of the target peptides and their corresponding proteins were calculated. The obtained data were then compiled for subsequent modeling and analysis.

[0026] Example 2: Construction of the Diagnostic Model

[0027] A weighted logistic regression model was used, and the complete analysis and prediction process was completed using R language. First, a subset of breast cancer and healthy control samples was selected. Sample labels were transformed, peptide feature matrices and diagnostic labels were separated, and the AUC values ​​of each peptide were calculated and sorted in descending order to select the optimal peptide. After optimization, a weighted logistic regression model was finally constructed to detect breast cancer. The specific model settings and application parameters are as follows: Objective function: Binary classification logistic loss function Sample balancing weights: 1 Number of iterations: 100 rounds Optimal decision threshold: 0.142.

[0028] The optimal markers used in this model correspond to the following UniProt ID combinations: P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679.

[0029] Solidify the complete mathematical formula: logit(P)=-3.164881+0.113829×P05546-0.036906×P0DJI8+0.007089×P19823+1.0117 15×P35858-1.449817×Q9Y6Z7+1.599488×P33908-0.596567×P00738+0.215522×P02679 P=1 / (1+exp(-logit(P))) The optimal diagnostic threshold was determined to be 0.142 using the Youden index method.

[0030] Independent test set validation metrics: test AUC=0.758, accuracy 0.688, sensitivity 0.738, specificity 0.675; training-test AUC difference 0.073, indicating model stability and no overfitting.

[0031] Instructions for use: Substitute the mass spectrometry quantitative values ​​of the peptides corresponding to the 8 UniProt numbers in the sample into the formula to calculate the P value. A P value ≥ 0.142 indicates high risk of breast cancer, and further imaging / pathological diagnosis is recommended; a P value < 0.142 indicates low-risk healthy individuals.

[0032] Example 3: Operating Procedure of the Matching Mass Spectrometry Reagent Kit

[0033] Thaw the serum, then digest it at 37°C with the kit's enzyme digestion buffer and the corresponding UniProt isotope internal standard. Reagent kit packing material for desalting and purifying peptide mixtures; Eight target peptide ion peaks corresponding to UniProt were acquired using LC-MS / MS MRM mode, and the absolute quantification of each protein was calculated. The quantitative values ​​are substituted into the patented fixed LR formula to calculate the probability of disease and output the screening results.

[0034] Interpretation of diagnostic results

[0035] Quantitative analysis: Quantitative data were generated by detecting the expression levels of eight UniProt protein peptides in the sample.

[0036] Results scoring: A breast cancer discrimination model was constructed using the saved weighted logistic regression. The model parameters were: objective function was a binary classification logistic loss function, sample balancing weight was 1, number of iterations was 100, and the optimal judgment threshold was set to 0.142. The calculated P-value was used to determine whether the sample belonged to the high-risk group for breast cancer.

[0037] As shown in the above embodiments, this invention provides a protein combination for breast cancer diagnosis and its application. The protein combination includes one or more of the following UniProt numbers: P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679. This invention employs an interpretable weighted logistic regression model to perform non-invasive breast cancer screening, achieving significant technical results. It demonstrates significant advantages in high diagnostic accuracy, non-invasive detection, early diagnostic capability, and rapid detection process, and has broad clinical application value. The method of this invention not only improves the efficiency and accuracy of breast cancer diagnosis but also provides important support for the development of personalized treatment and precision medicine.

[0038] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A protein assembly for breast cancer diagnosis, characterized in that, The protein combination includes one or more of the following UniProt numbers: P05546, P0DJI8, P19823, P35858, Q9Y6Z7, P33908, P00738, and P02679.

2. A breast cancer diagnostic model based on the breast cancer early screening diagnostic biomarker as described in claim 1, characterized in that, The breast cancer diagnosis model uses the absolute quantitative mass spectrometry values ​​of the peptides corresponding to each UniProt number as input features and outputs the probability of breast cancer risk in the range of 0 to 1. The formula for calculating the linear logit of the model is as follows: logit(P)=-3.164881+0.113829×P05546- 0.036906×P0DJI8+0.007089×P19823+1.011715×P35858-1.449817×Q9Y6Z7+1.599488×P33908-0.596567×P00738+0.215522×P02679; Formula for converting disease probability: P=1 / (1+exp(-logit(P))) The optimal diagnostic threshold is 0.142, and the judgment rule is: if the model output probability P≥0.142, it is judged as breast cancer positive, and if P<0.142, it is judged as healthy negative; the model training uses a binary classification logistic loss function, with 100 iterations and a sample balancing weight of 1.

3. The use of the protein combination of claim 1 as a diagnostic biomarker for breast cancer.

4. The use of the protein combination of claim 1 in the preparation of products for breast cancer diagnosis.

5. The application according to claim 4, characterized in that, The product in question is a reagent kit.

6. A reagent kit for breast cancer diagnosis, characterized in that, The kit contains a targeted mass spectrometry quantitative reagent for detecting the eight UniProt characteristic peptides corresponding to claim 1.