A biomarker combination and its application in predicting breast cancer risk

By detecting the protein and mRNA expression levels of breast cancer patients using a combination of biomarkers, and constructing a predictive model using a generalized linear regression model, the problem of insufficient sensitivity and accuracy in the early diagnosis of breast cancer was solved, and efficient early cancer screening was achieved.

CN119410775BActive Publication Date: 2025-09-02SHANGHAI AIPUTIKANG BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411548471.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-09-02
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

Existing technologies lack sufficient sensitivity and accuracy in the early diagnosis of breast cancer, making it difficult to achieve efficient early cancer screening.

Method used

A combination of biomarkers, including ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1, and UROD, was used to detect the expression levels of these proteins or mRNAs. A predictive model was then constructed using a machine learning algorithm based on a generalized linear regression model to predict breast cancer risk.

Benefits of technology

It achieves high sensitivity and high specificity in the early diagnosis of breast cancer, with an area under the ROC curve (AUC) of 0.96-1.00. The diagnostic sensitivity and specificity are both above 90%, providing convenience for early clinical diagnosis and intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119410775B_ABST
    Figure CN119410775B_ABST
Patent Text Reader

Abstract

The present invention discloses a biomarker combination and its application in predicting breast cancer risk, specifically disclosing the application of a biomarker combination in preparing a product for predicting breast cancer, wherein the biomarker combination is composed of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD. The marker combination provided in the present invention can be used as risk estimation and detection for breast cancer patients, has the advantages of high sensitivity and high specificity, and provides favorable technical support for predicting the occurrence and development of breast cancer. The development of corresponding prediction devices using the above-mentioned marker combinations can provide great convenience for early clinical diagnosis or intervention treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biological detection technology, and specifically relates to a biomarker combination and its application in predicting breast cancer risk. Background Art

[0002] In 2020, breast cancer overtook lung cancer as the most common cancer worldwide. Clinically, early detection of lesions can be achieved through ultrasound, CT scans, and biopsy. Regular self-examination can also help. However, these methods suffer from high false-positive rates and delayed detection. Therefore, a highly sensitive and accurate diagnostic method is urgently needed for early cancer screening.

[0003] Proteomics has played a significant role in revealing the complex molecular events of tumorigenesis, such as tumorigenesis, invasion, metastasis, and treatment resistance. Proteomic tumor diagnosis, with its advantages of high sensitivity, strong specificity, and well-defined underlying mechanisms, has been increasingly used for tumor detection in recent years. However, the research on these tumor markers is often based on a limited amount of experimental data, involving relatively limited cancer types and sample sizes. With the continuous development of the proteome in recent years, big data on body fluid proteomes has continued to increase. Therefore, it is urgently needed to collect body fluid proteome data and utilize big data analysis methods to develop a widely applicable and highly accurate tumor risk model to achieve early diagnosis and treatment for patients. Summary of the Invention

[0004] To address the above technical issues, the present invention provides a biomarker combination and its application in predicting breast cancer risk. The protein molecular marker combination of the present invention exhibits significant differences in expression levels in body fluid samples from cancer patients and healthy controls. Therefore, the marker combination provided by the present invention can be used for risk estimation and detection in breast cancer patients, offering the advantages of high sensitivity and specificity, providing advantageous technical support for predicting the development and progression of breast cancer. Developing a corresponding prediction device using this marker combination could greatly facilitate early clinical diagnosis or therapeutic intervention.

[0005] In order to solve the above technical problems, the first aspect of the present invention provides an application of a biomarker combination in the preparation of a product for predicting breast cancer, wherein the biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD.

[0006] In a preferred embodiment, the product for predicting breast cancer is a product for predicting early-stage breast cancer.

[0007] In a second aspect, the present invention provides a reagent for detecting a biomarker combination, wherein the biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD.

[0008] In a preferred embodiment, the reagent is used to detect the expression level of the biomarker combination; the expression level is the protein expression level and / or the mRNA transcription level.

[0009] In a preferred embodiment, the reagent is a biomolecule reagent that specifically binds to the biomarker or specifically hybridizes with the nucleic acid encoding the biomarker.

[0010] In a preferred embodiment, the biomolecule reagent is selected from the group consisting of primers, probes and antibodies.

[0011] In a preferred embodiment, the reagent is a reagent for genome, transcriptome and / or proteome sequencing.

[0012] A third aspect of the present invention provides a use of a reagent for detecting a biomarker combination in preparing a kit for predicting breast cancer; wherein the biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD.

[0013] In a preferred embodiment, the reagent is as described in the first aspect of the present invention.

[0014] In a preferred embodiment, the kit for predicting breast cancer is a kit for predicting early stage breast cancer.

[0015] A fourth aspect of the present invention provides a biomarker combination, which consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD.

[0016] The fifth aspect of the present invention provides a kit, which comprises the reagents as described in the second aspect of the present invention or the biomarker combination as described in the fourth aspect of the present invention.

[0017] A sixth aspect of the present invention provides a method for predicting breast cancer for non-diagnostic purposes, the method comprising detecting the expression level of a biomarker combination in a sample to be tested;

[0018] wherein the biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1, and UROD;

[0019] The expression level is protein expression level and / or mRNA transcription level.

[0020] In a preferred embodiment, the prediction of breast cancer is the prediction of early stage breast cancer.

[0021] A seventh aspect of the present invention provides a breast cancer risk prediction system, the prediction system comprising a detection module and an analysis and judgment module; the detection module detects the expression level of the biomarker combination in the sample to be tested and transmits the expression level data to the analysis and judgment module;

[0022] The analysis and judgment module processes the expression level data using Firmiana software, where the expression level data is preferably FOT (Fraction of total, defined as the iBAQ of the protein divided by the total iBAQ of all identified proteins in the sample), and is preset to a machine learning algorithm based on a generalized linear regression model to construct a prediction model to predict the probability of the sample having breast cancer and the probability of not having breast cancer, respectively, and to determine whether the expression level data meets a preset judgment condition to predict the risk of the sample having breast cancer, and output a prediction result; the judgment condition is that the probability of having breast cancer is greater than or equal to the probability of not having breast cancer;

[0023] When the expression level data meets the judgment condition, the prediction result is output as "having a risk of breast cancer"; when the expression level data does not meet the judgment condition, that is, the probability of having breast cancer is less than the probability of not having breast cancer, the prediction result is output as "not having a risk of breast cancer";

[0024] The biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD; and the expression level is protein expression level and / or mRNA transcription level.

[0025] In some embodiments of the present invention, the prediction system is used to process the expression level data through Firmiana software after the receiving or input is completed, and a machine learning algorithm based on a generalized linear regression model is preset to construct a prediction system.

[0026] In a preferred embodiment, the sample to be tested is a plasma sample.

[0027] In a preferred embodiment, the prediction system also includes a data collection module, which is used to collect expression level data of the biomarker combination in the sample breast tissue, and the expression level data is preferably FOT (Fraction of total, defined as the iBAQ of the protein divided by the total iBAQ of all identified proteins in the sample).

[0028] In a preferred embodiment, the prediction system is a system for predicting early breast cancer.

[0029] The eighth aspect of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the functions of the prediction system as described in the seventh aspect of the present invention, or implement the steps of the method as described in the sixth aspect of the present invention.

[0030] The ninth aspect of the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the functions of the prediction system as described in the seventh aspect of the present invention, or to implement the steps of the method as described in the sixth aspect of the present invention.

[0031] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.

[0032] The reagents and raw materials used in the present invention are commercially available.

[0033] The positive progress effect of the present invention is:

[0034] The protein molecular marker combination of the present invention exhibits significant expression differences in body fluid samples from cancer patients and healthy controls. Therefore, the marker combination provided by the present invention can be used for risk assessment and detection of breast cancer patients, with the advantages of high sensitivity and specificity, providing advantageous technical support for predicting the development and progression of breast cancer. The development of a corresponding prediction device using this marker combination could greatly facilitate early clinical diagnosis or therapeutic intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the area under the ROC curve of the marker combination in the training set.

[0036] Figure 2 Schematic diagram of the area under the ROC curve of the marker combination in the test set.

[0037] Figure 3 Schematic diagram of the area under the ROC curve of the marker combination in the external validation set.

[0038] Figure 4 Schematic diagram of the system for predicting breast cancer risk.

[0039] Figure 5 A schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION

[0040] The present invention includes plasma samples from 114 normal subjects and 101 breast cancer patients. The design and implementation of this study have been approved and supervised by ethics committees, and written informed consent has been obtained from all patients.

[0041] 1. Separation of plasma

[0042] Whole blood samples were collected in EDTA anticoagulant tubes, mixed by inversion, and centrifuged at 1,600 × g for 10 min in a 4°C low-temperature centrifuge. The supernatant (plasma) was collected into new EP tubes and centrifuged at 16,000 × g for 10 min to remove cell debris. The plasma was aliquoted into centrifuge tubes and frozen at -80°C for later use.

[0043] 2. Plasma sample pretreatment

[0044] To 2 μL of plasma sample, 100 μL of 50 mM ammonium bicarbonate was added and vortexed for 1 minute. The sample was heated at 95°C for 4 minutes to denature the protein. After cooling to room temperature, 2 μg of trypsin was added to the system. The system was shaken at 37°C for 18 hours, and then 10 μL of ammonia was added to stop the enzymatic hydrolysis. The peptide samples after enzymatic hydrolysis were desalted, dried, and frozen at -80°C until mass spectrometry analysis.

[0045] 3. Mass spectrometry detection of plasma samples

[0046] The Orbitrap Fusion Lumos three-in-one high-resolution mass spectrometry system (Thermo Fisher Scientific, Rockford, USA) was used in conjunction with a high-performance liquid chromatography system (EASY-nLC 1200, Thermo Fisher) to obtain mass spectrometry data of the whole protein corresponding to the peptide sample. The specific operation was as follows:

[0047] Nanoflow liquid chromatography was used, and the chromatographic column was a homemade C18 column (150 μm ID×8 cm, 1.9 μm / The column oven temperature was 60°C. The dry powdered peptide was reconstituted in loading buffer (0.1% formic acid in water) and applied to the column for separation. Elution was performed at 600 nL / min using a linear 6-30% mobile phase B (ACN and 0.1% formic acid). A 10-min liquid phase gradient was used with data-independent acquisition (DIA) mass spectrometry detection. DIA mass spectrometry parameters were as follows: positive ionization mode; primary mass spectrometry resolution of 30K, maximum injection time of 20 ms, AGC target of 3e6, scan range of 300-1400 m / z; secondary scan resolution of 15K, acquisition of 30 variable isolation windows, and collision energy of 27%. Data acquisition was performed on a liquid chromatography-tandem mass spectrometry system controlled by Xcalibur software.

[0048] 4. Data Analysis

[0049] All data were processed using Firmiana. Firmiana is a workflow based on the Galaxy system, consisting of multiple functional modules such as user login interface, raw data, identification and quantification, data analysis and knowledge mining. The DIA data were searched using DIANN (v12.1) against the UniProt human protein database (updated on 2019.12.17, 20406 entries). The mass difference of the parent ion is 20ppm, and the mass difference of the daughter ion is 50mmu. A maximum of two missed cleavage sites are allowed. The search engine sets cysteine ​​carbamidomethylation as a fixed modification and methionine N-acetylation and oxidation as variable modifications. The parent ion charge range is set to +2, +3 and +4. The false discovery rate (FDR) is set to 1%. The results of the DIA data were merged into the reference library using SpectraST software. A total of 327 libraries were used as reference libraries.

[0050] The quantitative results of the identified peptides were recorded as the average of the peak areas of the chromatographic fragment ions in all reference spectral libraries. Protein quantification was performed using the label-free intensity-based absolute quantification (iBAQ) method. We calculated the peak area values ​​as a fraction of the corresponding protein. The total fraction (FOT) was used to represent the normalized abundance of a specific protein in the sample. The FOT was defined as the iBAQ of the protein divided by the total iBAQ of all identified proteins in the sample. Proteins with at least one unique peptide and a 1% FDR were selected. The FOT of each protein was calculated and input into the generalized linear regression model as protein expression data.

[0051] The Firmiana algorithm selected in this example is a machine learning algorithm based on a generalized linear regression model. A prediction model is constructed to predict the probability of a sample having breast cancer and the probability of not having breast cancer. The code for constructing the prediction model is:

[0052] from sklearn.linear model import LogisticRegressionCV

[0053] import joblib

[0054] import pandas as pd

[0055] tumor_types = ['***','***','***','***','***'] #*** refers to the 37 protein molecular markers described in this invention

[0056] for tumor_type in tumor_types:

[0057] df_train=pd.read_csv('{}_train.csv'.format(tumor_type))

[0058] df_test=pd.read_csv('{}_test.csv'.format(tumor_type))

[0059] df_val=pd.read_csv('{}_validation.csv'.format(tumor_type))

[0060] X_train = df_train.iloc[:,2:]

[0061] y_train = df_train.iloc[:,1]

[0062] X_test = df_test.iloc[:,2:]

[0063] y_test=df_test.iloc[:,1]

[0064] X_val = df_val.iloc[:,2:]

[0065] y_val = df_val.iloc[:,1]

[0066] model=LogisticRegressionCV()

[0067] model.fit(X_train,y_train)

[0068] joblib.dump(model,f'{tumor_type}_model.joblib');

[0069] The receiver operating curve (ROC) was drawn to calculate the area under the ROC curve (AUC) for the relative expression levels of 37 protein molecular markers (ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD) in the plasma samples of breast cancer patients. Curve), where the training set includes 49 positive cases and 56 negative cases, AUC = 1.00, diagnostic sensitivity 100.00%, and specificity 100.00% (see Figure 1 The test set included 21 positive cases and 24 negative cases, with an AUC of 0.98, a diagnostic sensitivity of 97%, a specificity of 100%, a positive predictive value of 94%, and a negative predictive value of 100% (see Figure 2 The external validation set included 31 positive cases and 34 negative cases, with AUC = 0.96, diagnostic sensitivity 96.77%, and specificity 82.35%. (See Figure 3For analysis methods, see Karimollah Hajian-Tilaki, Receiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation, Caspian J Intern Med 2013; 4(2): 627-635. The FOT values ​​of 37 protein markers in the training set, test set, and external validation set are shown in Tables 1-3, 4-6, and 7-9.

[0070] For unknown samples, the expression levels of the above biomarkers are substituted into the model to obtain the breast cancer risk prediction of the sample and output the result. When the probability of having breast cancer is greater than or equal to the probability of not having breast cancer, the output prediction result is "with breast cancer risk"; when the probability of having breast cancer is less than the probability of not having breast cancer, the output prediction result is "without breast cancer risk".

[0071] From the above results, it can be seen that the combination of 37 protein molecular markers in the plasma of cancer patients can be used to predict tumor risk.

[0072] Table 1 FOT values ​​of 37 protein markers in the training set

[0073]

[0074]

[0075]

[0076] Table 2 FOT values ​​of 37 protein markers in the training set

[0077]

[0078]

[0079]

[0080] Table 3 FOT values ​​of 37 protein markers in the training set

[0081]

[0082]

[0083]

[0084]

[0085] Table 4 FOT values ​​of 37 protein markers in the test set

[0086]

[0087]

[0088] Table 5 FOT values ​​of 37 protein markers in the test set

[0089]

[0090]

[0091] Table 6 FOT values ​​of 37 protein markers in the test set

[0092]

[0093]

[0094] Table 7 FOT values ​​of 37 protein markers in the external validation set

[0095]

[0096]

[0097] Table 8 FOT values ​​of 37 protein markers in the external validation set

[0098]

[0099]

[0100]

[0101] Table 9 FOT values ​​of 37 protein markers in the external validation set

[0102]

[0103]

[0104] Example 2 System for Predicting Breast Cancer Risk

[0105] The system 61 for predicting breast cancer risk includes a data processing module 52 and a judgment and output module 53, and also includes a data collection module 51 ( Figure 4 ).

[0106] The data collection module 51 is used to collect the expression level data of the biomarker combination in the patient's breast cancer tissue sample and transmit it to the data processing module.

[0107] The data processing module 52 is configured to analyze the received or inputted biomarker combination expression level data using the data analysis method described in Example 4 to obtain a calculation result. The biomarker combination expression level data can be collected by the data collection module 51 or obtained from other sources.

[0108] The judgment and output module 53 is used to judge whether the calculation result meets the preset judgment condition, that is, the probability of breast cancer risk is greater than or equal to the predicted probability of no breast cancer risk, so as to predict the breast cancer risk and output the prediction result; wherein, in the judgment and output module, when the expression level data meets the judgment condition that the probability of breast cancer risk is greater than or equal to the predicted probability of no breast cancer risk, the prediction result is output as "having breast cancer risk"; when the expression level data does not meet the judgment condition that the probability of breast cancer risk is less than the predicted probability of no breast cancer risk, the prediction result is output as "not having breast cancer risk".

[0109] Example 3 Electronic Equipment

[0110] This embodiment provides an electronic device, which can be expressed in the form of a computing device (for example, a server device), including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for predicting breast cancer risk in Embodiment 1 of the present invention can be implemented.

[0111] Figure 5 The hardware structure diagram of this embodiment is shown. The electronic device 4 specifically includes:

[0112] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including the processor 91 and the memory 92), wherein:

[0113] The bus 93 includes a data bus, an address bus, and a control bus.

[0114] The memory 92 includes a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read-only memory (ROM) 923 .

[0115] Memory 92 also includes a program / utility 925 having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0116] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92, such as the data analysis method of embodiment 1 of the present invention.

[0117] The electronic device 9 can further communicate with one or more external devices 94 (e.g., a keyboard, pointing device, etc.). Such communication can be performed via an input / output (I / O) interface 95. Furthermore, the electronic device 9 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 96. The network adapter 96 communicates with other modules of the electronic device 9 via a bus 93. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 9, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.

[0118] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0119] Example 4 Computer-readable storage medium

[0120] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for predicting breast cancer risk in embodiment 1 of the present invention are implemented.

[0121] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0122] In a possible implementation, the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of the method for predicting breast cancer risk in Example 1 of the present invention.

[0123] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.

[0124] Finally, the above specific implementation method is only used to illustrate the technical solution of the present invention, rather than to limit it.

[0125] Full name of biomarker (refer to genecards database)

[0126] ADPRS: ADP-Ribosylserine Hydrolase

[0127] ARMC8: Armadillo Repeat Containing 8

[0128] ASCC3:Activating Signal Cointegrator 1Complex Subunit 3

[0129] B3GAT3: Beta-1,3-Glucuronyltransferase 3

[0130] BLVRB: Biliverdin Reductase B

[0131] CA1:Carbonic Anhydrase 1

[0132] CDC37: Cell Division Cycle 37

[0133] DLG1: Discs Large MAGUK Scaffold Protein 1

[0134] DSC3: Desmocollin 3

[0135] DUSP3: Dual Specificity Phosphatase 3

[0136] EIF5AL1: Eukaryotic Translation Initiation Factor 5A Like 1

[0137] EWSR1: EWS RNA Binding Protein 1

[0138] GCNT3:Glucosaminyl(N-Acetyl)Transferase 3

[0139] H3-4:H3.4 Histone

[0140] H3C1:H3 Clustered Histone 1

[0141] H3C15:H3 Clustered Histone 15

[0142] HBB:Hemoglobin Subunit Beta

[0143] HBE1:Hemoglobin Subunit Epsilon 1

[0144] HNRNPA0:Heterogeneous Nuclear Ribonucleoprotein A0

[0145] HSPA6:Heat Shock Protein Family A(Hsp70)Member 6

[0146] ITGA6:Integrin Subunit Alpha 6

[0147] MSL1:MSL Complex Subunit 1

[0148] NUP155:Nucleoporin 155

[0149] PDCD10:Programmed Cell Death 10

[0150] PLBD2:Phospholipase B Domain Containing 2

[0151] PXDN:Peroxidasin

[0152] RDH10:Retinol Dehydrogenase 10

[0153] RECK:Reversion Inducing Cysteine Rich Protein With Kazal Motifs

[0154] RUVBL1:RuvB Like AAA ATPase 1

[0155] SAMHD1:SAM And HD Domain Containing Deoxynucleoside TriphosphateTriphosphohydrolase 1

[0156] SLC30A7:Solute Carrier Family 30 Member 7

[0157] SMARCA5:SWI / SNF-Related Matrix-Associated Actin-Dependent RegulatorOf Chromatin Subfamily A Member 5

[0158] SPG21:Spastic Paraplegia 21

[0159] TCIRG1:T Cell Immune Regulator 1

[0160] TRIOBP:TRIO and F-actin binding protein

[0161] UFC1:Ubiquitin-Fold Modifier Conjugating Enzyme 1

[0162] UROD:Uroporphyrinogen Decarboxylase。

Claims

1. A reagent for detecting the protein expression level of a biomarker combination in a plasma sample, characterized in that: The biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1, and UROD.

2. The reagent according to claim 1, wherein The reagent is a biomolecule reagent that specifically binds to the biomarker combination; and / or, the reagent is a reagent for proteome sequencing.

3. The reagent according to claim 2, wherein The biomolecule reagent is an antibody.

4. A kit, characterized in that The kit comprises the reagent according to any one of claims 1 to 3.

5. Use of the reagent according to any one of claims 1 to 3 or the kit according to claim 4 in preparing a product for predicting breast cancer.

6. A breast cancer risk prediction system, characterized in that: The prediction system includes a detection module and an analysis and judgment module; the detection module detects the expression level of the biomarker combination in the sample to be tested and transmits the expression level data to the analysis and judgment module; the analysis and judgment module processes the expression level data through Firmiana software, which is preset to a machine learning algorithm based on a generalized linear regression model, to construct a prediction model, respectively predict the probability of the sample having breast cancer and the probability of not having breast cancer, determine whether the expression level data meets the preset judgment conditions, so as to predict the risk of the sample having breast cancer, and output a prediction result; the judgment condition is that the probability of having breast cancer is greater than or equal to the probability of not having breast cancer; When the expression level data meets the judgment condition, the prediction result is output as "having a risk of breast cancer"; when the expression level data does not meet the judgment condition, that is, the probability of having breast cancer is lower than the probability of not having breast cancer, the prediction result is output as "not having a risk of breast cancer"; Among them, the biomarker combination consists of ADPRS, ARMC8, ASCC3, B3GAT3, BLVRB, CA1, CDC37, DLG1, DSC3, DUSP3, EIF5AL1, EWSR1, GCNT3, H3-4, H3C1, H3C15, HBB, HBE1, HNRNPA0, HSPA6, ITGA6, MSL1, NUP155, PDCD10, PLBD2, PXDN, RDH10, RECK, RUVBL1, SAMHD1, SLC30A7, SMARCA5, SPG21, TCIRG1, TRIOBP, UFC1 and UROD; and the expression level is the protein expression level.

7. The prediction system according to claim 6, wherein: The sample to be tested is a plasma sample; and / or, the prediction system further includes a data collection module, which is used to collect expression level data of the biomarker combination in the plasma sample.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the functions of the prediction system according to claim 6 or 7 can be realized.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: The processor is configured to execute the computer program to implement the functions of the prediction system according to claim 6 or 7.

Citation Information

Patent Citations

  • Biomarkers for breast cancer detection

    CA3206126A1

  • Application of biomarker combination in preparation of kit for predicting lymphoma

    CN117051112A